-
-
A memo dated March says the installer is in good standing. The live check found a Chapter 11 filing from April.
-
Two documents price the same expansion $25M apart — 13.4%, above the workspace threshold.
-
Every result the search read, with why each was used or discarded.
-
Every decision signed into a PDF by Nutrient DWS, with a SHA-256 over the signed bytes.
-
8 claims extracted from this page by Nutrient DWS — 4 produced findings, 4 read cleanly.
-
Same screen, both themes.
-
Which domains the searches returned, and which the policy accepts.
Inspiration
A firm is three days from wiring money for a 250 MW solar portfolio. The deal sits on a folder of documents, and two of them matter.
Their own investment memo prices the expansion at $186M. The independent engineer's report says $211M.
Nobody put those two sentences side by side. They live on different pages, in different files, written by different people in different words: one says "installation cost is estimated at", the other says "total expansion installation cost". So the committee plans around $186M, and $25M of the gap lands on their equity after signing.
That same memo, dated March 20th, records the installer as "in good standing". True when written. On April 15th the installer filed for Chapter 11. The memo did not change. In June it still says good standing, and nothing inside the document will ever tell you otherwise.
Both failures are ordinary. Neither is anyone's fault. They are what happens when the amount of paper exceeds the number of people who can read it. And when the loss surfaces months later, there is a second cost: nobody can prove who knew what. The concern lived in an email thread. Under the current wave of disclosure mandates, "we reviewed the documents" is no longer a sentence a regulator or an auditor accepts without proof.
Summarizing documents with a model does not fix any of this. A summary has no page numbers, no per-field confidence, no check against the outside world, and no signature. What a credit committee accepts is evidence: which field, which page, which live source, which reviewer, which hash. So we built the evidence machine instead.
What it does
Upload two documents. About ten seconds later you have a list of what disagrees with what, what the world has since made false, and what nobody can verify at all. Every claim lands in one of five states: CORROBORATED, CONFLICTING, STALE, REVIEW_REQUIRED, UNVERIFIED, and every state comes with its evidence attached.
Document against document. Nutrient DWS reads each PDF three ways at once: the tables, the text layer, and key-value pairs that carry a native confidence score per field. Sparkline normalizes differently worded passages onto one canonical claim type, so the two cost sentences finally meet. On our demo bundle it pulls 16 claims out of two documents and reports the conflict as $186M against $211M: a $25M variance, 13.4%, high materiality, with both source sentences quoted and their pages cited.
Document against reality. Claims no second document could ever settle go to SerpApi. The memo says the installer is in good standing; the live check returns the installer's Chapter 11 petition of April 15, 2026, sourced from the bankruptcy court's own claims agent. Stale, critical. The interesting part is what the engine refuses: only authoritative domains can carry a verdict, and every result the evaluator looked at stays on screen with the reason it was accepted or rejected. The claims agent accepted. The law firm accepted. The Reddit thread rejected as non-authoritative. The review aggregator rejected because it predates the filing. The same single search also corroborates a different claim, so the engine visibly distinguishes agreement from conflict rather than just reporting "found a hit".
A person, and a signature. Every finding lands beside its source page in the embedded Nutrient viewer, with the evidence face-off and the full query trace. A human approves or rejects. That decision is rendered to a PDF through DWS conversion, digitally signed with DWS, and written to a ledger that records the SHA-256 of the signed bytes. Anyone can recompute the digest and check the row. Six months later, when the bankruptcy reaches the portfolio, there is a signed, dated record of who saw the flag and what they decided.
How we built it
Next.js 16, React 19, TypeScript, Tailwind v4, Node 22 with its built-in test runner.
We did not rebuild the hard parts. Nutrient DWS does four jobs: extraction with per-field confidence, Markdown-to-PDF conversion for the review record, the digital signature on it, and, through the Nutrient Web SDK, the document viewer, running as WASM in the browser from static assets with no session token and no server round trip. SerpApi does one job carefully: the live public-record check, behind an authoritative-domain source evaluator we wrote.
What we own is the judgment in between: the claim registry that normalizes wording onto canonical types, the comparator that decides what counts as a conflict (a numeric gap wider than 0.5% of the primary figure), the source evaluator, and the trust score that blends extraction confidence with cross-document agreement.
The pipeline is instrumented rather than animated. Every stage transition and every reasoning line is written to a run record as it happens, so the analyzing screen polls a real run instead of playing a timer. A pure adapter turns a stored run into the view model the screens render, so committed fixtures and live runs travel through exactly the same code. All provider calls happen on the server, so no key ever reaches the browser.
One full analysis costs about 7 DWS credits and 1 SerpApi search; a signed decision costs 2 more DWS operations. Identical queries are cached in-process, so two claims about the same counterparty share one search. The queries themselves were not guessed: we ran a three-run stability protocol over candidate public records and threw out the CAISO interconnection queue, whose status field truncates unpredictably in snippets, before locking the counterparty-status query.
Challenges we ran into
The rule that shaped everything: never claim more than the evidence. It would have been easy to make the demo look smarter, and we kept deleting the places where it did. A claim nothing can settle is reported UNVERIFIED, never quietly passed. A run whose live check is refused reports no trust score at all rather than a flattering one computed from the checks that did run, because the missing check is exactly the one that would have pulled the number down.
We shipped a worse number because it was the true one. We found the trust dial displaying a weighted blend, 0.4 × 0.88 + 0.6 × 0.62 = 0.724, that no code in the repository actually performed. The real formula is the product the scorer computes. Fixing it dropped the demo's score from 72 to 55, and the app now prints 0.62 × 0.88 = 0.55 directly under the dial. A product whose entire argument is that documents state things which do not survive checking cannot put numbers on screen that do not survive checking.
The hardest bug was on the audit trail itself. Merging ledger rows keyed on flag id silently deleted countersignatures. One flag carries two records, the decision and the approver's endorsement of it, so they collided in a Map and one was dropped. It only appeared once a real signature was written to disk, which is why every demo looked right. The arrival of an unrelated signature was deleting the row recording who endorsed a different decision. On an audit trail, that is the worst possible loss, and nothing anywhere would have said so.
What we learned
Confidence you can audit beats confidence you assert. Showing the counts behind every score, the results behind every verdict, and the bytes behind every signature changed this product more than any feature did. We also learned that the discipline is subtractive: the honest version of this product is mostly the flashy version with the unsupported claims removed.
What's next
Object storage so uploads and signed records survive a serverless deploy. Keeping the bounding boxes DWS already returns, so highlights sit on the PDF itself rather than a text rendition. Claim registries beyond this vertical: the router and comparator are generic, the nine claim types are not. More verification routers: permits, corporate registries, court dockets. A countersignature as its own signed record. And re-running on a schedule, so a document that goes stale is caught before someone opens it rather than when they do.
Where Nutrient DWS does the heavy lifting: it reads every document three ways with per-field confidence, renders each human decision to a PDF, signs it, and puts the viewer a reviewer judges from in the browser. Sparkline's guarantees are DWS guarantees.
Where SerpApi does the heavy lifting: it settles the claims no second document can, catching the drift between what a document says and what is true now, with every accepted and rejected source kept as evidence.
Built With
- api
- digital-signatures
- markdown
- next.js
- node.js
- nutrient
- nutrient-dws
- nutrient-web-sdk
- react
- serpapi
- sha-256
- tailwind-css
- typescript
- webassembly

Log in or sign up for Devpost to join the conversation.