Deduction triage and trade-promotion reconciliation for CPG brands.
A retailer pays a $17,000 invoice and sends $15,400. The remittance says MCB-107
and nothing else. Somebody has to work out whether that $1,600 was a promotion the
brand actually agreed to run — and if it wasn't, find the signed deal sheet that
proves it and file the dispute before the window closes.
Deductions run 2–15% of gross sales for most CPG brands. A meaningful share are invalid. Almost none of it is recovered, because the work is manual, the documents are scanned faxes, and the deadlines pass quietly.
Shortpay reads the remittance, reconciles every deduction against the trade calendar and the delivery paperwork, and produces a queue ordered by what is actually worth working today.
All data here is synthetic. Harvest Lane Foods is not a real company and no figure is derived from any real brand's data.
A deduction is a decaying asset. It is recoverable only until its dispute window closes — 180 days at UNFI and KeHE, 90 at Whole Foods. After that the money is gone no matter how strong the evidence. Most tools present a list sorted by size. This one weights by how close the deadline is, so a $900 claim closing in nine days outranks a $2,400 claim with four months left. In the demo dataset 20 deductions worth $11,048 have already expired — a quarter of all recoverable money, lost to working the backlog in the wrong order.
A dispute you cannot document is worth nothing. Distributors adjudicate on the backup you attach, not on the argument you make. A shortage claim needs a signed delivery receipt; a promotional chargeback needs the signed deal sheet. So every deduction is scored on whether the decisive document is actually retrievable, and a claim without it is discounted to 22% of face value. That drops undocumentable disputes down the queue automatically instead of letting an analyst discover the problem an hour in.
Together these give the ranking:
expected recovery = amount in dispute × win probability × evidence readiness
priority = expected recovery × deadline weight
remittance PDF
├─ vision extraction ──────── reads scans with no text layer
├─ text-layer parser ──────── exact and free where a text layer exists
↓
invoice matching ────────────── tolerates OCR damage, refuses ambiguous matches
↓
reconciliation
├─ promotions ─────────────── against the signed deal calendar: window, SKU, rate
├─ shortages ──────────────── against the signed delivery receipt
├─ fill-rate penalties ────── against the amended purchase order
├─ pricing claims ─────────── against the effective cost sheet
├─ cash discounts ─────────── against the payment date
└─ duplicates ─────────────── same charge twice on one remittance
↓
scoring ─────────────────────── validity, evidence readiness, days to deadline
↓
ranked queue + dispute packet
The dispute packet names the exact documents that partner requires for that deduction type, marks which are on file, and drafts the narrative with the specific dates and amounts filled in.
The dataset is generated with a known defect planted in each invalid deduction. Nothing in the pipeline reads those labels — they exist only so the result can be measured rather than asserted.
python -m backend.app.evaluate
Reconciliation — 214 deductions
| Precision | 93.9% |
| Recall | 90.6% |
| F1 | 0.922 |
| Recoverable dollars identified | $46,820 of $49,442 — 94.7% |
Recall by defect type:
| Defect planted | Recall |
|---|---|
| Duplicate charge | 8/8 · 100% |
| Promotion billed outside its window | 12/12 · 100% |
| Shortage contradicted by delivery receipt | 12/12 · 100% |
| Fill-rate penalty against an amended PO | 3/3 · 100% |
| Compliance chargeback with no case reference | 8/8 · 100% |
| Cash discount taken outside terms | 3/3 · 100% |
| Advertising charge with no authorisation | 3/3 · 100% |
| Rate billed above the signed rate | 2/2 · 100% |
| SKU never on the deal sheet | 22/24 · 92% |
| Price claim contradicted by cost sheet | 4/5 · 80% |
| Unsaleables with no disposition record | 0/5 · 0% |
Extraction — deterministic parser, 121 deductions across the 19 documents that have a text layer: 100% precision and recall on (invoice, reason code, amount) triples.
Extraction — vision, measured on 8 documents (2 of them scans with no text
layer) via OpenRouter against anthropic/claude-sonnet-4.5:
| Documents read | 8/8 |
| Precision / recall | 98.2% / 98.2% |
| Deduction triples | 55 correct, 1 missed, 1 spurious |
| Cost | $0.2145 total — ~$0.027/document |
| Latency | 13.5s/document |
Both scanned documents were read perfectly (6/6 and 7/7 deductions). The single
error across all 8 was a decimal misread — 58.80 transcribed as 58.00 — and the
arithmetic cross-check caught it without being told, flagging an 80-cent gap on both
the invoice line and the document total:
HL-4510: gross 12,704.40 less deductions 1,365.60 = 11,338.80,
but net paid reads 11,338.00
line deductions total 2,808.40 but the document total reads 2,809.20
That is the argument for the cross-check. A model that reads 55 of 56 figures correctly is useful; a model that reads 55 of 56 correctly and flags the one it got wrong is something a finance team can actually run unattended.
Spoilage scores zero and that is reported rather than hidden. Deciding whether an unsaleables claim is legitimate needs disposition records from the distributor's side, which this brand does not hold. A rule that read back its own planted label would lift the headline number and detect nothing real. It is a data-access problem, not a modelling one.
Real remittances are not a standard form. Each partner emits its own layout, column order and reason-code vocabulary, and many arrive as scanned faxes or portal print-outs.
The generator renders three distinct layouts, then degrades about 40% of them into image-only PDFs — rotated a fraction of a degree, speckled with toner noise, contrast-flattened and JPEG-recompressed. Those files carry zero extractable characters. The rule-based parser returns nothing on them; only the vision path can read them, which is exactly the situation a brand is in when a distributor faxes its remittance.
On the documents that do have a text layer, both extractors run. The text layer is effectively ground truth there, so it gives a continuous accuracy check on the vision path in production — no hand-labelled evaluation set required.
Every extraction also gets an arithmetic cross-check: gross minus deductions must equal net paid on every line. A mismatch means either the read is wrong or the document is internally inconsistent, and both are worth surfacing.
The 33 PDFs in data/remittances/ are fixtures. They are committed for two
reasons only: a fresh clone runs with no setup, and the evaluation is reproducible
against known ground truth. They are not a document store, and real remittances
would never belong in version control.
Anything that arrives afterwards is runtime data. POST /api/upload — or the drop
zone on the Extraction page — takes a PDF, writes it to data/inbox/ (gitignored;
object storage in production), and runs the same extract-and-reconcile path:
curl -X POST http://localhost:8000/api/upload \
-F "file=@some-remittance.pdf" -F "mode=vision"Two things change once a document is genuinely new.
There is no ground truth. For a fixture we know what was printed, so accuracy is measurable. For a document seen for the first time nothing can be checked against — so the arithmetic cross-check (gross − deductions = net paid) becomes the only automated signal that the read was correct. The upload response surfaces it.
The document is only half the input. Reconciliation compares a deduction to the brand's own records — the invoice, the signed deal sheet, the delivery receipt. A remittance for invoices the system does not hold extracts perfectly and reconciles to nothing, and the response says so explicitly rather than returning an empty queue:
known invoices: 0 unrecognised: ZZ-90000, ZZ-90001, ZZ-90002 …
5 invoice(s) on this remittance are not in our records, so their deductions
cannot be reconciled against a deal sheet or delivery receipt.
That same test showed why the vision path matters beyond scans: the rule-based
parser reads zero lines from that document, because its invoice regex is hardcoded to
this brand's HL- format. The model read all five ZZ- invoice numbers without
being told the format changed.
Requires Python 3.11+ and Node 18+.
git clone <this repo> && cd shortpay
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m backend.app.seed.generate # build the dataset
python -m backend.app.seed.remittance_pdf # render + degrade the PDFs
cd frontend && npm install && npm run build && cd ..
uvicorn backend.app.main:app --port 8000Open http://localhost:8000.
Without a key the app still runs: digital documents parse deterministically, scanned ones report that they need vision, and the whole reconciliation layer is unaffected.
Copy .env.example to .env — it is loaded automatically at startup and is
gitignored.
cp .env.example .envAnthropic first-party:
ANTHROPIC_API_KEY=sk-ant-...OpenRouter — the Anthropic SDK appends /v1/messages, so the base URL is the
surface root with no trailing /v1:
ANTHROPIC_API_KEY=sk-or-v1-...
ANTHROPIC_BASE_URL=https://openrouter.ai/api
SHORTPAY_MODEL=anthropic/claude-sonnet-4.5GET /api/health reports which model and endpoint are actually in use, so a
misconfigured route shows up immediately rather than at the first extraction.
.envoverrides your shell. dotenv's default is the other way round, but some hosts exportANTHROPIC_BASE_URLglobally — Claude Desktop does — which would silently beat the file and route to the wrong provider. For a self-contained app the checked-out.envis the configuration surface, so it wins.
The vision extractor asks for structured data two ways. Strict mode uses Anthropic's structured-outputs parameter, which constrains decoding so the reply is guaranteed to satisfy the schema. JSON mode puts the schema in the prompt and validates what comes back, with one repair round-trip if it does not fit.
Strict is better where it exists. Measured against OpenRouter, it does not.
OpenRouter serves POST /api/v1/messages in Anthropic Messages format and accepts
output_config — then ignores it. That is the worst failure mode, because the strict
path deliberately withholds the schema from the prompt on the assumption the API will
enforce it, so the model receives no schema at all and invents its own field names.
Because of that, auto skips the probe entirely when ANTHROPIC_BASE_URL is set and
goes straight to JSON mode. Set SHORTPAY_STRUCTURED_OUTPUTS=on to force the attempt
anyway (it will fail loudly), or off to pin JSON mode everywhere.
The probe state is module-level, not per-extractor. extract() builds a fresh
extractor per document, so a per-instance flag re-probed and re-paid on every single
file — that bug cost a wasted request per document and pushed latency from 13.5s to
71s before it was found.
Frontend development with hot reload:
cd frontend && npm run dev # proxies /api to port 8000backend/app/
domain.py deduction taxonomy, per-partner reason codes,
and the backup document that decides each dispute type
extraction.py vision + text-layer extractors against one schema
reconcile.py matching, promo reconciliation, scoring, dispute packets
evaluate.py accuracy against the planted ground truth
store.py loads the dataset and keeps the reconciled queue warm
main.py HTTP API
seed/
catalog.py the synthetic brand, its SKUs and its promotions
generate.py dataset builder with labelled defects
remittance_pdf.py three layouts + scan degradation
frontend/src/
components/PromoTimeline.tsx the deal window drawn against the ship date
components/Fuse.tsx dispute window burn-down
components/Queue.tsx the ranked work queue
components/Detail.tsx findings, evidence, dispute packet
Reason codes are never normalised away. Each partner's code is mapped onto a
canonical taxonomy for scoring, but the original is preserved and shown, because that
is what has to go on the dispute filing. Codes that don't map fall through to
UNKNOWN rather than being guessed — an unmapped code is itself a finding, since a
deduction nobody can explain is one the partner has to substantiate.
Rate variances recover the variance, not the invoice. When a promotion is billed at $8.40/case against a signed rate of $5.20, the promotion was real and only the $3.20/case overcharge is disputable. Claiming the full amount is how a brand loses credibility with a trade partner.
Invoice matching refuses ambiguity. Scanned documents produce transcription damage, so matching tolerates a dropped prefix or a single misread character — but only when exactly one candidate fits. A wrong match silently attributes a deduction to the wrong order, which is worse than no match.
Win probabilities are priors, not promises. They are the knob a finance team would tune against their own recovery history. They are visible in the UI for exactly that reason.
- Deduction prevention. The same reconciliation run against open orders before they ship would catch the compliance and fill-rate charges that are cheaper to avoid than to dispute.
- Recovery feedback. Every filed dispute has an outcome. Feeding those back turns the hand-set win probabilities into fitted ones per partner and defect type.
- Real portal integration. The packet is assembled but still filed by hand; the distributor portals are the last manual step.
- Evidence retrieval. The evidence store currently records whether a document
exists. Fetching it from the WMS, the 3PL and the deal-sheet archive is what turns
a
blockedclaim into areadyone.