Invoice Document Review Lab

Invoice Document Review Lab

Read Invoice Document Review Lab: invoice practice guidance with evidence limits, practical checks and editable source material.

The existing invoice reference tests controlled writes. This optional upstream lab tests a different question: can a proposed business record survive source checking and operator correction? It extends the same worked engagement; it does not replace its charter or authorize posting.

No customer access is needed. Start with the self-contained practice packet, then open the review surface. In a clone, serve the repository with a static server bound to loopback, for example python3 -m http.server 8123 --bind 127.0.0.1, and open /examples/invoice-exception/document-review/. Do not open the HTML as a local file: module imports require HTTP. Stop the server when finished; serve only a clean public clone, never a directory containing customer files or credentials.

Boundary and mechanism decision

The user is an accounts-payable reviewer. The accepted practice outcome is a source-supported invoice draft or a justified return to the manual queue. Posting, payment, supplier contact, policy changes, identity management and customer data are excluded. The fictional controller retains posting authority. The manual queue is the fallback.

Step Smallest sufficient mechanism Why and limit
Extract fixed labels Regular expressions, baseline v1 Cheap and inspectable; fails on narrative wording and revisions
Interpret variable wording Optional single model call Compare against the baseline; no tools, agent loop, graph or retrieval engine
Validate record shape and source quote Deterministic code Rejects malformed fields and invented quotes; cannot prove a quoted amount is the correct amount
Resolve contradictions and accept a draft Human source review Requires identity, currency, amount and revision checks; a schema pass is insufficient
Post to the ledger Not implemented Use the separate controlled-write reference and target-system authority; never connect this page directly to payments

The domain is deliberately small: source document and revision, proposal, review decision, and practice history. The browser stores only local practice records. Imported model output is untrusted; it cannot execute HTML, call tools or authorize an effect. A decision is terminal for the currently loaded case; reloading a case starts a new attempt and preserves prior history. Pause retains fields. Reject and escalate remain available without accepting an invalid record. Storage failures tell the user to export rather than claim persistence. ARC-004, ARC-005, HUM-001, HUM-002.

Run and compare

From the repository root:

npm run test:document-review
node examples/invoice-exception/document-review/run-evaluation.mjs

The default runs the real fixed-label baseline offline against eight fictional documents. It does not replay a made-up model score. Four development and four qualification cases cover plain labels, prose, unresolved totals, credit notes, instruction injection, missing currency, unreadable OCR and superseded estimates. Inspect sources, candidate and validation, and the separate grader.

For a live comparison, select a structured-output-capable model and set OPENAI_API_KEY in your shell. The command below sends only the eight built-in public synthetic documents to OpenAI, with no retries, eight calls maximum, a 30-second limit per call and 1,200 output tokens per response. It does not read environment files or accept customer documents. Provider charges apply; check current prices before choosing the model. Capture stdout to a file outside the repository if you want to import the report into the browser.

node examples/invoice-exception/document-review/run-evaluation.mjs --live YOUR_MODEL_ID

The adapter uses the Responses API structured-output format, inspected 2026-09-04. Model refusal, incomplete output, HTTP failure and timeout produce a failed case with the manual-queue terminal reason. Validation remains in trusted code after the response. A schema-constrained response is not a correctness guarantee.

The run log binds source, prompt, candidate and grader bytes; it records per-case correctness, unsafe drafts, latency and live token usage. Keep failed outputs. Do not tune on qualification cases and still call them an unseen holdout. These public partitions are inspectable teaching data, not independent qualification evidence. Neither an 8/8 run nor repository CI establishes population accuracy. EVA-001, EVA-002, EVA-003, EVA-006.

Measure the operator, not just the extractor

Try manual entry and proposals in counterbalanced order with different but comparable cases. Record source-reading time, field edits, escalations, abandoned attempts, corrections missed by the reviewer and final record accuracy. The browser exports elapsed time and edit events; elapsed time includes idle time and is not measured savings. Browser storage and imported reports are editable, unverified practice data, not immutable production audit evidence.

Calculate cost per accepted draft using current model prices, observed reviewer time, rework, exception handling, integration and ongoing support. The complete engagement's existing value case is a forecast, not the result of this exercise. If human review costs erase the benefit, keep manual entry or the rules baseline. Do not count escalation as a completed invoice.

Threats, checks and open proof

Failure Prevention and detection Recovery / executable check
Invoice text attempts to authorize payment No payment tool; outputs are data, never instructions Retain manual authority; injection case and forbidden-field test
Plausible but wrong amount Source shown beside draft; exact quote and field checks Correct or escalate; wrong-amount negative control must fail the grader
Missing or conflicting source Explicit escalation policy Manual queue; missing-currency, OCR and conflict cases
Invalid or oversized imported report Local size and source-revision checks; no execution Reject import and keep prior state; browser tests
Duplicate click, pause or storage failure Terminal decision state; disabled controls; visible save status Export history; review-state and browser checks

This is an experimental learning extension, not a deployable service or a model/agent release bundle. A production version still needs target authentication, tenant isolation, permissioned sources, durable state, tamper-evident audit, source freshness, validated model behavior, representative human trials, capacity, operating ownership and the release gates. Keep development-run logs separate from the canonical evaluation report and solution release required for a real model release. The candidate model has no filesystem or grader access; repository maintainers can inspect both, so this is not evaluator isolation against a malicious developer.