A working prototype that automates Acme Corp's four-stage manual invoice workflow (ingest → validate → approve → pay) using Grok as the reasoning engine, with hard business rules enforced in plain Python wherever a mistake would be costly.
Acme's current process — manual extraction, validation against a legacy inventory system, VP approval via email, manual payment — has a 30% error rate and a 5-day turnaround, costing roughly $2M/year.
This prototype collapses that into a single automated pass per invoice (seconds, not days) while keeping a human-auditable trail at every step:
- Every invoice's extracted fields, validation flags, and approval reasoning are printed and retained — no black-box decisions.
- The one rule Acme can't afford the model to get wrong (invoices over $10K need extra scrutiny) is enforced in Python, not left to the LLM to remember or apply consistently.
- Every approval is checked twice: an initial decision, then a self-critique pass before anything is treated as final — the same "sanity check yourself" step a careful human reviewer does.
The data/output/batch_summary.csv produced by the batch runner (below)
is meant to be handed directly to a stakeholder as evidence the system
catches the failure modes Acme is currently losing money on.
data/invoices/* ──▶ ingestion.py ──▶ validation.py ──▶ approval.py ──▶ payment.py
(Grok: messy (Python: check (Grok: reason (mock: pay
text → clean against SQLite + self-critique or log
JSON fields) inventory) before deciding) rejection)
Each stage is a plain Python module with one job and no hidden shared state — every function takes explicit inputs and returns a plain dict, so any stage can be tested, replaced, or run standalone.
| Stage | File | LLM involved? | What it does |
|---|---|---|---|
| 1. Ingestion | ingestion.py |
Yes — Grok | Reads txt/json/csv/xml/pdf, extracts vendor, amount, items, due_date from messy/inconsistent text |
| 2. Validation | validation.py |
No — pure Python | Checks items/quantities against inventory.db; flags unknown items, stock mismatches, zero-stock items, negative quantities |
| 3. Approval | approval.py |
Yes — Grok, twice | Applies the $10K scrutiny rule in code, gets an initial Grok decision, then runs a self-critique pass before finalizing |
| 4. Payment | payment.py |
No | Mock-pays approved invoices; logs the rejection reason otherwise |
main.py wires all four stages together for a single invoice.
batch_test.py runs every sample invoice through the full pipeline and
prints/exports a summary table — the fastest way to see the system's
behavior across every edge case at once.
console_ui.py is purely cosmetic: color-coded console output, no
pipeline logic lives there.
pip install -r requirements.txt
export XAI_API_KEY="xai-..." # required — the pipeline will refuse to start without it
python3 setup_db.py # creates inventory.db (run once, or any time you want a clean slate)Single invoice:
python3 main.py --invoice_path=data/invoices/invoice_1001.txtEvery sample invoice at once, with a summary table + CSV export:
python3 batch_test.py- The $10K scrutiny threshold is enforced in Python, not prompted.
An LLM can be talked into ignoring or misremembering a policy; a
Python
ifstatement can't. The model only ever sees the result of that check ("this invoice requires high scrutiny because..."). - Self-critique before finalizing.
approve_invoice()makes an initial decision, then explicitly asks Grok to double-check its own reasoning against the scrutiny level before it's treated as final. Both the initial and final decisions are kept (_initial_decision) for audit purposes. - Five input formats, not just PDF/text. txt, json, csv, xml, and
pdf are all supported by the same
extract_invoice_fields()call — format-specific parsing is isolated toread_raw_text(), so Grok only ever sees plain text.