The AI that pays your invoices, and knows exactly when not to.
AP Autopilot extracts, matches, and risk-scores every invoice automatically. But the decision to move money never touches a language model. It passes through a deterministic policy engine, and every step lands in a hash-chained ledger your controller can replay months later.
Manual AP is slow. A chatbot with a bank login is worse.
Manual AP
- Invoices sit in an inbox for days waiting on someone to notice them
- PO matching is a person cross-referencing three systems by hand
- Approval hierarchies live in someone’s memory, not in any system
- By the time fraud is noticed, the money is already gone
A model with a bank login
- The same reasoning that reads the invoice also decides whether to pay it
- A cleverly worded PDF is one prompt injection away from an approval
- “Trust the agent’s judgment” is not an audit trail a controller can sign off on
- When something goes wrong, the only record is a conversation transcript
AP Autopilot separates reasoning from authority: no process that reads an invoice can also pay one.
Seven steps. One of them a language model never touches.
Every invoice becomes one durable Temporal workflow the moment it arrives. It survives a crash, a deploy, a three-day wait on an approver, and resumes exactly where it left off.
Ingest
A durable workflow starts. Nothing after this point can be lost to a restart.
Extract
Vendor, amount, PO number, bank details parsed out of the document.
Validate
Checked against the accounting system’s PO / goods-receipt records, and against our own history for duplicates.
Score risk
Vendor history, bank-detail changes, round-number and split-invoice patterns, scored deterministically.
Gate
A Go rule engine decides: auto-approve, escalate, or reject. Same inputs, same answer, every time.
no LLM past this pointPay / wait
Cleared invoices are paid immediately. Everything else waits, durably, for a named human.
Record
Every step is appended to a hash-chained ledger. Nothing here can be quietly edited later.
The full decision flow: branches, waiting, and all.
Traced directly from the workflow’s own source. Rounded boxes are automated steps, the six-sided box is the one decision point, and filled pills are where an invoice’s story ends. The gold dot marks every step that also gets appended to the hash-chained audit ledger.
Run a real invoice through this exact path on /demo.
Everything above the gate can be probabilistic. The gate itself and everything below it (approve, escalate, pay) is deterministic Go code with no model in its call stack.
Built to plug into your stack, not replace it.
Every one of these is reached through the Model Context Protocol behind one stable tool contract, a swappable pattern we’ve built and run against sandboxed equivalents, not a claim of official partnership. Change providers without touching the workflow logic that calls them.
The same dashboard that runs this pitch runs your invoice mix.
These four numbers come from actually running a 26-invoice adversarial corpus through the live system, not a hand-picked demo run.
Typical time to a decision
“Manual AP” reflects a typical email-and-spreadsheet approval cycle, not a measured benchmark. AP Autopilot’s figure is measured, for the subset of invoices that clear autonomously. Everything else still moves to a human in seconds, not days.
Bounded autonomy isn’t a prompt. It’s an architecture.
Exactly one component in the entire system holds credentials to the payment rail, and it has no language model anywhere in its call stack.
Policy decisions run through a small, deterministic, unit-tested Go service, auditable without reading a single token of model output.
Every event is append-only and hash-chained to the one before it. We tested this by corrupting our own database. It caught it instantly.
Every external system is reached over MCP behind a stable tool contract. Change providers without touching the workflow that calls them.
Average invoice amount, prior bank details, submission cadence: it all persists per vendor and feeds every future risk score. A first invoice from a vendor is scored differently from the fortieth, and a bank-detail change against that memory is exactly what trips the gate.
The infrastructure choices, and why each one is there.
Durable execution: a three-day approval wait survives a worker restart with zero custom recovery code.
The policy gate. Deterministic, fast, and easy to unit-test in complete isolation.
Every external tool call (accounting, payments, Slack) behind one uniform, swappable contract.
The append-only ledger, tenant isolation via row-level security, and every read model the dashboard shows.
The live dashboard your team actually watches: approval queue, audit explorer, trust metrics.
Built end to end. Tested end to end.
Every number on this page came from actually running the system, including a bug we found and fixed by watching it hang, not from a slide deck. If you’re evaluating AP automation and want to see how the pieces actually fit together, get in touch.