Independent benchmark · Finance workflows · No paid placement
Independent benchmarks for AI on real finance work — and what it actually costs.
LedgerRate scores AI models on real finance work — invoice extraction, ledger reconciliation, AP triage, ad-hoc reporting — and publishes the true cost alongside every score. Accuracy ranks; cost never buys a better position.
Edition 1 is in the works
Benchmark editions publish monthly — a ranked model index, cost-vs-quality charts for every finance workload, a plain-English rate card, and the raw logs behind every number. Subscribe to get the first one.
How LedgerRate scores
A model's score on a workload is its pass rate against deterministic, published pass criteria — reported with a 95% confidence interval, on the verbatim prompt every model receives. Rankings use scores alone.
CPO = cost per attempt ÷ score
Beneath every score sits the economics: CPO, the expected cost of one acceptable result. A cheap model that fails often costs more per outcome than an expensive one that doesn't. If a model never passes a workload, its CPO is DNF — reported as such, never as a number.
Every edition runs on published, versioned methodology (read it) with raw logs released. No paid placement, ever (independence policy).
The benchmarks
Task designs and verbatim prompts →Real finance tasks, not trivia: each workload publishes its exact prompt, deterministic pass criteria, and a rate-card unit a finance team can price.
- Order-to-cash apply.cash.v1
Apply 20 customer payments to an open AR ledger of 30 invoices: multi-invoice payments, valid and invalid early-pay discounts, short-pays, overpays, loosely formatted remittance memos, and payments referencing unknown invoices (15 scenarios).
Apply 1,000 remittances - Controls & audit audit.expenses.v1
Audit 40 expense lines against a company T&E policy whose limits vary per scenario — meal per-diems, receipt thresholds, hotel nightly caps, mileage rates, prohibited categories, and EUR conversions — flagging 12 planted violations by rule id, with boundary-exact amounts as clean traps (15 scenarios).
Audit 1,000 expense lines - Credit & treasury compute.covenant.v1
Compute LTM Adjusted EBITDA, net leverage, and fixed charge coverage from four quarters of financials under a credit agreement's own definitions — capped restructuring add-backs (greater/lesser of a fixed amount and a percentage), litigation add-backs allowed or excluded, cash netting capped or prohibited — and state compliance with each covenant (15 borrower scenarios).
Test a quarter's covenant package - Procure-to-pay extract.invoice.v1
Extract structured data (vendor, dates, currency, line items, totals) from 15 synthetic invoices with planted edge cases: multi-page, credit note, foreign currency, discount lines.
Process 1,000 invoices - Procure-to-pay match.threeway.v1
Three-way match 40 supplier invoices against purchase orders and goods receipts, flagging 12 planted exceptions by type: price variance beyond tolerance, quantity over receipt, missing receipt, and item mismatch — with boundary-exact tolerances, split receipts, and partial invoicing as clean traps (15 scenarios).
Match 1,000 invoices - Data & reporting query.nl2sql.v1
Answer 15 natural-language finance questions of graded difficulty by writing SQL against a synthetic finance database (GL, AP, AR, dimension tables).
Answer 100 ad-hoc finance questions - Record-to-report reconcile.ledger.v1
Reconcile a month of 200 bank transactions against a shuffled general-ledger extract containing 12 planted discrepancies: missing entries, duplicates, amount mismatches, and date drift (15 monthly scenarios).
Reconcile a month of transactions - Procure-to-pay triage.ap-inbox.v1
Sort 15 batches of 20 synthetic accounts-payable emails into invoice, statement, dunning notice, PO response, or other.
Sort 1,000 AP emails
Deploying AI in your own workflows?
The benchmark tells you what AI costs per successful outcome; the guides cover everything around that number — vendor-neutral, for finance teams:
- How to evaluate AI for a finance workflow — a pilot playbook
Before you buy or build, run a two-week evaluation that measures what actually matters — cost per acceptable result, not demo quality.
- Where AI fits in the finance function — seven workflow patterns
A map of the finance jobs AI can do today, organized by the shape of the work rather than the hype around it — and where each pattern is proven versus experimental.
- Controls and risk management for AI in finance workflows
How to deploy AI in a controlled finance environment — review design, audit trails, data handling, and change management that will survive your auditors.
- Understanding AI pricing — from tokens to cost per outcome
A finance professional's guide to how AI is actually billed — tokens, caching, batch tiers — and the arithmetic that converts a price sheet into a budgetable cost per unit of work.
- The vendor question checklist — ten questions that cut through an AI sales call
Bring these to the demo. The questions that separate measurable products from confident decks, with the answers you should expect to hear.
Newsletter: email to subscribe (signup form coming soon).