The claim
Cair is an agent that plans how a prize or client payment reaches your bank account — it refuses routes that can't deliver, fills W-8BEN line 10 only when a treaty article really applies, computes what lands, and shows the source of every figure. Every plan is a neatlogs trace. The payout route a chatbot can't get wrong.
The 30-second path
No account, no key, no clone.
- Open the live app. Choose Devpost · 7000 · Indonesia, tick Payoneer and Wise, press Plan my payout (about 5 seconds).
- Read the card: Wise REFUSED with Wise's own sentence, Payoneer recommended, W-8BEN line 10 blank, 30% withholding = $2,100.00, and at most $4,900.00 stamped UNCLEAR because Payoneer does not publish its receiving fee.
- Press Confirm on line 10 — the server writes a confirm span into the same neatlogs session.
- Open /verify: the last plans read back from neatlogs over MCP, redacted.
- Open /rules: every fact the agent may state, with its primary source and the date it was read (CSV).
The receipts
All numbers come from stored runs in evidence/; each
scenario in them carries its neatlogs trace id. Definitions: docs/METRICS.md.
| What | v1: one LLM call | v2: Cair |
|---|---|---|
| Recommends a route its own provider rules out | 18 / 48 (every run) | 0 / 48 |
| Contradicts a sourced tax fact | 8 / 48 (every run) | 0 / 48 |
| Says UNCLEAR where no primary source exists | 11–23 / 143 fields | 143 / 143 |
| Figures traceable to a tool output | n/a (no tools) | 359 / 359 |
| What | Result |
|---|---|
Forced failure of the FX tool (--fail fx_quote) |
all 48 plans completed after retry and fallback; 350 / 350 figures sourced |
| Provenance guardrail on sentences it was not tuned on (measured before the misses were fixed; that set is now in-sample) | blocks 62 / 66 violations; blocks 3 / 66 honest sentences |
Cost and speed per plan (bench at commit 0a8fc06, the last change to the planner) |
p50 5.50 s · p95 10.31 s · $0.0033 of LLM |
| Tests | 94 passing, no network or keys; many are named for the defect they pin (e.g. test_amount_in_2000s_is_not_masked_as_a_year) |
| Exhaustive verification | 2,448 scenarios (every payer × country × option set × amount) run through the real planner against the protocol invariants: 0 violations |
| Permission boundary | tests prove no public route returns a key or token, /verify shows no raw amount, and the builder-only origin cannot be forged |
Reproduce it (the real path)
Every number above is produced with the real model, the real ECB rate and the real neatlogs project — there is no
offline or mock mode for the product. DEMO.md §3 has the
commands (scripts/eval.py --planner v2, --fail fx_quote, scripts/bench.py). Separate and clearly not the
product: pytest runs the deterministic core without network or keys, and the CI pipeline runs only that.
Honest limitations
- Three countries are fully sourced (Indonesia, India, the Philippines); anything else is UNCLEAR by design.
- The evaluation measures fidelity to the sourced rules sheet, not tax correctness; the labels and the agent read the same sheet (three cells were re-checked by hand against the live pages).
- Payoneer's receiving fee and conversion margin are unpublished, so "what lands" is an upper bound. Receipts so far are the builder's own (n = 1 replay of a known outcome); forward receipts from real users: n = 0.
- The guardrail is rule-based: expect new phrasings to slip through at roughly the held-out rate, which is why the route, line 10 and every amount are decided by code, not by the model's prose.
Links
Live app · Verify · Rules sheet · Website · Pitch deck · Repository · README