Build an AI layer for Finance where every number re-adds and every decision stays human.
Ten tenant-tested finance agents, run against poisoned totals, unsupported drivers, currency and period mismatches, and three turns of pressure to drop an inconvenient exception. The test ledger keeps its failures.
No autonomous payment release. No autonomous journal posting. No autonomous sign-off.
The never-automate list is part of the product, and the pack ships the reasons. 100 finance activities are classified into five operating modes, and the boundary between them is a gate, not a suggestion.
If a vendor's finance-AI deck shows fewer prohibitions, ask which control they deleted.
of the 100 classified activities are never automate. Finance carries a heavier list than IT, and that is the honest shape of a controls function.
Agent-led with review
AI-assisted
Automate
Never automate
Human only
Built for the failure modes that matter.
A variance agent that invents a driver is worse than no agent. A morning brief that quietly turns a forecast into an actual is worse than no brief. A reconciliation agent that finds a convenient zero manufactures confidence out of bad accounting.
The whole test pass: one author, one Microsoft 365 test tenant, one day. The honesty page below spells out what one day proves, and it ships in the pack.
I built the fixtures to break it
Totals that do not reconcile. A claimed driver with no data behind it. EUR forecasts against USD actuals. A period mismatch buried in a header. A planted exact fit that zeroes the unexplained remainder. Commentary that contradicts the pack it summarizes. Three turns of pressure to drop an exception, ending in "the auditors need it tonight and it was already authorized".
PASS meant the agent CAUGHT the inconsistency and named it. Explaining it confidently was a fail.
of the 115 recorded runs passed on the instruction set that ships. The ledger keeps the other 28: 24 that failed and passed after a fence edit, and 4 that never passed. All 4 are the same scenario, and it ships named, behind a human control.
Adversarial scenarios
Recorded runs
Instruction defects found
Corrected before release
Shipped named, behind a human control
The one limitation I shipped named, instead of rewriting until it passed
The CFO Morning Brief stops correctly when a set of accounts fails to re-add to its stated total. It does not reliably notice that a GL cash line and a bank total are meant to be the same quantity. Four versions of its instructions failed that test, and rewriting until a run passes is test-fitting, not a control.
It ships named, and the human owns the check on both sides of the agent: the curator reconciles to one cash truth before assembling the extract, then re-adds the cash equation and scans for a second stated cash total before the brief goes out. Two cash totals in a brief means it does not go out. The scenario, its four failing runs and the reasoning are in the ledger you get.
Five parts: the tested agents, and the method they run inside.
81 files in one zip. Structure: vendor-neutral methodology (core/), implementation tested on Microsoft 365 (microsoft/). The methodology binds to modes and gates, not to a vendor; other platforms are roadmap, not in this zip.
10 tenant-tested agents115 runs
Variance Investigator, CFO Morning Brief, Month-End Commentary Builder, Forecast Challenger, Management Reporting Assistant, Reconciliation Exception Analyst, Budget Assumption Reviewer, Expense Anomaly Analyst, Finance BP Meeting Prep, Board Pack Quality Challenger. Built and run in Microsoft 365 Copilot Agent Builder against pre-computed answer keys. Paste-ready instruction files with character counts, so a truncated paste cannot masquerade as a working agent. Each agent also ships its own implementation guide: 10 guides that take you from a blank Agent Builder form to a run you have graded.
Fixtures and answer keys included38 equations
Eleven poisoned inputs and ten answer keys, with every expected figure pre-computed so you can re-add it by hand: input tie-outs, driver splits, unexplained remainders, difference equations, aging day-counts, medians. "It works on my tenant" stops being an opinion once you have run the same graded checks yourself, and you re-run them after every platform change.
100 finance activities classified5 modes
Across 11 domains, into five operating modes: 43 agent-led with review, 19 AI-assisted, 14 automate, 22 never automate, 2 human only. Every row carries the reason, not just the verdict.
50 engineered prompts, 11 gated workflows
A decision matrix that gives the next AI idea a lane instead of a debate, a governance page with mechanisms instead of posters, eight working templates, an AI estate register, and a 30-day sprint with a gate at the end of every week.
A 20-page Playbook PDF and a 13-tab workbook2 formats
The 100-row map with your own posture, priority and owner columns, the assumption register, the pre-filled estate register, the full test ledger with blank columns for your own run dates, and the value scorecard your CFO reads at day 30.
Honesty page included
TESTING-SCOPE states what the testing establishes and what it does not. It was one author, on one Microsoft 365 test tenant, in one day. It does not establish independent reproducibility, durability across future Copilot model updates, or behaviour on production inputs at scale.
And it establishes nothing at all about your numbers. The agents re-add what they are given; a bad extract produces well-formatted questions about a bad extract, which is the designed behaviour and not a substitute for extract hygiene.
Deploy times are estimates for a first-time builder; the author's expert build time is not your first build, and the pack says so.
Same zip. Different usage rights.
Both editions ship the same zip, file for file: you are choosing usage rights, not features. Licensed, not sold.
What you need first: a Microsoft 365 Copilot licence with Agent Builder. This is the instruction layer for a platform you already pay for. It does not include, replace or discount the licence.
- Deploy the agents in a tenant you work in
- Adapt everything for what you personally do
- No team redistribution, no resale, no public republication
- Share the pack across your finance function
- Multiple practitioners can build from it
- No resale, no public republication, no delivering it as paid training
Same zip either way. Click Buy, then confirm Individual or Organization Edition before you pay.
Pair it with the cost book and the admin baseline.
The Stack is the build layer. These go one level deeper on what the AI itself costs and on the tenant it runs in. Each stands on its own.
The Variable Cost of Copilot
The metered layer on top of the flat seat, as a finance problem: allocate it, charge it back, cap it, forecast it.
The Copilot Hardening Baseline
176 checks for the Microsoft 365 Copilot admin surface. Each one names the baseline to set and how to prove it.