Build an AI layer for procurement where every claim re-adds and every agent knows what it may read.
Ten governed implementations, each shipping its read scope written down and its write paths set to NONE. The fixtures carry a savings register that re-adds correctly and is still wrong, bids restated on bases that do not match, and three turns of pressure to drop an inconvenient finding.
No autonomous award. No autonomous release above the threshold. No autonomous sign-off.
Most AI guidance treats access as plumbing. In procurement it is the control itself. One agent that can read every live bid can carry one bidder's price into another bidder's clarification, and nobody in the room will see it happen. That is not a data-loss incident with a report to file. It is a tender you cannot defend.
So every agent ships with its read scope written down and its write paths set to NONE, the estate register records what each one may read before its first real use, and one scenario exists purely to try to make an agent reach outside its scope.
of the 100 classified activities are held back from automation, 24 never automate and 4 human only. Seven of them cost something, and gates that cost nothing are decoration.
Agent-led with review
AI-assisted
Straightforward automation
Never automate
Human only
The seven gates that cost something
The award decision. Purchase order release above the delegation-of-authority threshold. Goods receipt certification. The sanctions hit determination. Supplier onboarding approval. Supplier bank detail changes. Inspection release.
Those are not late-roadmap items. Every one of the ten implementations is written to stop short of them, and each build sheet names the gate it stops at. Nothing in this pack prepares a safety authorization either: AI assembles the evidence, a competent person qualifies and signs.
Built for the failure modes that matter.
A savings register that ties out perfectly is the most convincing way to certify a double count. An agent that checks the total and stops has done the damage.
A bid tabulation that ranks two bids on different bases hands the award to whoever excluded the most. A supplier dossier that reads complete on a company with no public filings is worse than a blank page, because someone will sign against it.
The fixtures do not stop at tidy inputs
A register that re-adds correctly and is still wrong. One saving claimed twice across two initiatives. A baseline quietly restated against the same source document. A currency mismatch with no rate. Bids where one excludes freight and duty and another annualizes a nine-month volume. An index-linked increase the published index does not support. And three turns of pressure to drop an inconvenient finding.
PASS means the agent CAUGHT the inconsistency and named it. Explaining it confidently is a fail.
the amount one savings register in the fixtures is out by, while re-adding perfectly to its own stated total. If your agent returns a clean tie-out on that file, you have learned something about the agent rather than about the register.
One fixture where correct arithmetic produces the wrong answer
A three-way-match exception clears a 2.0 percent tolerance and breaches the policy as written by 10.00, because the policy says whichever limb is lower. An agent that clears it shows its working while doing so, which is what makes the trap gradeable rather than merely unfair. You can see exactly where it stopped reading.
You do not take anyone's word for it, which is the point.
The test protocol ships as its own document: every scenario, the fixture it runs on, and what a PASS has to produce. Ten of the scenarios are clean inputs, so you can see the shape of a good answer before you start breaking things.
Build one agent, grade it against its key, and you have a written result to put in front of the person who asks why you trust it.
Scenarios, P-T1 to P-T56
Poisoned inputs
Three-turn pressure sequences
Controls where PASS is printing the zero
Read-scope probe
7 numeric fixtures carrying 90 pre-computed equations66 restated in the keys
Grading is arithmetic you re-add by hand, not impression. Every expected figure is worked out for you: input tie-outs, normalized bid columns, baseline differences, unexplained remainders, index calculations, tolerance limbs.
10 answer keys ship with the packnot behind a support ticket
Never paste one into the agent's chat. The key is the ruler, and an agent that has seen the expected answer proves nothing. You read the key, the agent does not.
9 of the 56 are controlsPASS is the zero
The finding is genuinely absent and a PASS means printing the zero instead of manufacturing something to look useful. Over-reporting destroys the credibility of the real findings, which is why the controls are in the protocol rather than left to judgement.
Run WF-1 on your own build before first real use
Then re-run it after any instruction edit and after platform changes. A fix without a named fixture re-run does not count, and the workflow is the calendar entry that keeps that honest.
Five parts: the governed builds, and the method they run inside.
83 files in one zip, 62 PDFs across 190 pages, 20 paste-ready text files and one workbook. The zip separates the vendor-neutral methodology (core/) from the implementation written for Microsoft 365 Copilot (microsoft/), so the classification, the gates and the test protocol stay usable if you build somewhere else.
10 governed implementationsa build sheet and a spec each
Leakage and Tail Spend Analyst, Bid Normalizer, Savings Claim Validator, Negotiation Prep Pack Builder, Supplier Risk Dossier Builder, Contract Obligation Extractor, Supplier Scorecard Assembler, Price Increase Challenger, Benchmark and Should-Cost Challenger, Three-Way Match Exception Analyst. Paste-ready instruction files with character counts, so a truncated paste cannot masquerade as a working agent. Each one takes you from a blank Agent Builder form to a run you have graded yourself, and each one carries its read scope and its write paths in the file rather than in a policy nobody opens.
Fixtures and answer keys included90 equations
Seven numeric fixtures with every expected figure pre-computed, 66 of them restated in the ten answer keys so you re-add them by hand. The poisoned inputs are the point: a double count, a moved baseline, a currency mismatch with no rate, bids on incompatible bases, an index-linked increase the index does not support.
100 procurement activities classified5 modes
Into five operating modes: 49 agent-led with review, 15 AI-assisted, 8 straightforward automation, 24 never automate, 4 human only. Every row carries the reason, not just the verdict. 30 of the 100 are rows where the arithmetic has to reconcile to a stated baseline.
50 engineered prompts, 12 gated workflows
A decision matrix that gives the next AI idea a lane instead of a debate, a nine-section governance pack with mechanisms instead of posters, 13 working templates, an AI estate register pre-filled for all ten agents, and a 30-day sprint with a gate at the end of every week.
A 20-page Playbook PDF and a 15-tab workbook2 formats
The 100-row map with your own posture, priority and owner columns, the paste shapes the agents expect for a spend extract, a bid tabulation, a savings register, a match input and performance counts, your procurement profile with its tail threshold and delegation levels, the pre-filled estate register, an agent test log with blank columns for your own run dates, and the value scorecard your CPO reads at day 30.
AI prepares. Humans decide. Read scope is a control.
That is the whole methodology in one line, and every gate in the pack is downstream of it. The agents assemble, normalize, challenge and package. A named person awards, releases, certifies and signs.
Deploy estimates come from the same build pattern on comparable agents. They are estimates for one experienced buyer following the build sheet, and they exclude your own prep: extract shapes, the contract coverage list, baseline methodology, tolerance thresholds. The prep is usually the longer half, and the pack says so. Time your first build and replace the estimate with your own number.
Same zip. Different usage rights.
Both editions ship the same zip, file for file: you are choosing usage rights, not features. Licensed, not sold.
What you need first: a Microsoft 365 Copilot licence with Agent Builder. This is the instruction layer for a platform you already pay for. It does not include, replace or discount the licence.
- Deploy the agents in a tenant you work in
- Adapt everything for what you personally do
- No team redistribution, no resale, no public republication
- Share the pack across your procurement function
- Multiple practitioners can build from it
- No resale, no public republication, no delivering it as paid training
Same zip either way. Click Buy, then confirm Individual or Organization Edition before you pay.
Pair it with the agent controls and the tenant baseline.
The Stack is the build layer. These two are the surface it runs on: what an agent is allowed to reach, and how the tenant around it is set. Each stands on its own.
Securing Agents in Microsoft 365 Copilot
A control-by-control verification kit for the agent surface itself: what an agent can reach, who published it, and how you prove it.
The Copilot Hardening Baseline
176 checks for the Microsoft 365 Copilot admin surface. Each one names the baseline to set and how to prove it.