The round before this one tested whether a product idea described a real problem. It did, and the solution it proposed was not supported: 92% of merchants were very uncomfortable with an AI acting unsupervised, which was the idea's main selling point. The findings went to Claude, which rewrote the document. The rewrite fixed the autonomy problem, and fixing it meant making five more decisions, each one chosen once, written as a statement, and never compared to anything. A person rewriting it would have made them too. This lesson tests the alternatives to all five, across 104 questions in seven simulations. v2's choice came top in none of them.
What you'd normally do
Go with the first idea that seems right. Argue two of the choices out in a meeting, usually the two someone feels strongly about, and let the rest through unexamined because there is no time and no way to check them. The choices nobody argued about are no more likely to be right than the ones they did.
What you bring
Version two of the concept, rewritten against the last round's findings. It reads as a settled plan: draft-and-approve rather than autonomy, invoice chasing first and listings second, a plain-English log, £99 a month with no contract, a read-only trial, and a pain-led message rather than an AI-led one. Each of those is a decision, and the document does not say what else was on the table.
For testing decisions, bring anything that has committed to something:
- A product spec or PRD
- A pricing page or a proposed price change
- A roadmap, which is a decision about order
- A launch plan
- A design with a layout, a default, and a flow already chosen
What comes back
A decision is a point where the plan picked one option. The alternatives are the options it passed over, which usually appear nowhere in the document. A decision and its alternatives make one claim carrying an option set, and the whole set is tested at once.
It compares v2 to what the last round tested, and drops the options that round already ruled out:
Full autonomy and "AI that does your admin" are dropped — round 1 refuted them.
Then it lists the decisions v2 makes and the options it could have taken instead:
flowchart LR
subgraph C["`**Control model**`"]
direction LR
C1["`Read-only`"]:::question ~~~ C2["`Draft and approve · v2`"]:::document ~~~ C3["`Per-task permissions`"]:::question
C4["`Earned autonomy`"]:::question ~~~ C5["`Supervised sending`"]:::question
end
subgraph T["`**Trust safeguard**`"]
direction LR
T1["`Plain-English log · v2`"]:::document ~~~ T2["`Approve every action`"]:::question ~~~ T3["`Money limits`"]:::question
T4["`View-only bank`"]:::question ~~~ T5["`Task switches`"]:::question ~~~ T6["`Undo`"]:::question
T7["`Weekly summary`"]:::question ~~~ T8["`Named human`"]:::question
end
subgraph P["`**Price and trial**`"]
direction LR
P1["`£99/month · v2`"]:::document ~~~ P2["`£49/month`"]:::question ~~~ P3["`£29 plus add-on`"]:::question
P4["`5% commission`"]:::question ~~~ P5["`Free tier`"]:::question ~~~ P6["`30-day trial`"]:::question
end
subgraph H["`**Headline**`"]
direction LR
H1["`Stop moving information · v2`"]:::document ~~~ H2["`The office manager`"]:::question ~~~ H3["`Nothing sent until you send`"]:::question
H4["`Get your invoices paid`"]:::question ~~~ H5["`It drafts, you approve`"]:::question ~~~ H6["`It works, you control`"]:::question
H7["`Done and waiting`"]:::question ~~~ H8["`Finally looked after`"]:::question
end
subgraph A["`**Acquisition message**`"]
direction LR
A1["`Referral, first month free · v2`"]:::document ~~~ A2["`A plumber in Leeds`"]:::question ~~~ A3["`Start in watch-only`"]:::question
A4["`500 sceptical owners`"]:::question ~~~ A5["`Ask your accountant`"]:::question ~~~ A6["`500 who approve it all`"]:::question
A7["`A café owner`"]:::question ~~~ A8["`Trade association`"]:::question
end
class C,T,P,H,A simulation
Five decisions and the options under each. Blue is what the document chose. Grey is what it passed over without saying so.
Any option from each decision combines with any option from the others. Five control models, eight safeguards, six pricing shapes, eight headlines and eight acquisition messages is 5 × 8 × 6 × 8 × 8, or 15,360 possible products.
Those are the alternatives the agent proposed, not every alternative that exists. A wider set of options would make a bigger space, and the document would still describe one point in it.
Three of those products, to show what changing the choices does:
| One combination | v2 as written | Another | |
|---|---|---|---|
| Control | Read-only | Draft and approve | Earned autonomy |
| Safeguard | A named human | A plain-English log | Approve every action |
| Price | £29 plus an add-on | £99 a month | 5% commission |
| Headline | Nothing is sent until you send | Stop moving information | The office manager |
| Acquisition | Trade association | Referral, first month free | A plumber in Leeds |
All three are products you could ship. The document does not say which one an owner would buy.
Testing every combination is not possible and not necessary. Decisions are treated as separable unless there is a reason to think two of them change each other's answer, so each decision's options are scored against each other, and the best option from each composes into the best product. Price and trial shape run together as one set, because nobody judges a price without knowing what they get to try first.
Seven simulations, 104 questions, about fifteen minutes on a model of UK SMEs. Five test a decision, one checks whether approving everything recreates the work, and the seventh tests v2 exactly as written, end to end, so there is a baseline to compare the assembled winner against.
They run in parallel.
Then every option has a score.
| Decision | What v2 chose | Where it ranked | What won |
|---|---|---|---|
| Control model | Draft and approve · 70% | 3rd of 5 | Earned autonomy · 76% |
| Trust safeguard | Plain-English log · 57% | 5th of 8 | Approve every action · 78% |
| Price and trial | £99 a month · 30% | 4th of 6 | 5% commission · 72% |
| Headline | Stop moving information · 66% | 2nd of 8 | The office manager · 67% |
| Acquisition | Referral, first month free · 62% | 5th of 8 | A plumber in Leeds · 82% |
Two of the five are close enough to argue about. The headline lost by a point, and the message that beat it splits the audience 44% to 44% on whether it sounds patronising, while the one v2 chose has the best clarity score of anything tested at 63%. Control is a sequence rather than a choice: earned autonomy scores highest at 76%, but 89% would never accept unsupervised action and 67% are very likely to try the product if draft-and-approve is the default. It is the right default and the wrong permanent state: 64% would give up and do the work themselves if made to approve everything forever.
The other three are not close.
Pricing is the clearest. 30% say they are very likely to pay the £99 a month v2 commits to, 38% say so for £49, and 72% for a 5% commission on what the tool recovers. 81% say commission feels fairer than a monthly fee, and 63% say two tasks do not justify a subscription at all. Merchants object to paying before seeing anything work, rather than to the amount.
The safeguard v2 picked came fifth of eight. A plain-English log scores 57% very effective against approve-every-action at 78% and money limits at 72%. But no single safeguard is enough on its own: with all eight, 65% say they are very likely to connect their systems, and with none, 96% say they are very unlikely. The number shipped matters more than which one, and v2 shipped one.
The acquisition message came fifth too. The referral offer v2 builds its distribution on scores 62%, while a named case study, a plumber in Leeds getting £4,200 of late invoices paid, scores 82%. Watch-only mode scores 77%.
The seventh simulation tested v2 exactly as written, end to end, and it passed. 61% said they were very likely to pay £99 a month for it and 88% would recommend it. On its own, that result sends the plan to the build queue.
The price test says 30% for the same £99, from the same audience, on the same day. Both numbers are real. The same price converts twice as well when it is described as a whole product as it does when it is described as a price, and only the option test puts it beside the commission model that beats it. A round that tested the composed concept alone would have confirmed the plan and missed that.
UK SMEs, 400 simulated respondents. A model of your own users would answer differently.What you changed because of it
£99 a month, no contract.
5% of what it recovers. 72% are very likely to sign up for that against 30% for £99, and 81% say it feels fairer than a monthly fee.
A plain-English log and per-task toggles.
All eight safeguards from day one. The log came fifth of eight on its own, and it takes the full set before 65% say they are very likely to connect.
Draft and approve, with the owner signing off every action.
Draft and approve on day one, earned autonomy after sixty days. 64% would give up and do the work themselves if made to approve everything forever.
Referral is built in from day one: both owners get a free month.
Lead with a named case study. A plumber in Leeds getting £4,200 paid scores 82% against the referral offer's 62%.
Four of the five decisions changed. The headline is the one that stayed, because the message that beat it won by a point and splits the audience on whether it sounds patronising. Nobody had questioned any of the four.
The four changes have not been tested together as one product, so the best version this round found is still a prediction.