Vision documents are written to persuade, so the assumptions underneath them rarely get checked. This lesson takes one, an idea called ChatGPT for SMEs, and tests every claim and assumption it depends on against a model of UK small business owners: twelve claims, five simulations, about fifteen minutes. The document's description of the problem was supported, with 90% of merchants agreeing that their tools not talking to each other is the most frustrating part of admin. Its main selling point was not. "No setup, no prompting, no configuration" describes an assistant that acts on its own, and 92% of merchants are very uncomfortable with that. The simulations also show why they are uncomfortable, and the reason rules out solving it with better onboarding or clearer marketing.
What you'd normally do
Send it to four people who already agree with you. Or build enough of it to find out in three months. Or book six interviews, wait a fortnight for calendars, and get six opinions shaped by whoever replied. None of those separate what the document claims from what it assumes, so you cannot tell which part is wrong.
What you bring
A write-up of an idea called ChatGPT for SMEs:
The biggest untapped market for AI is not the enterprise, it's the 300 million small businesses that run the global economy. A plumber with four staff, a café owner, a two-person accountancy firm — none of them are writing prompts, connecting APIs, or building agent workflows.
Meanwhile the tools they actually use, their invoicing software, their bank, their booking system, their inbox, all sit disconnected. The owner is the integration layer. They spend evenings moving information between systems by hand.
A system that connects to the handful of tools every SME already uses, understands the standard shapes of small business operations — chasing late invoices, quoting jobs, scheduling staff, filing taxes, replying to customers — and just does them. No setup, no prompting, no configuration. Less like software and more like the capable office manager they could never afford to hire.
For this example we started with a one page vision document. For problem discovery you can start with anything that describes what you think is wrong:
- A single sentence saying what problem you think people have
- A one page vision document like this one
- A multi page spec or PRD
- A dump of support tickets, sales call notes or user interview notes
- A link to a competitor, a review page or a forum thread
What comes back
The agent reads the document and pulls twelve claims out of it. Some are things the document states. Some are things it assumes without stating, which have to be true for the stated ones to matter.
It also says what it cannot help with:
Our SME audience is UK in-person merchants — not the global 300 million, and not the two-person accountancy. Your profile describes web presence and listings management, while this describes a whole-business operational layer. I will test what is written here.
Yours to answer: which market and vertical you launch in, your price floor, whether those integrations are actually obtainable.
It holds the product description and the price fixed, so the simulations test the twelve claims rather than variations on the product itself. It then lists what will not be simulated this round but could be later: other price points, per-job and per-seat pricing, a version that asks for approval on each action, merchants outside the UK and outside retail, and switching from an existing bookkeeper.
The twelve claims are grouped into five simulations.
flowchart LR D["`**Input** · vision document An operational layer for small businesses`"]:::document D --> K1["`**1** · Owners recognise being the integration layer`"]:::question D --> K2["`**2** · Evening admin is real`"]:::question D --> K3["`**3** · Disconnected tools cost meaningful time`"]:::question K1 --> S1["`**Simulation** Admin pain`"]:::simulation K2 --> S1 K3 --> S1 D --> K4["`**4** · The product as described appeals`"]:::question D --> K5["`**5** · £99/month is payable`"]:::question D --> K6["`**6** · No setup beats configurable control`"]:::question K4 --> S2["`**Simulation** Concept and price`"]:::simulation K5 --> S2 K6 --> S2 D --> K7["`**7** · Bank, invoicing and inbox access is acceptable`"]:::question D --> K8["`**8** · An AI acting without asking is acceptable`"]:::question D --> K9["`**9** · An AI error is a reputational risk they refuse`"]:::question K7 --> S3["`**Simulation** Autonomy and access`"]:::simulation K8 --> S3 K9 --> S3 D --> K10["`**10** · Appeal differs by job`"]:::question D --> K11["`**11** · At least one job is never handed over`"]:::question K10 --> S4["`**Simulation** Job order`"]:::simulation K11 --> S4 D --> K12["`**12** · Office manager beats done-for-you and AI-led`"]:::question K12 --> S5["`**Simulation** Framing`"]:::simulation
Twelve claims from the document, grouped into the five simulations built to test them.
The simulations run and take about fifteen minutes. The agent then produces a synthesis with an answer for every claim.
| Claim | Simulation | Result |
|---|---|---|
| 1 · Owners recognise being the integration layer | Admin pain | Supported · 90% |
| 2 · Evening admin is real | Admin pain | Partly · 52% |
| 3 · Disconnected tools cost meaningful time | Admin pain | Supported · 67% |
| 4 · The product as described appeals | Concept and price | Supported · 87% |
| 5 · £99/month is payable | Concept and price | Not supported · 71% |
| 6 · No setup beats configurable control | Concept and price | Not supported · 88% |
| 7 · Bank, invoicing and inbox access is acceptable | Autonomy and access | Not supported · 87% |
| 8 · An AI acting without asking is acceptable | Autonomy and access | Not supported · 92% |
| 9 · An AI error is a reputational risk they refuse | Autonomy and access | Supported · 88% |
| 10 · Appeal differs by job | Job order | Supported |
| 11 · At least one job is never handed over | Job order | Supported · 93% |
| 12 · Office manager beats done-for-you and AI-led | Framing | Partly |
The document's description of the problem is supported. 67% say they spend a lot of time moving information between tools by hand. The evening admin claim is the weakest of the three at 52%, but still partly supported.
The claims about how the product should work were not supported, and claim 9 explains the other two. 88% strongly agree that an AI-sent email or quote containing an error would do reputational damage that is difficult to repair. Someone who believes that is not going to grant access to their bank and inbox, or let the system act without asking, because the cost of a mistake is a customer relationship and the gain is an evening of admin. 87% say the access is too risky and 92% are uncomfortable with unsupervised action, which is what you would expect based on the first result.
This changes what you do next. If merchants were only nervous, clearer onboarding or a page about security might solve it. They are weighing a specific consequence, so the only thing that changes the answer is a product where a mistake cannot reach the customer unreviewed. The same pattern appears in the other simulations. Invoice chasing is 85% appealing against 8% rejection while replying to customers draws 29% rejection, and 88% say that choosing which tasks the AI handles unsupervised would make them much more likely to use it.
Testing only the obvious claims would have missed this. The concept appeals to 87% and 74% call £99 fair value, and those two results on their own would have supported building the product that the other claims rule out.
UK POS merchants, a model of UK SMEs taking payments in person. A model of your own users would answer differently.What you changed because of it
No setup, no prompting, no configuration. It just does them.
The AI does the work, the owner signs it off, and a plain-English log records what happened. 75% say that account of its actions would give them enough control.
Chasing late invoices, quoting jobs, scheduling staff, filing taxes, replying to customers.
Invoice chasing first, listings second. Customer replies and tax filings stay behind a review step, as a product constraint rather than a messaging choice.
The operational layer for the whole business, working across every system the owner touches.
It starts with one task and earns the next. 88% say choosing which tasks it handles unsupervised would make them much more likely to use it.
Enterprises will build their own AI stacks. The other 99% of businesses will buy one.
They will buy it one approved task at a time, and the first ones will be recruited by the 80% who recommend it before they buy it.
The market thesis and the office manager line were both supported. The delivery model was not, and replacing it changes more than the copy. A product that waits for approval needs an audit trail, a review step in front of anything customer facing, and a way to show its reasoning before it acts.