On this page
In short
- Build answers one question: does the thing exist, and does it work? "Works" means for real users on real data, not in the demo.
- We agree the success metric and check the data before any architecture. Many stalled builds were designed around data nobody had checked.
- Rules before models: if a simpler product delivers most of the value without AI, we say so and build that.
- Evaluation, security testing and cost per request are designed before the build, and a graduated launch finds failures before everyone else does.
Building got cheap. A convincing prototype that once needed a team now needs one person and an AI coding tool. What didn't get cheap is a product that holds up when real users arrive with their own data, their own permissions and their own mistakes. Build is the stage where a prototype becomes that product, or proves it can't.
01What is the Build stage for?
Build is the third stage of Stuck-to-Scale, and its question is plain: does the thing exist, and does it work?
Both halves carry weight. "Exists" means in production, not in a demo: sign-in, permissions, admin tools, real data, and the monitoring to see what it's doing. "Works" means for real users, measured against the success metric agreed at Define, at a cost per request that makes sense for the business.
What Build isn't: a demo, a proof of concept dressed up as a product, or a product shipped without evaluation or observability. Each of those can look finished, and each one moves the hard part to the day after launch.
02What does being stuck at Build look like?
These are patterns we've seen across the AI products we've looked at, and what each one usually means.
| Signal | What it looks like | What it usually means |
|---|---|---|
| "Almost done", again | Every stretch of work ends one feature away | The scope was never fixed, so there's no "done" to reach |
| The demo works; real data breaks it | Messy records, missing fields, permissions nobody tested | The prototype was built on clean sample data |
| Model quality is treated as the product | Every fix is a better prompt | The product's value rests on one component nobody can test |
| Users won't change how they work | The feature exists; nobody uses it | The AI asks for a change in routine that users won't make |
| Launch is a single switch | Everyone at once, on the first day | Failures will be found by the users you most wanted to impress |
03What do we do before any code?
The first stage of an AI Product Build is discovery, and it produces four things before anyone designs a system.
- The business objective and the success metric, carried over from Define and agreed before any architecture, so the architecture serves the metric rather than the other way round.
- A workflow and user map, with the intervention point named: the exact moment in someone's day that the product changes.
- A data readiness assessment. Does the data you have actually support the outcome you want? This is the check that saves the most builds. A product designed around data that turns out to be partial, messy or out of reach fails late, after every decision has been built on top of it.
- A constraints register: compliance, security, integration, speed and cost ceilings, stated up front instead of discovered at launch.
Insight
The intervention point is the product. If you can't name the moment in the user's day that changes, you're building a feature list.
04Rules or a model: how do we decide?
The biggest architecture decision in an AI product is which parts need a model at all. Our test is simple, and we run it on every product: remove the AI entirely. If a simpler product still delivers most of the value, an AI-native architecture isn't justified, and we'll say so. It's written up as the Remove-the-AI Test.
Where a model does earn its place, we design its edges before we build it: which tools it can use, when it hands over to a person, and where a human stays in the loop. An agent without boundaries isn't a feature. It's a risk with a nice interface.
apprn is our clearest case. The event everything depends on is deterministic: an artist can only start a service by entering the one-time code sent to the customer with that booking. That's a rule, not a model.
Decision · apprn's spine is a rule
- What we chose: a rule for the core event, and AI only where judgement helps, in agentic features inside the product.
- What we gave up: the easy story of AI in every part of the product.
- Why: the event that decides commission and pay has to be checkable. A model's best guess isn't evidence in a dispute between an owner and an artist.
- What it made possible: every agentic feature sits on a record nobody argues with.
We go further into what breaks when a prototype meets real users, and when to start again rather than patch, in Your AI-built prototype works. Here's what breaks when real users arrive.
05How do we know it works?
Two stages of the build exist to answer that, and neither is optional.
Prototype and validation. A working prototype running on real data, and an evaluation set sized to the problem, measured against the success metric from discovery. It ends in a go or no-go. A failed prototype at this point is a cheap discovery, not a failed engagement: it's the cheapest moment the answer can arrive.
Testing and assurance. Functional and integration tests, as you'd expect. Then the parts AI products usually skip:
- model evaluation, with a regression suite that runs before every deploy;
- prompt and adversarial testing, including prompt injection and attempts to pull data out;
- benchmarks for speed and for cost per request, so the product's economics are known before its users are.
A model is code whose behaviour changes when its inputs do. It gets a regression suite like everything else.
A demo proves it can work. An evaluation set proves it does.
06How do we release without breaking trust?
In steps. Graduated Launch runs internal, then beta, then a soft launch, then live, and each step opens only once the one before it holds. More users arrive at each step, so each kind of failure is found by the smallest group that can find it.
The launch comes with the means to see what's happening: automated deployment, tracing along the agent's decision path, and dashboards for cost and quality. Then documentation and handover: the architecture, runbooks, and the model and prompt documentation with the reasoning behind each choice, plus product documentation written to be read by machines as well as people. The code lives in your repository from day one. The aim of Build is a product you can run without us.
Once real users have it, the question changes: do the right people want it badly enough to keep using it? That's PMF.
07Questions founders ask about Build
Can you take our AI-built prototype to production?
Often, yes, and it's one of the things an AI Product Build is for. Sometimes the better call is to keep the prototype as a specification and build again. We decide that from the data and the architecture, not from what's already been spent.
Does our product have to use AI?
No. If rules deliver most of the value, we'll build rules and tell you why. AI goes where judgement is genuinely needed, and nowhere else.
Who owns the code?
You do. It lives in your repository from the first day, with the documentation to run it.
What do you need from us during Build?
Access to the data, a technical counterpart, quick decisions, and your constraints stated up front. Most delays in a build trace back to one of those four.
Where this fits
