Partners Running an agency? Partner with us
The Graylemon Journal Build logs, teardowns and method, by email. Subscribe All resources
Follow the studio LinkedIn Instagram
Partner with usBook a call (opens in a new tab)
Guide · Scoping

How to scope an AI product before you build it

Scoping an AI product means making seven decisions before anyone designs a system: who it's for and which moment in their day it changes, what success looks like and what would make you stop, whether the data exists, which parts need a model at all, what the first release leaves out, which constraints can't move, and how you'll know it works before everyone sees it. Get those in writing and the build has something to be measured against. Skip them and the build becomes the way you find them out, which is the most expensive way there is.

The Scope Cut, drawn on a blueprint grid: a pile of ideas sorted into four trays, for build first, later, cut and don't build.
Mihir PatelFounder, Graylemon
On this page

In short

  • A scope is a set of decisions in writing, not a feature list. The feature list is what's left after the decisions.
  • Decide the success metric and the stop line before any architecture, so the architecture serves the metric.
  • Check the data before you design around it. Most stalled AI builds were designed around data nobody had checked.
  • Remove the AI from the idea and see what's left. If a simpler product does most of the job, build that first.
  • The first release is defined as much by what it leaves out as by what it includes.

Building an AI product got cheap. A convincing prototype that once took a team now takes one person and an AI coding tool. What didn't get cheap is building the wrong thing well. Scoping is where that gets caught: a short stretch of deciding, done before the building starts, that turns "we should build an AI thing for this" into a first release with a buyer, a measure and a boundary.

01What does scoping an AI product actually decide?

Scoping isn't estimating, and it isn't writing requirements. It's the part of the work where the questions that decide whether the product can succeed get answered before money goes into building it. For an AI product there are seven of them:

  1. The buyer and the moment. Who pays, who uses it, and the exact moment in their day that the product changes.
  2. Success and the stop line. The one measure that says it's working, and the result that would make you stop.
  3. The data. Whether the data the product depends on exists, can be reached and is good enough.
  4. Where a model earns its place. Which parts need AI, which need rules, and where a person stays in the loop.
  5. The first release. What's built first, what's later, what's cut and what's never built.
  6. The constraints. Compliance, security, integrations, speed and cost ceilings, stated up front.
  7. How you'll know it works. The evaluation, the tests and the release steps, designed before the build.

The output is a scope document short enough to be read in one sitting and specific enough to be argued with. If it can't be argued with, it isn't specific enough yet.

02Why do AI products need scoping more than other software?

Because the parts that decide success are the parts a demo hides. In ordinary software, a feature either works or it doesn't, and you can usually see which. In an AI product, four things behave differently.

How AI products differ from ordinary software when you scope them
What changesIn ordinary softwareIn an AI product
CorrectnessA feature works or it doesn'tA model is right some of the time, so "works" needs a measure
Dependence on dataThe data is what users type inQuality rests on data you may not hold yet
CostMostly fixed once builtEvery request has a cost, so usage changes the economics
FailureErrors are visibleWrong answers can look confident, and inputs can steer the model

Each of those has to be decided before the build, not discovered after it. A prototype on clean sample data answers none of them, which is why so many AI products look finished at the demo and stall at launch. We go further into what breaks at that point in AI prototype to production.

A demo proves the idea can work. A scope says what "works" means before you build it.

03Who is it for, and what moment in their day changes?

Start with the buyer, named precisely enough that you'd know where to find them. Then separate the buyer from the user if they're different people, because they usually want different things: the clinic owner wants fewer empty slots; the receptionist wants fewer phone calls.

Then name the intervention point: the exact moment in someone's day that the product changes. Not "helps with scheduling", but "when a patient messages to book, the booking happens in that chat instead of a phone call back". The intervention point is the product. If you can't name the moment, you're building a feature list.

Evidence matters here more than enthusiasm. The strongest evidence that a problem is real exists before you build anything: a workaround people already run, and what it costs them. We set out how to weigh it, rung by rung, in Is it worth building?

Decision · who otlo is for

  • What we chose: the builder, the person who runs a community, as otlo's customer, not the member.
  • What we gave up: the larger audience. Members outnumber the people who run communities.
  • Why: the builder carries the cost of the workaround, so the builder is the one who'd pay to stop carrying it.
  • What it decided: every screen that followed, and the positioning line: community intelligence for the people who build communities.

04What does success look like, and what would make you stop?

Agree one success metric before any architecture, so the architecture is chosen to serve it rather than the other way round. It should describe a change in the user's behaviour or the business, not a property of the model: "patients book in chat without being asked to" rather than "the model understands booking requests".

Then write down the stop line: the result that would make you drop or change course, set before you look. We call these kill criteria. They matter more in AI products than most, because it's always possible to believe one more prompt, one more model or one more data source will fix it. A stop line written in advance can't be argued away afterwards.

Two cautions from what we've seen:

  • Don't make accuracy the goal. A model can score well on a test and change nothing about how people work. Accuracy is a means; the success metric is the end.
  • Don't pick a metric you can't measure at launch. Build the tracking into the scope, so the first week of real use produces the number.

Insight

The success metric is the most important architecture decision, and it isn't technical. It decides what the system has to be good at, and what it's allowed to be bad at.

05Does the data exist to support it?

This is the check that saves the most builds. An AI product designed around data that turns out to be partial, messy or out of reach fails late, after every decision has been built on top of it. So before anyone designs a system, check four things about the data the product depends on:

  1. Does it exist? Is the information the model needs recorded anywhere today, or does it live in people's heads?
  2. Can you reach it? Who owns it, what system holds it, and what permission do you need to use it?
  3. Is it good enough? Look at real records, not a clean sample. Missing fields, inconsistent formats and duplicates are the norm.
  4. Is there enough of it for what you want to do? Some features need history before they can work at all. Those belong in "later", not in the first release.

When the data isn't there, the scope changes, not the ambition. Often the first release becomes the thing that collects the data the later feature needs.

06Which parts need a model at all?

The biggest architecture decision in an AI product is which parts need a model. Our test is simple, and we run it on every scope: remove the AI entirely and see what's left. If a simpler product still delivers most of the value, an AI-native architecture isn't justified, and the scope should say so. We've written it up as the Remove-the-AI Test.

Where a model does earn its place, scope its edges before you build it: what it can do, which tools it can use, when it hands over to a person, and what happens when it's unsure. An agent without boundaries isn't a feature; it's a risk with a nice interface.

Decision · apprn's spine is a rule

  • What we chose: for apprn, a rule for the event everything depends on. An artist can only start a service by entering the one-time code sent to the customer with that booking.
  • What we gave up: the easy story of AI in every part of the product.
  • Why: the event that decides commission and pay has to be checkable. A model's best guess isn't evidence in a dispute between an owner and an artist.
  • Where the AI went: into agentic features inside the product, where judgement genuinely helps.

07What goes in the first release?

Every idea on the list goes into one of four trays. That's the Scope Cut:

  • Build first: what the first paying user needs for the product to do its one job.
  • Later: good ideas that depend on something the first release will produce, such as data, users or evidence.
  • Cut: things that would be nice, cost real effort, and don't move the success metric.
  • Don't build: things that shouldn't exist in this product at all, however often they're requested.

The first release is defined as much by what it leaves out as by what it includes. A useful rule: if an idea can't be traced to the success metric or to the stop line, it doesn't go in the first tray. And write down why each idea landed where it did. The reasons are what stop cut features from creeping back in during the build.

Every feature cut in scoping is one nobody has to design, build, test or maintain.

08Which constraints can't move later?

Some things can't be changed cheaply once the build starts, so they belong in the scope. Keep them in a short register, each with an owner:

  • Compliance: where the data may live, who may see it, and what must be logged. Regulated sectors need legal review before the build, not after.
  • Security: sign-in, permissions for each role, and what happens when someone tries to steer the model with the text they send it.
  • Integrations: the systems it has to talk to, and whether those systems will let it.
  • Speed: how fast a response has to be for the moment you're changing. A booking in chat and a report overnight have very different limits.
  • Cost: the most a single request can cost for the product's economics to work, so model choices are made against a ceiling.

Stated up front, each of these shapes the design. Discovered at launch, each one is a rebuild.

09How will you know it works before everyone sees it?

Design the proof into the scope. Three pieces:

  1. An evaluation set: real inputs, each paired with a good output, scored against the success metric. It runs before every deploy, so a change that fixes one case and breaks others never reaches users.
  2. Tests for the parts AI products usually skip: inputs designed to mislead the model, attempts to pull out data it shouldn't share, and benchmarks for speed and cost per request.
  3. A release in steps: internal, then beta, then a soft launch, then live, with each step opening only once the one before it holds. That's Graduated Launch.

And decide where the code and data will live before the build starts. In our work, both sit in the client's own accounts from the first day.

Decide, too, what happens when the evaluation fails. A result below the bar before a release means the release waits, and someone owns the fix. A result below the bar after launch means the step rolls back to the one before. Writing that down in advance matters more for AI products than for most software, because a model can get worse after a change that looked harmless, and the team that notices first should already know what to do.

10Worked example: a clinic booking assistant

Illustrative example This is the example on our home page, and it's illustrative, not client work. A clinic wants an AI assistant for bookings. Here's how the seven decisions shape it.

  • The buyer and the moment: the clinic owner pays; patients and front-desk staff use it. The moment is a patient messaging to book: today that means a call back, and missed appointments leave empty slots. Staff already phone every patient the day before, which is the workaround, and the evidence the problem is real.
  • Success and the stop line: success is patients booking in chat without being told to, and paid bookings turning up for their slots. It goes to market only if the beta clinics keep it switched on after the pilot.
  • The data: the clinic's calendar and its services list. There's no history of no-shows in a usable form yet, which matters for the next decision.
  • Rules or a model: payment at the time of booking cuts no-shows with no AI at all, so it's a rule. The conversation is where a model earns its place, and an off-the-shelf model does that job.

Then the six ideas on the table go into the four trays:

The clinic booking assistant's six ideas, sorted into the Scope Cut's four trays
IdeaTrayWhy
Booking over WhatsApp, with remindersBuild firstPatients are already there, so there's nothing to install
Payment at the time of bookingBuild firstCuts no-shows, with no AI at all
A custom-trained modelDon't buildAn off-the-shelf model does this job
An admin dashboard with 14 reportsCut, to 2Nobody reads the other twelve
No-show predictionLaterIt needs months of real booking data first
A voice assistantLaterNot until the text flow is proven

Two ideas make the first release, and they're built as one flow in the chat patients already use. The constraints register notes where patient data may be stored and who can see a booking. The proof is designed in: conversation tests on real booking requests, checks against inputs that try to steer the assistant, every step it takes traced, and a release that goes to internal use and beta clinics before anyone else. Try sorting the same six ideas yourself in Cut It.

11What are the most common scoping mistakes?

  • Scoping the technology instead of the problem. "We need a RAG pipeline" is an answer to a question nobody has asked yet. Start from the buyer and the moment; the architecture follows.
  • Letting the prototype set the scope. A working prototype makes everything in it feel decided. It isn't: each part still has to earn its place in the first release.
  • Leaving the stop line vague. "We'll see how it goes" isn't a stop line. Without a result written down in advance, every result looks like a reason to carry on.
  • Assuming the data. Designing around data that turns out to be incomplete, locked in another system or not yours to use is the most expensive mistake on this list, because it's found last.
  • Scoping by committee. A room where everyone can add and nobody can cut produces a first release that contains everything and ships late.

Each one is avoided the same way: decide in writing, in order, before the build starts, with the person who can say no in the room.

12What goes in the scope document?

One document, short enough to read in one sitting, with every decision and the reason for it. Use this as the checklist before any build starts:

  1. The buyer, the user, and the intervention point, in one sentence each.
  2. The evidence the problem is real: the workaround, and what it costs.
  3. The success metric, and how it will be measured from the first week.
  4. The stop line, written before anything is built.
  5. The data check: exists, reachable, good enough, enough of it.
  6. Which parts are rules, which are a model, and where a person stays in the loop.
  7. Every idea in one of four trays, with the reason it's there.
  8. The constraints register: compliance, security, integrations, speed and cost, each with an owner.
  9. The evaluation set, the tests and the release steps.
  10. Where the code and data will live.

Insight

A scope is finished when someone who wasn't in the room could read it and say what the first release is, why, and what would make you stop.

13Questions founders ask about scoping

How long should scoping an AI product take?

Long enough to make the seven decisions with evidence, and no longer. Run it to a fixed scope and a fixed end date, so it can't drift into open-ended exploring. What matters is that each decision is written down before the build starts.

We already have a prototype. Do we still need to scope?

Yes, and the prototype helps. It's the best specification you'll ever have: it shows what you meant. Scope against it, then decide whether to extend it or build again, based on the data and the architecture, not on what's already been spent.

Can we scope and build at the same time?

You can, but the build then becomes the way you discover the scope, which is the most expensive way to find it. Decide the success metric, the data and the first release before the build; everything else can be refined as you go.

Who should be in the room when you scope?

The person who decides, someone who knows the data, and someone who knows the users' day. Scoping goes wrong in committees, because a committee can always agree to keep every idea.

What if the data isn't ready?

Then that's the first decision the scope records, and it often changes the first release. Sometimes the first release becomes the thing that collects the data, with the model coming later. That's still progress: a product that gathers the evidence beats one designed around data that isn't there.

Does the first release need AI at all?

Not always. If the Remove-the-AI Test shows rules and a good interface deliver most of the value, build that first and add the model where it earns its place. Plenty of good AI products started as good products without it.

Work through these decisions with the AI product scope template, a free Word file with every field explained.

Where this fits

AI Product Build starts with exactly this: the business objective and success metric, the workflow and intervention point, a data readiness check and a constraints register, before any architecture.

  • Stage: Define
  • Strategy
  • Launching a product
  • Source: method, own ventures, illustrative example

Search Graylemon

↑↓ Move↵ OpenEsc Close