AI sprints: where the measuring happens. The talking before it is free.
Four steps get you from a first call to a scope and a fixed number, and you can stop at any of them. What follows is a sprint: we score the tooling you run today, score the AI approach against the same test set, and hand you the gap with the arithmetic shown. One sprint answers one decision.
How a consultation works.
Nothing starts until the scope and the cost are agreed in writing. Until then this costs you a call and an hour of your team's time.
01
The intro call
You describe the workflow that is costing you the most, and what you have already tried on it. We ask what your team runs today and how you would know if it got better. No deck, and nobody joins to qualify you.
Costs nothing, and plenty of these end with us saying you do not need a sprint yet.
02
The scoping note
We write back with the decision we think is actually blocking you, which service settles it, what data we would need read access to, and who we would need an hour with. It is short enough to forward to whoever signs.
You get a fixed number in that note. Not an hourly rate, and not a range.
03
The go or no-go
You approve the note as written, or you change the scope and we requote once. If your security review needs to happen first, it happens here — before any access is granted, not after work has started.
Nothing begins until the scope and the number are agreed in writing.
04
The sprint starts
Read access is granted on the narrowest scope that lets us measure, we take the hour with the people who run the workflow, and the clock starts. From here it is the three stages below.
Most first sprints run two to four weeks, mostly set by how fast access lands.
What happens inside the sprint.
Three stages, in this order, every time. There is no version that skips the baseline and goes straight to a recommendation — without the first stage the second one has nothing to be a gap from.
01
Measure what you run today
A score for the system you already have.
We pull real work out of your ticket queue and your internal channels, grade a sample of it by hand, and score your current tooling against that set. This is the baseline, and it is the number most AI proposals never bother to establish.
- A test set built from your own work, not a benchmark
- A graded baseline for the tooling you run today
- The workflows ranked by what they cost you now
02
Score the AI approach against it
The gap between the two, priced.
The same test set, the same grading, run against the AI approach — and then both sides priced, including what it costs to keep running once the demo is over. The gap between those two scores is what your budget would actually be buying.
- The AI approach scored on the identical set
- Run cost at your real volume, not at demo volume
- Where it fails, and how loudly it fails when it does
03
Hand over the findings
The verdict, with the arithmetic shown.
Every opportunity ranked by payback, the arithmetic visible on each row, and a build plan for the one at the top — written for your engineers, with the accuracy bar and the failure signals spelled out. Then a walkthrough with whoever has to act on it.
- The ranked table, with the working shown per row
- A build plan for the top row, written for engineers
- The test set and runbook, so you can re-run it without us
What holds in every sprint.
One sprint, one decision.
A sprint is scoped to the question that is blocking you, not to a calendar. If a second question turns up mid-sprint, it is written up alongside the rest — it does not quietly become extra work.
You keep everything it produces.
The test set, the grading notes and the runbook come with the findings. Your team can re-run the measurement after we are gone, which is the point of building it from your data in the first place.
No production access, and nothing has to ship first.
Read access on the narrowest scope that lets us measure, under an NDA, and your data never trains anything. We do not need a code change to start.
“Do not build this” is a valid result.
Plenty of candidate use cases do not clear the bar once you price them against the tools you already run. When that is what the numbers say, that is what the findings say.