Model selection for AI consultancies: measure your client's tasks, not the leaderboard
Public leaderboards cannot choose a model for your client, and a one-shot demo repeats their mistake. We are piloting a service that runs your client's real tasks through candidate models with repeated trials and error bars, then breaks ties on cost.
~4 min spoken. Keeps playing while you work in another tab.
This post describes a service TerraNet is piloting. The software, working name Separable, is in development. Pilots are run by hand today, and the request form is at the end.
The decision your clients are paying you to make
Every AI engagement includes the same question: which model should this workflow run on? The client usually cannot answer it, which is part of why they hired you. And the answer goes stale each time a vendor ships.
In practice, that choice is often settled by things that have nothing to do with the client's work: which productivity suite they already pay for, which vendor the consultancy partners with, and a tier list absorbed from blog posts. When there is a pilot, it tends to be a handful of prompts run once and judged by eye. The recommendation that comes out of it sounds confident. It is rarely measured.
Why the leaderboard cannot make it for you
In September we published LLM evaluation and the noise floor, which looked at how much the public rankings can actually tell a buyer. The short answer is very little:
- Of the 77 benchmarks in our catalogue, only 2 publish run-to-run confidence intervals. For the rest, you cannot even ask whether a rank difference is real.
- On DeepSWE, one of the two that does, only 3 of 27 adjacent rank steps are statistically separable at 95% confidence. Most of the ranking is sorted noise.
- The same article found that Claude Sonnet 5 costs 263 times as much per task as DeepSeek-V4-Flash, yet their intervals overlap completely. A ranking that separates them is choosing between prices, not capabilities.
A leaderboard also measures someone else's tasks, in someone else's harness. Your client's invoice matching, support triage or document extraction was not on it.
A demo repeats the same mistake
The obvious answer is to test on the client's own work. It is the right answer, with one catch that the original article spelled out: a small set of tasks run once separates nothing. If a public benchmark running hundreds of trials can separate only a handful of its own rank steps, a dozen prompts run once in a sales meeting cannot separate two frontier models.
As that article put it, the case for evaluating your own tasks "is not that it is inherently more precise; it is that it measures the right thing." Precision has to come from somewhere else: repeated trials, explicit error bars, and treating differences within the margin of error as ties.
What we will measure for you
That is the service we are piloting. You bring one client workflow; we do the measurement:
- Scope. You name one workflow for one client, for example sorting inbound support email into six queues.
- Examples. You send 20 to 50 real examples with known outcomes, drawn from the client's own history.
- Harness. You tell us the scaffold you intend to deploy, so the result belongs to the system your client will actually get.
- Candidates. You pick 3 to 5 model configurations, including at least one cheap one. We estimate the API spend before anything runs.
- Trials. Every task runs repeatedly per model until the intervals settle or the spend cap is reached.
- Verdict. Every model runs the same tasks, so models are compared task by task. Where the difference between two models is within the margin of error, they are declared a tie, never ranked. Ties are broken on cost per task and latency, measured on the same runs.
The first version covers tasks with a checkable outcome: classification, extraction, structured output and routing. Open-ended drafting has no ground truth to measure against, and grading it with another model reintroduces the noise this exists to remove.
What you hand your client
The deliverable is a short decision memo, not a dashboard: what was tested, what won, what tied, and why the recommendation follows. A verdict looks like this:
And the memo that goes with it:
RECOMMENDATION Model B for the intake-classification workflow.
WHY A, B and C are statistically tied on your 22 tasks:
compared task by task after 40 trials each, the
differences are within the margin of error.
B costs $0.11 per task against $4.10 for A and
returns in 4.1 s against 9.8 s.
WHAT WOULD More trials of these tasks would not separate A
CHANGE IT from B: your tasks differ more than the models do.
About 60 more tasks would, at roughly $1,300 in
API spend. We do not recommend buying that answer;
the cost gap dominates any plausible accuracy gap.
The memo is written to go into your statement of work. "We measured it on your data" is a line a vendor-partner shop recommending whatever it resells cannot use.
What a pilot is, and what it is not
- It is done by hand. There is no software to log into yet. We run the trials with our own scripts and deliver the memo.
- It is paid, and scoped per engagement. The form asks for a budget range so we can size it; nothing is charged until a pilot is agreed.
- Data handling is agreed before anything is sent. Your client's examples leave their hands only on terms you both accept.
- It does not grade open-ended writing. If your engagement turns on drafting quality, this is not yet the tool for it, and we will say so.
Who this is for
AI consultancies and integrators, roughly five to fifty people, who choose models for small and mid-sized clients and would rather show a measurement than a recommendation. If you have picked a model for a client in the last few months, we would like to hear how you did it, even if a pilot is not right for you yet.
Pilot program · Separable is in development
Request a pilot for one of your clients
The software is not built yet. The measurement is: send us one client workflow and twenty to fifty real examples with known outcomes, and we run your candidate models through repeated trials by hand, then deliver the decision memo. Pilots are paid and scoped per engagement; the budget question below helps us size them.