Editorial illustration for Model selection for AI consultancies: measure your client's tasks, not the leaderboard
AI analysis / Latest briefings
From TerraNet · A service we are piloting

Model selection for AI consultancies: measure your client's tasks, not the leaderboard

Public leaderboards cannot choose a model for your client, and a one-shot demo repeats their mistake. We are piloting a service that runs your client's real tasks through candidate models with repeated trials and error bars, then breaks ties on cost.

By TerraNet Technologies4 min read
Editorial illustration for Model selection for AI consultancies: measure your client's tasks, not the leaderboard
llm evaluation
model selection
ai consultancy
confidence intervals
benchmark noise
Listen to this article

~4 min spoken. Keeps playing while you work in another tab.

This post describes a service TerraNet is piloting. The software, working name Separable, is in development. Pilots are run by hand today, and the request form is at the end.

The decision your clients are paying you to make

Every AI engagement includes the same question: which model should this workflow run on? The client usually cannot answer it, which is part of why they hired you. And the answer goes stale each time a vendor ships.

In practice, that choice is often settled by things that have nothing to do with the client's work: which productivity suite they already pay for, which vendor the consultancy partners with, and a tier list absorbed from blog posts. When there is a pilot, it tends to be a handful of prompts run once and judged by eye. The recommendation that comes out of it sounds confident. It is rarely measured.

Why the leaderboard cannot make it for you

In September we published LLM evaluation and the noise floor, which looked at how much the public rankings can actually tell a buyer. The short answer is very little:

  • Of the 77 benchmarks in our catalogue, only 2 publish run-to-run confidence intervals. For the rest, you cannot even ask whether a rank difference is real.
  • On DeepSWE, one of the two that does, only 3 of 27 adjacent rank steps are statistically separable at 95% confidence. Most of the ranking is sorted noise.
  • The same article found that Claude Sonnet 5 costs 263 times as much per task as DeepSeek-V4-Flash, yet their intervals overlap completely. A ranking that separates them is choosing between prices, not capabilities.

A leaderboard also measures someone else's tasks, in someone else's harness. Your client's invoice matching, support triage or document extraction was not on it.

A demo repeats the same mistake

The obvious answer is to test on the client's own work. It is the right answer, with one catch that the original article spelled out: a small set of tasks run once separates nothing. If a public benchmark running hundreds of trials can separate only a handful of its own rank steps, a dozen prompts run once in a sales meeting cannot separate two frontier models.

As that article put it, the case for evaluating your own tasks "is not that it is inherently more precise; it is that it measures the right thing." Precision has to come from somewhere else: repeated trials, explicit error bars, and treating differences within the margin of error as ties.

What we will measure for you

That is the service we are piloting. You bring one client workflow; we do the measurement:

  1. Scope. You name one workflow for one client, for example sorting inbound support email into six queues.
  2. Examples. You send 20 to 50 real examples with known outcomes, drawn from the client's own history.
  3. Harness. You tell us the scaffold you intend to deploy, so the result belongs to the system your client will actually get.
  4. Candidates. You pick 3 to 5 model configurations, including at least one cheap one. We estimate the API spend before anything runs.
  5. Trials. Every task runs repeatedly per model until the intervals settle or the spend cap is reached.
  6. Verdict. Every model runs the same tasks, so models are compared task by task. Where the difference between two models is within the margin of error, they are declared a tie, never ranked. Ties are broken on cost per task and latency, measured on the same runs.

The first version covers tasks with a checkable outcome: classification, extraction, structured output and routing. Open-ended drafting has no ground truth to measure against, and grading it with another model reintroduces the noise this exists to remove.

What you hand your client

The deliverable is a short decision memo, not a dashboard: what was tested, what won, what tied, and why the recommendation follows. A verdict looks like this:

Illustrative verdict: three models tie on the client's tasksFour example models shown as 95% interval bands on task success. Models A, B and C overlap and are marked as a tie; model D is separably lower. The recommendation goes to B on cost and speed. Illustrative data, not a real run.Three models tie on the client's tasks, so cost decidesTask success with 95% intervals. Amber bracket: cannot be told apart. Illustrative data, not a real run.40%50%60%70%80%Model A$4.10 per task · 9.8 sModel B$0.11 per task · 4.1 sModel C$1.35 per task · 6.3 sModel D$0.06 per task · 3.2 sSeparable: below the tieTIEVerdict: A, B and C tie on success. B costs 1/37 of A and returns 2.4 times faster. Recommend B.Illustrative example: 40 trials per model across 22 tasks. Not a measurement of any real model.
An illustrative verdict, not a real run. When candidates cannot be told apart on the client's own tasks, the recommendation is made on cost and latency, and the memo says so.

And the memo that goes with it:

RECOMMENDATION   Model B for the intake-classification workflow.

WHY              A, B and C are statistically tied on your 22 tasks:
                 compared task by task after 40 trials each, the
                 differences are within the margin of error.
                 B costs $0.11 per task against $4.10 for A and
                 returns in 4.1 s against 9.8 s.

WHAT WOULD       More trials of these tasks would not separate A
CHANGE IT        from B: your tasks differ more than the models do.
                 About 60 more tasks would, at roughly $1,300 in
                 API spend. We do not recommend buying that answer;
                 the cost gap dominates any plausible accuracy gap.
An illustrative memo excerpt with example numbers. Every pilot memo says what it would take to settle a tie - more trials or more tasks, and what each would cost - so you and your client can decide whether the answer is worth buying.

The memo is written to go into your statement of work. "We measured it on your data" is a line a vendor-partner shop recommending whatever it resells cannot use.

What a pilot is, and what it is not

  • It is done by hand. There is no software to log into yet. We run the trials with our own scripts and deliver the memo.
  • It is paid, and scoped per engagement. The form asks for a budget range so we can size it; nothing is charged until a pilot is agreed.
  • Data handling is agreed before anything is sent. Your client's examples leave their hands only on terms you both accept.
  • It does not grade open-ended writing. If your engagement turns on drafting quality, this is not yet the tool for it, and we will say so.

Who this is for

AI consultancies and integrators, roughly five to fifty people, who choose models for small and mid-sized clients and would rather show a measurement than a recommendation. If you have picked a model for a client in the last few months, we would like to hear how you did it, even if a pilot is not right for you yet.

Pilot program · Separable is in development

Request a pilot for one of your clients

The software is not built yet. The measurement is: send us one client workflow and twenty to fifty real examples with known outcomes, and we run your candidate models through repeated trials by hand, then deliver the decision memo. Pilots are paid and scoped per engagement; the budget question below helps us size them.

AI Tools

    Model selection for AI consultancies: measure your client's tasks, not the leaderboard | TerraNet Technologies