The standard you can't write down

Your firm now drafts proposals with AI. It triages support with AI, summarises client calls with AI, screens applications with AI. And if you ask how anyone knows whether that work is being done your way — to the standard your name is on — the honest answer, almost everywhere, is a vibe check: whoever happens to glance at the output that day decides whether it feels right.

That was tolerable when AI wrote first drafts nobody shipped. It stops being tolerable at exactly the point most organisations have now reached, where the model’s output goes to a client with light edits or none, and the volume is far past what anyone can read.

Satya Nadella named the general problem in an essay that the forward-thinking end of this industry has been chewing on ever since: “Companies need to turn their workflows, domain knowledge, and accumulated judgment into AI systems that improve with each use. Private evals should capture whether a model is actually improving against outcomes that matter to the business (not just external benchmarks!)” His larger claim is that this learning loop on top of models — not the model itself — “becomes the new IP of the firm”: a hill-climbing machine that, unlike most assets, compounds. And underneath it, the sentence I would put on the wall: “You can offload a task, or even a job, but you can never offload your learning.”

Take that seriously and the implication is uncomfortable in a specific way.

Why “private” is the load-bearing word

Every published measure of AI quality — the leaderboards, the benchmark suites, the vendor evals — was written by strangers, about tasks that are not your business, scored against a standard that is not yours. That is not a criticism of benchmarks; it is what benchmarks are for. They tell you whether a model is generally capable. They cannot tell you whether it handles your edge cases the way your best people would, because nobody who wrote them has ever seen your edge cases or met your best people.

Meanwhile, the system doing your work will not hold still. Models get upgraded on the vendor’s schedule, not yours. Prompts get tweaked by whoever touched them last. Vendors get swapped when procurement finds a better price. Each of those changes alters the work being done in your name — silently, unless you have your own measure. There is no changelog entry that says “the summaries now hedge more than your firm would.”

So an organisation that cannot measure AI against its own standard has, in the only sense that matters, handed its standard to whoever owns the model. That is what “private” means in private evals, and it is why this is not an engineering nicety. The eval is where an organisation’s judgement either becomes an owned asset or quietly stops being consulted.

Stated as a recipe, it sounds almost embarrassingly simple. Take the fifty cases your senior people agree were handled well and the twenty they agree were botched. Encode that judgement. Now every model upgrade, prompt change, and vendor swap gets scored against your bar rather than a stranger’s leaderboard — and when the AI stops sounding like your firm, you find out from a number, not from a client.

Simple to state. So why doesn’t every organisation have one?

The three walls

Because building a private eval runs through three walls, and only the third one is a machine-learning problem. We have hit all three ourselves — Harmonica’s product is an AI facilitator, and “was this conversation facilitated well?” is about as judgement-laden as quality questions get — so what follows is field notes, not theory.

One scope note first: if the question has a checkable answer — did the model extract the right invoice total — none of this applies; write the test and move on. The walls only rise where good is a judgement call.

The first wall: your standard is tacit. Your best people know good work when they see it. Ask them to write down the criteria and the list they produce will not match the calls they actually make. This is not a flaw in your people; it is what expertise is. Judgement mostly lives as recognition, not rules — a senior lawyer does not consult a checklist to know a contract clause is dangerous, and a senior consultant does not deduce from principles that a slide will die in front of a board. Everyone who has seriously tried to build evals has hit this wall: the rubric you draft in a conference room is a guess about your own judgement, and it is usually wrong in ways you only discover when you hold it against real cases and watch it disagree with the very people it was meant to encode.

The machine shortcut fails at the same wall. When researchers benchmarked model-generated rubrics against expert-written ones (RubricBench, 2026), the same judge scored roughly 26 points lower with the model’s rubric than with the experts’ — a gap that held for every frontier model tested and did not close with more compute. The generated rubrics missed half or more of the constraints the experts considered essential. And to produce their gold-standard rubrics, the benchmark’s authors had experts annotate independently and a senior reviewer reconcile the disagreements: a deliberation, in miniature, because nothing else produced rubrics worth trusting.

The second wall: your experts disagree. Sit three senior people down with the same twenty cases and ask them to grade the work, and you will discover the firm does not have one standard. It has three, held by people who each assumed theirs was the house view. One partner prizes thoroughness and reads brevity as corner-cutting; another prizes judgement about what to omit and reads thoroughness as not knowing the point. Both have been “the standard” for the teams under them for years, and nobody noticed because the two standards never had to grade the same artifact side by side.

This is the moment most eval efforts quietly die — not because the disagreement is fatal, but because nobody signed up to run the negotiation it exposes. The tempting escape is to average: score everything with all three rubrics and take the mean. But an averaged standard is a standard nobody actually holds, and your people can tell. Encoding judgement forces a deliberation the organisation never scheduled. Hold that thought.

The third wall: AI judges lie confidently. At some point, someone proposes the obvious shortcut: have a model do the judging. It reads every output, applies the rubric, produces scores at a volume no human panel could match. And an uncalibrated model judge will do all of that while measuring almost nothing — scoring fluently, consistently, and wrong.

We know because it happened to us. We built automated quality scoring for our own facilitator, shipped prompt improvements against it, and then discovered the scoreboard itself was broken in three distinct ways. We published the post-mortem, including the part where fixing the measurement took longer than the improvements it was supposed to measure — and mattered more. Very few vendors in this category will show you their failure. We think it is exactly the standing this work requires, because the failure taught us the discipline: a judge must be checked against your experts’ actual verdicts, its systematic errors measured and corrected, before its scores mean anything. We now run a hard internal rule — no judge gates a real decision until it has passed that bar. A number nobody has calibrated is worse than no number, because people trust it.

The pattern is older than AI. Any system that compresses a domain into governing principles and then stops checking those principles against reality will drift: GDP standing in for national welfare, league tables standing in for a school, the correlated-risk models still producing confident numbers in 2007. Each of them ran correctly. None of them had anything whose job was to break them. An uncalibrated model judge is that machine again, running faster and cheaper, pointed at the standard your firm answers for.

The hard part is a deliberation problem

Now look back at the first two walls. Extracting tacit judgement from experts who cannot fully articulate it. Surfacing where they genuinely disagree. Letting that disagreement be contested rather than averaged away. Arriving at a standard the group has actually seen, argued with, and settled on.

That is not a machine-learning problem. That is a deliberation problem — and it is the exact problem we have been building for.

Harmonica runs a loop we built for how groups think together, whatever the subject: people express what they think in one-to-one conversations with an AI facilitator; the system maps the themes, the tensions, and the places agreement is forming across every conversation at once; then it reflects that map back and lets the group contest it before anything is called settled. Run that loop with your senior people on the question “what does good look like here” — with the fifty good cases and the twenty botched ones as the material — and watch what the walls turn into.

The tacitness wall becomes the express move: instead of asking experts to legislate criteria in the abstract, the facilitator walks each one through real cases — why is this one good, what specifically would you change here, what made you wince — and the criteria surface the way tacit knowledge actually surfaces, through reactions to concrete work, not through introspection in a conference room.

The disagreement wall becomes the map and the contest: the partners’ three standards show up as visible fault lines rather than silent averaging, the group sees exactly where they divide and gets to argue about it. Some disagreements dissolve on contact; they were vocabulary, not values. Some are real, and the group decides, explicitly, which pole is the house view, or that both are acceptable in different contexts, and that decision is recorded with its reasoning attached.

And what settles out the other end is not a workshop artifact. It is your rubric — grounded in real cases, contested by the people whose judgement it encodes, with the labeled examples and the recorded reasoning that judge calibration needs as ground truth. It is also yours in the literal sense: rubric, labels, and reasoning leave as your files, in your repository, not as rows in someone else’s database — a standard you could not take with you when you swap vendors was never really private. The third wall still takes the unglamorous statistical work, and that is the part we have run on ourselves for months. But the input it depends on — a standard your experts actually agree is theirs — now exists, which is the thing most eval efforts never get.

The same mechanism that captures what a group thinks captures what an organisation judges, because they were always the same problem. The working example is our own field: expert facilitators, whose sense of good facilitation is about as tacit as judgement gets, turning it with us into a measurable, open standard.

Where the raw material comes from

Calibration runs on captured judgement: real cases, real verdicts, the reasoning behind them. Most organisations have none of that written down anywhere. What we decided, why, where we disagreed, what our best people mean by “good”: it gets said out loud in meetings and then evaporates.

So the capture is part of the work rather than a precondition for it. The loop above is where it happens: experts reacting to concrete cases, disagreement surfaced instead of averaged, decisions recorded with their reasoning attached. Memory is the capture; the eval is the converter that turns captured judgement into an owned, compounding capability.

It also does one job the eval can never do for itself. A rubric is your current judgement, reified — so it cannot notice when that judgement moves. Recalibrate a judge against this quarter’s verdicts and it will happily track a standard that has quietly slid, reporting green the whole way. Only a record of past positions can catch the standard itself drifting, which is why the rubric and the record belong together rather than in sequence.

Some organisations already run the capture continuously, under another name: retrospectives run as a rhythm rather than a project, decisions that settle on a map instead of in a corridor, a record that accumulates instead of evaporating. Where that exists, the eval starts from a corpus instead of a blank page. Where it doesn’t, this is a good reason to start, and it pays for itself without an eval on top.

Where this stands

I will be straight about the status of all this, because straightness is the house style. The loop is live; the calibration methodology is real and battle-tested on our own product; the expert-facilitator standard is a working example. “We will build your private eval with you” as a packaged offering is younger than any of those — which is why the next step is not a product page with a price on it. It is a conversation.

If your organisation is wrestling with any wall of this — a standard you can’t write down, experts who turn out to disagree, an AI judge you’re not sure you can trust — talk to us. Bring your twenty botched cases. We have been on the other side of every one of these walls, and we published what it cost us.

Book a 30-minute call

Loading calendar…