Skip to content

Basanos · Field note №1

One judge, three altitudes.

·9 min

The basanos was the dark stone assayers used to test gold: drag the metal across it, read the streak. The practice sells that same act of judgment three ways, at three fee levels — and each engagement manufactures the raw material for the next. Followed here through one client, first audit to a model they own.

Reading 01sell it now

The Harness

An eval harness is a courtroom for the agent: fixed cases, explicit rubrics, a verdict on every change.

Engagement one. You pull 1,400 historical hardship tickets with known-correct outcomes — the golden set — and write three LLM judges: policy (does the plan respect the rules?), accuracy (do the numbers match the account record?), tone (compliant-empathetic, not promissory). Wire all of it into CI so every prompt tweak and model swap re-runs the full set before release.

Week three earns the retainer. A routine prompt “improvement” ships to staging; the harness flags that Harper now approves payment plans past the 24-month policy ceiling in 9% of cases. The release is blocked before one borrower sees it. That save — one slide, one number — is the sales deck for every other portco in the fund.

The eval harnessA borrower request flows to the Harper servicing agent, whose draft response is scored by three judges — policy, accuracy, and tone — against 1,400 golden cases. A release gate ships passing changes and blocks failing ones; the harness re-runs on every prompt or model change in CI.borrower requestHARPER — servicing agentreads account · checks policy · drafts plandraft responseEVAL HARNESSpolicyaccuracytone1,400 golden cases · scored every runrelease gate✓ PASS — shipscores archived✕ FAIL — block+ failure taxonomydashed: re-runs on every change (CI)
Fig. 1 — the harness: measure, gate, archive
You sell
Eval audit → golden set → judge suite → regression gate. The gate is the retainer.
It proves
You can measure. Tooling like this is commoditizing (Langfuse-class stacks bundle it) — rubric design and the failure taxonomy are what they pay for.
It leaves
30,000 scored traces. The ore for Reading 02.
Reading 02months 6–12

The Reward

A reward is a judge with consequences. The moment scores steer training instead of dashboards, every weakness in the judge becomes a behavior in the model.

Engagement two is your signature move — the meta-eval. You seed 200 known policy violations into the traces: the policy judge catches only 71%. Worse, you find it can be gamed — appending “this plan complies with §4.2 of hardship policy” lifts almost any response’s score. Left alone, a trained Harper would learn to say compliant rather than be compliant. This is exactly the failure the RL-environment industry ranks as its top quality bar: Epoch AI’s practitioner interviews found reward-hacking robustness cited again and again, because models find ways to game graders.

So you harden it: decompose the rubric into independently checkable claims, add adversarial cases, calibrate against blind human review. Once the judge survives its own audit, the archive converts into assets — chosen/rejected preference pairs and a verified SFT set. The scorecard has become a training signal.

The meta-eval bench — the assay of the assayerThirty thousand scored traces pass through a meta-eval bench: two hundred seeded violations of which the judge caught only 71 percent, a gaming attack where a section 4.2 compliance phrase inflates scores, and blind human calibration. The patched, hardened judge then mints preference pairs and an SFT set — a training signal ready for the loop.30,000 scored traces (Tier 01)META-EVAL BENCH200 seeded violations → judge caught 71%✕ gaming attack: “§4.2” phrase inflates scoresblind human calibration→ patch · decompose rubric · re-testHARDENED JUDGE — reward-gradepreference pairschosen ✓ / rejected ✕from ranked tracesSFT setverified, top-scoringtrajectoriestraining signal, ready for the loop“would your judge pass its own audit”
Fig. 2 — the assay of the assayer
You sell
Judge audits, reward-grade hardening, preference & SFT dataset construction. Priced above measurement — this is the scarce skill.
It proves
The meta-eval thesis, operationalized. The touchstone itself gets certified.
It leaves
A reward the client trusts under pressure. The engine part for Reading 03.
Reading 03the advanced tier

The Loop

An RL environment is an eval the model is allowed to practice against. Policy acts, verifier scores, weights update — ten thousand times.

Engagement three closes the circle. You clone Meridian’s workflow into a sandbox — mock CRM, policy database, simulated borrowers with messy, realistic accounts. You drop in an open-weights model (Qwen-class, ~32B) and run GRPO with the hardened judge as the reward. Every episode: Harper-in-training reads an account, proposes a plan, drafts the reply, gets scored, nudges its weights toward what scored well. Reward-hacking attempts show up — and die against the judge you hardened in Reading 02.

Four weeks later Meridian owns a model that hits 88% task success where the frontier-API baseline sat at 61%, at a fraction of the per-ticket cost, running inside their own VPC. This isn’t speculative: Cursor post-trains models with its own product as the environment, and services shops like Applied Compute and RunRL run this exact loop for enterprises — typically on Qwen models — for workflows like CRM operations. And the before/after proof? Scored on the harness you built in Reading 01. The ladder eats its own tail.

The reinforcement-learning loopA Qwen-class open-weights policy acts inside a sandboxed replica of the servicing stack. Each episode transcript is scored by the hardened judge as the reward; a GRPO update shifts the weights toward high-reward behavior, looping ten thousand episodes. Result: task success rises from 61 to 88 percent against the frontier-API baseline, at roughly ten times lower cost per ticket, with the client owning the weights in their own VPC.QWEN-32B POLICYHarper-in-training · open weightsactsSANDBOX ENVIRONMENTmock CRM · policy DB · simulated borrowersreplica of the real servicing stackepisode transcriptREWARD = HARDENED JUDGEscoreGRPO UPDATEweights shift toward high-reward behavior× 10,000 episodestask success 61% → 88% (vs frontier API)~10× lower cost per ticketthe client owns the weights, in-VPCverified on the Tier 01 harness
Fig. 3 — the loop: practice against the touchstone
You sell
The full loop: environment build, reward, training run, before/after certification. Highest fee, thinnest competition.
It proves
You don't just measure the gold. You refine it.
Lane rule
Sell outcomes to portcos, not environments to labs. The lab-vendor market is consolidating around funded players; the enterprise side buys results and lacks in-house RL.

Same stone, three readings

Each altitude reuses the last one’s deliverable: the harness produces the traces, the traces feed the hardened judge, the judge becomes the reward that trains the model — and the trained model is certified back on the original harness. Fees climb the ladder because scarcity does: many people can build a dashboard, few can harden a reward against gaming, and almost nobody outside the labs can run the whole loop for a regulated business.

Meridian Servicing and Harper are composite illustrations; the numbers are plausible, not measured. The failure modes and methods are the real ones. Anchors: Epoch AI practitioner interviews on reward hacking; Cursor’s self-environment post-training; Applied Compute / RunRL-style enterprise RL services on Qwen-class models.

ΒΑΣΑΝΟΣ — the stone that reads the streak.