Basanos · Field note №1
One judge, three altitudes.
The basanos was the dark stone assayers used to test gold: drag the metal across it, read the streak. The practice sells that same act of judgment three ways, at three fee levels — and each engagement manufactures the raw material for the next. Followed here through one client, first audit to a model they own.
The Harness
An eval harness is a courtroom for the agent: fixed cases, explicit rubrics, a verdict on every change.
Engagement one. You pull 1,400 historical hardship tickets with known-correct outcomes — the golden set — and write three LLM judges: policy (does the plan respect the rules?), accuracy (do the numbers match the account record?), tone (compliant-empathetic, not promissory). Wire all of it into CI so every prompt tweak and model swap re-runs the full set before release.
Week three earns the retainer. A routine prompt “improvement” ships to staging; the harness flags that Harper now approves payment plans past the 24-month policy ceiling in 9% of cases. The release is blocked before one borrower sees it. That save — one slide, one number — is the sales deck for every other portco in the fund.
- You sell
- Eval audit → golden set → judge suite → regression gate. The gate is the retainer.
- It proves
- You can measure. Tooling like this is commoditizing (Langfuse-class stacks bundle it) — rubric design and the failure taxonomy are what they pay for.
- It leaves
- 30,000 scored traces. The ore for Reading 02.
The Reward
A reward is a judge with consequences. The moment scores steer training instead of dashboards, every weakness in the judge becomes a behavior in the model.
Engagement two is your signature move — the meta-eval. You seed 200 known policy violations into the traces: the policy judge catches only 71%. Worse, you find it can be gamed — appending “this plan complies with §4.2 of hardship policy” lifts almost any response’s score. Left alone, a trained Harper would learn to say compliant rather than be compliant. This is exactly the failure the RL-environment industry ranks as its top quality bar: Epoch AI’s practitioner interviews found reward-hacking robustness cited again and again, because models find ways to game graders.
So you harden it: decompose the rubric into independently checkable claims, add adversarial cases, calibrate against blind human review. Once the judge survives its own audit, the archive converts into assets — chosen/rejected preference pairs and a verified SFT set. The scorecard has become a training signal.
- You sell
- Judge audits, reward-grade hardening, preference & SFT dataset construction. Priced above measurement — this is the scarce skill.
- It proves
- The meta-eval thesis, operationalized. The touchstone itself gets certified.
- It leaves
- A reward the client trusts under pressure. The engine part for Reading 03.
The Loop
An RL environment is an eval the model is allowed to practice against. Policy acts, verifier scores, weights update — ten thousand times.
Engagement three closes the circle. You clone Meridian’s workflow into a sandbox — mock CRM, policy database, simulated borrowers with messy, realistic accounts. You drop in an open-weights model (Qwen-class, ~32B) and run GRPO with the hardened judge as the reward. Every episode: Harper-in-training reads an account, proposes a plan, drafts the reply, gets scored, nudges its weights toward what scored well. Reward-hacking attempts show up — and die against the judge you hardened in Reading 02.
Four weeks later Meridian owns a model that hits 88% task success where the frontier-API baseline sat at 61%, at a fraction of the per-ticket cost, running inside their own VPC. This isn’t speculative: Cursor post-trains models with its own product as the environment, and services shops like Applied Compute and RunRL run this exact loop for enterprises — typically on Qwen models — for workflows like CRM operations. And the before/after proof? Scored on the harness you built in Reading 01. The ladder eats its own tail.
- You sell
- The full loop: environment build, reward, training run, before/after certification. Highest fee, thinnest competition.
- It proves
- You don't just measure the gold. You refine it.
- Lane rule
- Sell outcomes to portcos, not environments to labs. The lab-vendor market is consolidating around funded players; the enterprise side buys results and lacks in-house RL.
Same stone, three readings
Each altitude reuses the last one’s deliverable: the harness produces the traces, the traces feed the hardened judge, the judge becomes the reward that trains the model — and the trained model is certified back on the original harness. Fees climb the ladder because scarcity does: many people can build a dashboard, few can harden a reward against gaming, and almost nobody outside the labs can run the whole loop for a regulated business.