The method
One judge, three altitudes.
The same act of judgment sold three ways — and each engagement manufactures the raw material for the next. Fees climb the ladder because scarcity does.
The harness — gate the ship.
An eval harness is a courtroom for the agent: fixed cases, explicit rubrics, a verdict on every change. We pull historical tickets with known-correct outcomes, write judges for each failure dimension, and wire the whole set into CI so every prompt tweak and model swap is re-scored before release.
The gate is the point. A dashboard tells you how the agent did; a gate decides whether the change ships. Everything the gate scores is archived — and that archive of scored traces is the raw material for the next altitude.
The reward — score the trace.
A reward is a judge with consequences. The moment scores steer training instead of dashboards, every weakness in the judge becomes a behavior in the model — so before we trust a judge with training, we attack it.
The meta-eval seeds known violations and measures the real catch rate. It runs gaming attacks against the rubric. It calibrates against blind human review. Then we patch, decompose the rubric into independently checkable claims, and re-test until the judge is reward-grade. Only then does the trace archive convert into preference pairs and SFT data.
The loop — shape the model.
An RL environment is an eval the model is allowed to practice against. We clone the client's workflow into a sandbox, drop in an open-weights model, and run GRPO with the hardened judge as the reward — ten thousand episodes of act, score, update.
Reward-hacking attempts appear on schedule, and die against the judge hardened one rung down. The finished model is certified on the harness built two rungs down. The ladder eats its own tail — every altitude raises the stakes on the same judge.
How judges fail
The failure taxonomy.
01
Leniency drift
The judge softens over time as inputs shift; catch rates decay without anyone changing a line.
02
Rubric gaming
Responses learn the rubric's surface features — structure, keywords, tone — instead of its substance.
03
Magic phrases
An appended compliance sentence flips verdicts without changing behavior. Saying compliant beats being compliant.
04
Distribution shift
A judge calibrated on yesterday's tickets misreads today's — new products, new edge cases, new failure modes.
Principles
- 01Seed known violations — measure what the judge actually catches, not what it claims.
- 02Attack before you trust. Every judge gets gamed; find the attack before the model does.
- 03The judge is part of the system under test.
- 04Decompose rubrics into independently checkable claims.
- 05Calibrate blind, against humans.
- 06Certify the loop on the harness that started it.
The full assay: Meridian Servicing.
A loan-servicing portfolio company of a mid-market PE fund. Its agent, Harper, reads a borrower’s account, checks hardship policy, proposes a payment plan, and drafts the reply.
Reading one — the harness. 1,400 historical hardship tickets with known-correct outcomes became the golden set; three judges — policy, accuracy, tone — scored every run in CI. Week three, a routine prompt “improvement” made Harper approve plans past the 24-month policy ceiling in 9% of cases. Blocked in staging. That save was the sales deck for every other portco in the fund.
Reading two — the reward. 200 seeded violations exposed the policy judge: it caught 71%, and appending “this plan complies with §4.2 of hardship policy” lifted almost any response’s score. Left alone, a trained Harper would learn to say compliant rather than be compliant. The judge was decomposed, attacked, calibrated blind — and once it survived its own audit, 30,000 archived traces converted into preference pairs and a verified SFT set.
Reading three — the loop. Meridian’s workflow was cloned into a sandbox — mock CRM, policy database, simulated borrowers. A Qwen-class open-weights model ran GRPO with the hardened judge as reward. Four weeks later: 88% task success where the frontier-API baseline sat at 61%, at roughly a tenth of the per-ticket cost, running inside their own VPC. Certified — on the harness from reading one.
61%
Baseline task success
88%
After the loop
~10×
Lower cost per ticket
Drag your agent across the stone.
A two-week eval audit. Scored traces, a regression harness, and a straight answer.
hello@basanos.ai · Replies within 48h