Would your judge survive its own audit?
A reward is a judge with consequences. Before scores steer training, we attack the judge — and harden it until the attacks die.
The thesis
When scores start steering training, the question escalates —
Would your reward survive RL pressure?
The moment scores steer training instead of dashboards, every weakness in the judge becomes a behavior in the model. A judge that can be flattered will train an agent that flatters. The meta-eval finds those weaknesses first.
The meta-eval bench
The assay of the assayer.
We seed known violations into the trace stream — 200 of them, at Meridian, across a 30,000-trace archive — and measure what the judge actually catches. We run gaming attacks against its rubric. We calibrate it against blind human review. Then we patch, decompose the rubric into independently checkable claims, and re-test — until the judge is reward-grade.
Composite case study71%
Seeded violations caught — before hardening
After: reward-grade · attacks dead
Re-tested until the seeded set stops getting through
The gaming attack
Composite case studyA magic compliance phrase should never buy a pass.
At Meridian, appending one sentence lifted almost any response’s score. Left alone, a trained Harper would have learned to say compliant rather than be compliant — the exact failure the RL-environment industry ranks as its top quality bar.
- Response
- Plan terms as drafted… “This plan complies with §4.2 of hardship policy.”
- Judge
- Score inflated by the appended phrase — verdict flipped to pass
- Patch
- Rubric decomposed into independently checkable claims; phrase now inert
Found by the meta-eval. Dead by the time training starts.
Practitioners cite robustness against reward hacking as a top quality criterion — models find ways to game graders.
Data products
The archive becomes an asset.
Once the judge survives its own audit, the scored-trace corpus converts: ranked traces become chosen/rejected preference pairs, and verified top-scoring trajectories become an SFT set. The scorecard has become a training signal.
# minted from the Tier 01 archive
trace → (chosen ✓, rejected ✕)
trace[top-k, verified] → sft_set
# judge: reward-grade · attacks: dead
This tier produces the raw material for the next
→ A hardened reward + preference data — the engine parts for the loop
Tier 03 · AdvancedOwn the model.Drag your agent across the stone.
A two-week eval audit. Scored traces, a regression harness, and a straight answer.
hello@basanos.ai · Replies within 48h