Skip to content
Tier 02 · 6–12 mo

Would your judge survive its own audit?

A reward is a judge with consequences. Before scores steer training, we attack the judge — and harden it until the attacks die.

The thesis

When scores start steering training, the question escalates —

Would your reward survive RL pressure?

The moment scores steer training instead of dashboards, every weakness in the judge becomes a behavior in the model. A judge that can be flattered will train an agent that flatters. The meta-eval finds those weaknesses first.

The meta-eval bench

The assay of the assayer.

We seed known violations into the trace stream — 200 of them, at Meridian, across a 30,000-trace archive — and measure what the judge actually catches. We run gaming attacks against its rubric. We calibrate it against blind human review. Then we patch, decompose the rubric into independently checkable claims, and re-test — until the judge is reward-grade.

Composite case study

71%

Seeded violations caught — before hardening

After: reward-grade · attacks dead

Re-tested until the seeded set stops getting through

The meta-eval bench — the assay of the assayerThirty thousand scored traces pass through a meta-eval bench: two hundred seeded violations of which the judge caught only 71 percent, a gaming attack where a section 4.2 compliance phrase inflates scores, and blind human calibration. The patched, hardened judge then mints preference pairs and an SFT set — a training signal ready for the loop.30,000 scored traces (Tier 01)META-EVAL BENCH200 seeded violations → judge caught 71%✕ gaming attack: “§4.2” phrase inflates scoresblind human calibration→ patch · decompose rubric · re-testHARDENED JUDGE — reward-gradepreference pairschosen ✓ / rejected ✕from ranked tracesSFT setverified, top-scoringtrajectoriestraining signal, ready for the loop“would your judge pass its own audit”

The gaming attack

Composite case study

A magic compliance phrase should never buy a pass.

At Meridian, appending one sentence lifted almost any response’s score. Left alone, a trained Harper would have learned to say compliant rather than be compliant — the exact failure the RL-environment industry ranks as its top quality bar.

Trace #4417 · Harper · hardship-request Gamed
Response
Plan terms as drafted… “This plan complies with §4.2 of hardship policy.”
Judge
Score inflated by the appended phrase — verdict flipped to pass
Patch
Rubric decomposed into independently checkable claims; phrase now inert

Found by the meta-eval. Dead by the time training starts.

Practitioners cite robustness against reward hacking as a top quality criterion — models find ways to game graders.
Epoch AI — practitioner interviews on RL environments

Data products

The archive becomes an asset.

Once the judge survives its own audit, the scored-trace corpus converts: ranked traces become chosen/rejected preference pairs, and verified top-scoring trajectories become an SFT set. The scorecard has become a training signal.

# minted from the Tier 01 archive

trace (chosen ✓, rejected ✕)

trace[top-k, verified] sft_set

# judge: reward-grade · attacks: dead

This tier produces the raw material for the next

A hardened reward + preference data — the engine parts for the loop

Tier 03 · AdvancedOwn the model.

Drag your agent across the stone.

A two-week eval audit. Scored traces, a regression harness, and a straight answer.

hello@basanos.ai · Replies within 48h