Skip to content
Tier 03 · Advanced

Own the model that runs your workflow.

RL fine-tuning of open-weights models — Qwen-class, GRPO — in a sandboxed clone of your actual workflow, with your hardened judge as the reward.

The loop

An eval the model is allowed to practice against.

Every episode: the policy reads an account, proposes a plan, drafts the reply, gets scored, and nudges its weights toward what scored well — ten thousand times. Reward-hacking attempts show up, and die against the judge hardened in Tier 02.

The reinforcement-learning loopA Qwen-class open-weights policy acts inside a sandboxed replica of the servicing stack. Each episode transcript is scored by the hardened judge as the reward; a GRPO update shifts the weights toward high-reward behavior, looping ten thousand episodes. Result: task success rises from 61 to 88 percent against the frontier-API baseline, at roughly ten times lower cost per ticket, with the client owning the weights in their own VPC.QWEN-32B POLICYHarper-in-training · open weightsactsSANDBOX ENVIRONMENTmock CRM · policy DB · simulated borrowersreplica of the real servicing stackepisode transcriptREWARD = HARDENED JUDGEscoreGRPO UPDATEweights shift toward high-reward behavior× 10,000 episodestask success 61% → 88% (vs frontier API)~10× lower cost per ticketthe client owns the weights, in-VPCverified on the Tier 01 harness

The sandbox

Your workflow, cloned. Your data, staying put.

We replicate the stack the agent actually lives in — mock CRM, policy database, simulated customers with messy, realistic accounts — so the model trains on the job it will do, not a benchmark that resembles it.

Training runs inside your VPC. Nothing leaves: not the traces, not the policy corpus, and not the weights — because you own them.

Prerequisites

You can’t buy rung three without rungs one and two — the loop is built from their deliverables.

  • In place: Scored trace corpus (Tier 01 — the harness)
  • In place: Hardened, reward-grade judge (Tier 02 — the meta-eval)
  • Next: Training run — GRPO against the judge, in the sandbox

Results

Composite case study

Frontier API baseline

61%

Qwen-32B after GRPO · hardened reward

88%

+27pts

Task success vs frontier API

~10×

Lower cost per ticket

Weights: yours, in-VPC

Certified on the Tier 01 harness

This is a real services category

Cursor post-trains models with its own product as the environment; services shops run this exact loop for enterprises — typically on Qwen-class models — for workflows like CRM operations. Most organizations have no in-house RL expertise, so they buy it.

  • RunRL
  • Osmosis
  • Applied Compute
  • Adaptive ML
  • Cursor (self-environment post-training)

Landscape — not clients or endorsements

Start at the stone — book the audit.

Rung three is built from rungs one and two. Every engagement starts with a two-week eval audit.

hello@basanos.ai · Replies within 48h