Own the model that runs your workflow.
RL fine-tuning of open-weights models — Qwen-class, GRPO — in a sandboxed clone of your actual workflow, with your hardened judge as the reward.
The loop
An eval the model is allowed to practice against.
Every episode: the policy reads an account, proposes a plan, drafts the reply, gets scored, and nudges its weights toward what scored well — ten thousand times. Reward-hacking attempts show up, and die against the judge hardened in Tier 02.
The sandbox
Your workflow, cloned. Your data, staying put.
We replicate the stack the agent actually lives in — mock CRM, policy database, simulated customers with messy, realistic accounts — so the model trains on the job it will do, not a benchmark that resembles it.
Training runs inside your VPC. Nothing leaves: not the traces, not the policy corpus, and not the weights — because you own them.
Prerequisites
You can’t buy rung three without rungs one and two — the loop is built from their deliverables.
- In place: Scored trace corpus (Tier 01 — the harness)
- In place: Hardened, reward-grade judge (Tier 02 — the meta-eval)
- Next: Training run — GRPO against the judge, in the sandbox
Results
Composite case studyFrontier API baseline
61%
Qwen-32B after GRPO · hardened reward
88%
+27pts
Task success vs frontier API
~10×
Lower cost per ticket
Weights: yours, in-VPC
Certified on the Tier 01 harness
This is a real services category
Cursor post-trains models with its own product as the environment; services shops run this exact loop for enterprises — typically on Qwen-class models — for workflows like CRM operations. Most organizations have no in-house RL expertise, so they buy it.
- RunRL
- Osmosis
- Applied Compute
- Adaptive ML
- Cursor (self-environment post-training)
Landscape — not clients or endorsements
Start at the stone — book the audit.
Rung three is built from rungs one and two. Every engagement starts with a two-week eval audit.
hello@basanos.ai · Replies within 48h