Every AI agent in production produces traces — every attempt, every correction, every success. Almost all of that signal is thrown away the moment the conversation ends.
The teams that keep it have to duct-tape a stack together to use it: a training API, an RL engine, sandboxed environments, importers for their own logs. Each piece is its own install, its own config, its own glue code. Several well-funded companies now sell that duct-tape as a closed, metered service.
We collapsed the loop into one open-source CLI.
serve turns a GPU box — rented or your own — into a Tinker-API training server, built on SkyRL, Berkeley’s open-source RL stack. import converts your agent’s chat logs (OpenAI or Anthropic format) into training data and pipes them straight into training. train runs supervised fine-tuning or reinforcement learning — including agents acting inside Docker sandboxes, scored by tests that must pass.
Nothing leaves your machines. The weights are yours.
GSPO — the sequence-level loss that keeps agent RL stable — normally costs two forward passes and a backward per training step through hosted APIs: fetch logprobs, compute the loss client-side, ship it back. On your own GPU it is one fused call, exposed as a flag: --loss gspo.
fused 1-pass 12.9s/step · 586 tok/s
2-pass custom 16.8s/step · 451 tok/s
23% faster steps · 30% higher throughput
The benchmark script and raw results ship in the repository, so the two numbers above are an afternoon to reproduce, not an act of faith.
A Qwen3-0.6B trained on GSM8K with rlcli train rl --loss gspo on a single rented L4: flat near 1.6% accuracy for twelve steps — the cold-start regime, where a model that never succeeds gets no signal — then ignition, then a noisy climb to a ~60% plateau with a 76% peak. Forty steps, about a million tokens, $2.65 of GPU time.
The rollout text tells the same story more vividly than the curve: by step sixteen the model is checking its own arithmetic mid-answer — “Wait, is there an error here? Let me check.” — where at step one it could barely finish a sentence.
The same CLI trains agents, not just answers. A Harbor-style task is a Docker container, an instruction, and a test script; the agent gets a bash tool, works the task over multiple turns, and the test’s verdict is the reward. Any task you can express in that format becomes a training environment, running on your local Docker daemon — no cloud sandbox account.
Our acceptance run: a 4B agent working terminal tasks across concurrent local sandboxes, learning by fused GSPO, zero leaked containers.
At small scale, hosted per-token training APIs are cheaper than renting a GPU — our $2.65 demo run would have cost about a dollar at published per-token rates. The crossover comes with volume: the bigger the model and the longer the loop, the earlier owning the stack wins, and for large models the per-token spread runs roughly an order of magnitude.
Continual learning — retraining on your traces every night — lives on the far side of that crossover by definition. That is why the closed platforms meter it, and why an open alternative matters.
Your traces → verifier models that grade them → failures synthesized into new training environments → continual RL → a model that keeps getting better at your agent’s job.
We are building that loop in the open. The foundation ships now.
pip install polygramme-rlcli · Docs: docs.polygramme.com · Source opens with the launch — watch this page.