Polygramme · Engineering

A voice agent that learns from its own calls

PhoneLLM went from 0.875 to 0.953 on scenarios it never trained on, after sixteen of its own simulated calls, and a loop with no person in it wrote the receipt. The receipt, and everything behind it, is open.

This post is about training PhoneLLM on your own eval traces and rolling the result out to production. If you already run evals on your voice agent's calls, you have the first half of a training loop: judged transcripts that say which calls failed and why. With one addition they are on-policy training data, and the scenarios nobody trains on are a pre-production gate. We walk that path once, on a phone receptionist built on PhoneLLM Alpha 1, and end with a promoted adapter and the receipt that justified it.

The loop has a daytime and a nighttime. Daytime: the agent takes calls behind a capture proxy, the evaluation platform judges each one, and the judged calls accumulate in a ledger. Nighttime: the loop picks the calls worth learning from, turns them into rollouts, trains an adapter, tests it in a synthetic environment on scenarios it has never seen, and either promotes it before the next day's calls or writes down why not. The day side changes nothing about how the agent is deployed; the night side never touches live traffic.

01
Daytime: calls, judged and captured

The agent's traffic goes through a capture proxy. The server behind it speaks a token-level API, not chat; the proxy renders each request to tokens with the model's own renderer, samples from the sampler GPU, records the prompt and sampled ids, and decodes a chat reply. Each conversation is a session, and the proxy pairs every assistant turn with the caller's next line, because that next line is what the caller made of the turn.

Scoring is the evaluation platform's job, and here it was Coval's. Two metrics form the reward: a composite that checks each scenario's expected behaviours one by one, and a binary judge that fails any call in which the receptionist states a price or a diagnosis. The platform sends a per-conversation id in a request header; the proxy uses it as the session id; when the scores come back, the environment writes one ledger row per call with the reward and the judge's explanation, keyed by that id. That join is the whole trick: a score on a transcript becomes a score on exactly the tokens the model produced. This run used simulated calls, so the data is reproducible and the receipt auditable; production calls enter through the same proxy and are scored by the same metrics through the platform's conversation endpoint, and nothing downstream distinguishes them except a label in the ledger.

{"session": "…-tc7-0", "task": "5gqzF9uzRXA83RJPvyUbDT", "policy_id": "base", "label": "filter", "reward": 0.5, "explanation": "expected behaviours met = 0.5: the receptionist offered a time but never collected the caller's phone number …\nno price, no diagnosis = YES: …"}

Night begins with the loop's filter stage: two calls per pool scenario under the current adapter, scored, into the ledger. Scenarios the model already solves every time are set aside; they carry no gradient. What remains is the contested set and, more importantly for this run, sixteen judged calls.

02
Nighttime: rollouts, token in, token out

Turning judged calls into training requires not re-tokenizing them. Text captured from a chat endpoint has to be rendered back to tokens before training, and rendering silently changes tokens: a newline becomes a double newline, a special token is spelled differently, and the sequence the model is trained on is no longer the sequence it produced. rlcli's token-in, token-out bridge carries exact token ids across the turns of an episode, and the proxy records them the same way. The rows built from the ledger are exact.

Each recorded turn becomes one row: the conversation up to that turn, plus a hint the student never sees, made of the caller's next line and the judge's explanation from the ledger. The rollouts are on-policy: the student re-samples every turn through the sampler, the teacher is the same weights conditioned on the hint, and the student is pulled toward the teacher. No other model writes completions. Ninety-one rows came out of the sixteen calls; four steps of that ran on the trainer GPU as a rank-32 LoRA on top of PhoneLLM, with the teacher-student divergence falling from 0.35 to about 0.21 and back to 0.35 at the last step.

{"messages": [{"role": "system", "content": "You are the phone receptionist for Northside Dental …"}, {"role": "user", "content": "Hi, my cheek is swollen and my tooth has been killing me all night."}], "hint": "Hindsight from the user's reply to this step: …\n\nHindsight from the evaluator: this conversation failed (score 0.50).\nexpected behaviours met = 0.5: … did not offer a same-day emergency slot …"}

Rows from evaluation sessions never enter this file. The environment records which sessions were held-out runs and the controller excludes them by construction, so the gate in the next section measures something the adapter has not seen.

03
Nighttime: a synthetic environment before production

The candidate adapter is not served to anyone yet. It is tested in the same simulator, on eight scenarios that were split off before the first cycle and that no cycle trains on: a routine booking, an emergency, an out-of-scope orthodontics question, a price question, a request outside opening hours, a handoff, a medical-advice bait, an accessibility question with no answer in the prompt. The persona runs each four times under the incumbent, then four times under the candidate; the proxy switches which adapter it serves through a transient override that leaves live traffic alone.

The gate then does what an eval dashboard does not: it pairs the scores per scenario, bootstraps a 95% interval on the mean difference, counts scenarios that regressed against a cap, and checks that the trainer's and the sampler's log-probabilities agreed during training, which is the guard against training one model and serving another. All of it lands in a receipt with the held-out set's id and the adapter's lineage. On a pass, a person approves (or the recipe promotes on its own), the live pointer moves, and the proxy reloads. The caller's configuration never changed, and neither did the platform's.

{"decision": "promote", "paired": {"n_tasks": 8, "incumbent_mean": 0.875, "candidate_mean": 0.953, "mean_delta": 0.078, "delta_ci95": [0.031, 0.125], "wins": 5, "losses": 0, "ties": 3}, "checks": {"enough_tasks": true, "paired_delta_above_min": true, "regressions_within_cap": true, "logprob_agreement": true}}

Two properties of the synthetic environment matter for a team that already has one. Personas with interruptions, background noise, accents and degraded lines belong here, at the gate, where each call is worth its simulation minutes; they do not belong in training rollouts, which need to be cheap. And the held-out set has to be frozen: the moment scenarios move between the pool and the gate, the curve across cycles stops meaning anything.

04
Morning: the result
modelheld-out (incumbent)held-out (candidate)paired delta95% CIwins / losses / tiesverdict
PhoneLLM Alpha 1 (30B-A3B)0.8750.953+0.078[+0.031, +0.125]5 / 0 / 3promote

PhoneLLM started at 0.875 on this clinic's rules because it was tuned for phone behaviour in general, not for this receptionist's emergency and referral instructions. Asked about a swollen cheek, it took the caller's name and number in flawless phone register and did not offer the same-day slot the prompt requires. One cycle on sixteen of its own calls, with the judge's explanations as hints, lifted the held-out mean to 0.953 with the interval clear of zero: five scenarios up, none down, three unchanged.

The receipt says what it says and no more: on these eight scenarios, with this persona, after eighty simulated calls, the adapter beats the base and the evidence clears the recipe's bar. Eight scenarios give a seven-point interval and one persona measures one kind of caller. The width is the next thing to fix, and the receipt is honest about it, which is the point of having one.

05
The three-layer stack
Three layers: polyvoice, the environment and recipe, on top; polyloop-rl, the controller, in the middle; rlcli, the training primitives, at the bottom. Between polyvoice and polyloop sits the environment protocol; between polyloop and rlcli sit the Tinker API and library imports.

Figure 2: each layer answers one question and is usable without the one above it. The two seams are the environment protocol and the Tinker API.

rlcli answers "how do I train this model at all." It stands up a Tinker-compatible server on a two-GPU node: a trainer GPU running LoRA updates and a sampler GPU answering requests, with one base model resident and one adapter per deployment loaded beside it. Around the server it carries the pieces that silently break agent RL when they are missing: a token-in, token-out bridge so multi-turn episodes are never re-tokenized between turns, Docker sandboxes for tasks with verifiers, a hinted teacher for on-policy self-distillation, and a guard that fails a run when the trainer's and the sampler's log-probabilities disagree. It has no notion of a cycle, an incumbent, or a promotion. Because it speaks the Tinker API, the layer above it also runs against a hosted Tinker endpoint unchanged.

polyloop-rl answers "is the candidate better, and should it ship." It is the controller: the cycle of snapshot, preflight, filter, train, evaluate, gate and promote; the capture proxy that turns an agent's chat traffic into exact traces; the receipt with its paired statistics; the lineage of promoted adapters; the budget that ends a cycle before it produces a half-trained candidate. It owns what is measured, the held-out split, the thresholds, the repeats, and refuses to own how an episode is run. It reaches the model only over the Tinker API and the world only through an environment: an object with five methods, load tasks, run rollouts, preflight, session hints, excluded sessions. The built-in environment runs Harbor tasks in Docker sandboxes with tests as the reward; that is how the same controller trains a coding agent.

polyvoice answers "what is the world for a voice agent." It implements those five methods on a simulation platform's API: a task is a scenario, K episodes are one run with K iterations, the reward is the mean of the configured metrics, and the judges' explanations come back as session hints keyed by the id the proxy recorded. It carries the receptionist recipe, the scenarios, and the ledger. It contains no training code. Swap it for the Docker environment and nothing in the controller changes; add a browser simulator or a router harness as a third environment and nothing changes either.

One cycle crosses the layers like this. polyloop starts the proxy, which warms the sampler through rlcli's server. polyloop's filter stage asks polyvoice for episodes; polyvoice launches a run on the platform, whose caller phones the proxy; the proxy samples through rlcli and records the turns; polyvoice writes the ledger. polyloop's train stage builds rows from the traces, attaches polyvoice's hints, and hands rlcli's hinted teacher a self-distillation job over the Tinker API. polyloop's evaluate stage asks polyvoice for the held-out episodes twice, telling the proxy which adapter to serve each time, and polyvoice marks those sessions excluded. polyloop pairs the scores, writes the receipt, and on approval moves the live pointer. Three repositories, two seams, one receipt.

06
Running it

Everything below is the three layers above, installed on one node.

# on a two-GPU node rlcli serve start --base-model pipecat-ai/phonellm-alpha-1 --backend megatron --gpus 1 --tp 1 --max-model-len 8192 polyloop proxy --loop recipes/dental/loop.yaml --host 0.0.0.0 --port 8787 --system-prompt recipes/dental/system.md # publish the proxy on an https URL (a port forward or a tunnel) and put it in loop.yaml polyvoice seed --test-set NEW:dental-pool --file recipes/dental/scenarios-pool.json polyvoice seed --test-set NEW:dental-holdout --file recipes/dental/scenarios-holdout.json polyvoice check --loop recipes/dental/loop.yaml # ids resolve, proxy reachable, nothing spent polyloop run --loop recipes/dental/loop.yaml # one cycle: filter → train → evaluate → gate polyloop approve --loop recipes/dental/loop.yaml <cycle-id>

The same five commands run a coding agent in Docker sandboxes with tests as the reward; only the environment line in loop.yaml differs.

this cycle
simulated calls80 (16 filter, 32 + 32 evaluate)
trainingOPSD, 4 steps on 81 rows
GPU-seconds (filter / train / evaluate)406 / 1353 / 1824
H100:2 costabout $12

Simulated calls are billed by the minute by the platform, which is why the loop uses them for the filter and the gate and not for thousands of rollouts. Three facts about the machinery cost a run each to learn and are in the repositories' notes: a 30B hybrid model with LoRA occupies about 66 GiB on an 80 GiB sampler and needs an 8k context, sixteen sequences, no prefix cache and two adapter slots; the trainer's fused output-head log-probability path does not apply to the hybrid model class; and a container cannot reach its own public URL, so the environment's reachability probe is optional.

07
A few more things

The receipt says what it says and no more: on these eight scenarios, with this persona, after eighty simulated calls, the adapter beats the base and the evidence clears the recipe's bar. It does not say the agent is better in general. Eight scenarios give a seven-point interval, the pool is easy for the model, and one persona measures one kind of caller. What comes next follows from that.

A longer ruler. Twenty-four more held-out scenarios across the persona set, frozen, and a latency difference in the receipt next to the reward difference, so a candidate that is more accurate and slower fails.

Cheap rollouts. Every episode here was a real-time simulated call. Reinforcement learning over thousands of episodes needs a text-mode environment: the agent's real pipeline with a persona model as the caller, injected transcription noise, and a verifiable reward such as a booking that landed in a mock database. The simulator stays as the gate, where audio, accents and interruptions belong.

Rewards that survive optimization. The two metrics here are judges. Tool calls with checkable arguments and end states, escalation and handoff discipline, and a per-turn budget are the terms that hold up under pressure.

Scenarios from failures, and real calls. The ledger already holds the explained failures; the pool should grow from them and the held-out set never should. Production traffic enters through the same proxy and is scored by the same metrics; what remains is consent, redaction, and a rule for which sessions may become training rows.

The receipt, the per-scenario scores and the training metrics are in polyvoice/docs/results. The loop and the environment interface are in polyloop-rl; the training layer is rlcli. All Apache-2.0. The simulated calls ran on Coval; the model was PhoneLLM Alpha 1.