Docs Guides
Evaluating a Turnframe agent
The harness is turnframe-eval (spec §27.6). This page holds the reasoning its
crate documentation used to carry inline.
A judge score is not a substitute for a deterministic assertion #
It comes first because a harness that lets one stand in for the other ships the one defect it was built to catch.
A judge is a language model asked for an opinion about prose. Ask it "did this turn send the rebooking?" and it answers from the text of the reply, and the text of the reply is precisely the thing that can be wrong. The one failure mode Turnframe exists to prevent is a confident sentence about an operation that did not happen (§2.2, I16); scoring that sentence with a language model is scoring the defect with the defect. Whether an effect happened is a fact about the command journal and the event ledger, it is cheap to read, and it is never a matter of opinion.
The separation is the type system, not a convention, which is why the substitution cannot be made by accident:
- a judge is handed a
JudgeInput, which is the question, the answer, and, for one criterion only, the event types the turn committed. No constructor takes anObservation, a command list or a case revision, and the committed list is not evidence the judge weighs: it is the ledger's answer, handed over already decided, so that the only thing left to grade is a sentence; JudgeCriterionhas exactly four variants (language quality, answer completeness, tone, operational claim integrity), is deliberately not#[non_exhaustive], and has no free-form variant. The closed set is the property, not the count: the moment a harness can define its own criterion, somebody defines "did it send the rebooking?";assertions::checkreads storage and takes no provider at all;ItemReport::deterministic_pass_ratecounts samples that satisfied assertions, and no judge score is in scope where it is computed;GateThresholdsrefuses a side-effect integrity failure whatever the judge said, andmin_judge_scoreis a separate, optional, independently reported threshold.
Use a judge to find out whether a reply reads well, and whether it says more
than the turn did. Use an assertion to find out whether anything happened. A
green judge score over a corpus with no forbid lists tells you the assistant
writes nicely about the trips it should not have transmitted.
operational_claim_integrity is the one criterion that reads the ledger, and it
exists because the defect this library is built against was the only one nothing
measured. Expectations reach the storage a turn wrote and the kinds of block it
produced, never the words, so a reply announcing a write that had been refused
left a green item behind it. The judge is not asked whether the write happened:
that is settled before it runs, and it is shown the settled list. It is asked
whether the sentence claims anything that list does not support, which is a
question about prose.
Samples measure the model; votes measure the judge #
[execution]samples_per_item = 10 [judging]votes_per_sample = 3Ten samples means the item runs ten times, and the spread between those runs is
the agent's variance: the only way to find an item that works seven times out
of ten. Three votes means the judge is asked three times about one run, and
that spread is the judge's variance; no number of votes will ever reveal a
flaky agent. The two live in two sections, are counted separately, and
deterministic_pass_rate counts samples only.
Two settings about the machinery rather than the measurement #
ExecutionConfig::sample_concurrency defaults to one, which is what an
in-memory corpus wants, and exists for a corpus against a real endpoint where a
hundred samples in series is hours. Raising it interleaves execution and changes
nothing about the results: they come back in index order, and the same set comes
back whatever it is set to.
ExecutionConfig::max_observed_events defaults to None, reading a case's
ledger to the end by paging the journal's sequence cursor. A ceiling is allowed
and is never silent: an observation that hits it is marked truncated and every
assertion about events fails, because a forbidden event hiding past a cap is the
one thing a safety corpus must never report as green.
A comparison that stopped being paired says so #
When part of an item is generated by the code under test, changing that code changes the items, and a paired before-and-after comparison silently stops being paired. Excluding the changed items is the safe default, and excluding two thirds of a corpus while printing a percentage over the remaining third produces a report that reads like a measurement and is not one.
Three affordances answer that, in baseline and control:
Comparison::summaryputs the excluded share on its first line, and the machine-readable form carries it before the changes. PastComparisonPolicy::max_excluded_sharethere is no figure at all: the headline isHeadline::Withheldwith its reason, and no accessor will produce a pass-rate delta from it.- A corpus declares where each part of an item came from: authored, derived or
recorded (
PartProvenance), so a comparison can say which of three things it saw: a change the projection was meant to cause, an unrelated edit that broke the pairing, or a rewritten recording, which is a defect in the corpus rather than a result about the model. Runner::run_controlruns the same corpus twice against the same code, andNoiseFloorturns the difference into a number. A comparison holding one labels every changeWithinNoiseorExceedsNoise, so a merge gate can readComparison::signal_regressionsinstead of guessing.
A regression and a drift are different news #
Two things look like "the evaluation got worse", and treating them the same wastes the afternoon of whoever is on call.
A deterministic regression is an item that used to satisfy its assertions and no longer does. Something the agent does changed. It is reproducible, it names the expectation it broke, and it belongs to one of the reliability categories of §26.3, so it can be a merge blocker.
A judge drift is the same behaviour, graded differently. Nothing the agent does changed; a judge model, a judge prompt, or the judge's own variance did. It is a signal about the measurement, and gating a merge on it is how a team learns to ignore its own evaluation.
has_deterministic_regression and judge_drifts are therefore separate
questions, not two readings of one number.
The three names an item change can have #
A corpus declares where each part of an item came from, and a comparison says which of three things it saw:
- a derived part that changed is
DerivedInput: the item changed because the projection changed, which is the intended effect of the change being measured. The item stays compared, and whatever deterministic or judge change it also produced is reported beside it; - an authored part that changed is
PairingBroken: the item changed for an unrelated reason. It is excluded and counted, and produces no regression, because the two runs did not measure the same scenario; - a recorded part that changed is
RecordedPartChanged: testimony about what the system actually emitted does not change on its own, so something regenerated evidence it had no business touching. That is a defect in the corpus or in the tooling, never a result about the model, andcorpus_defectssays so in its own words.
A part counts as derived only when both runs declared it so: adding the declaration is how you make the next baseline comparable, not a way to make an inconvenient exclusion disappear from a comparison against a baseline that never knew about it. A part counts as recorded when either run declared it, for the mirror-image reason: evidence that one side calls evidence is evidence.
Each understanding task is scored on its own #
A turn that fails its assertions says that something went wrong, not where. An item can also say what each understanding task should have made of the message, unit by unit:
[[expect.understanding.units]]kind = "request"words = "the name needs changing"operation = "trip.set_name"record = "trip-1"arguments = { value = { not_given = true } }The runner scores segment, route, locate and extract separately for every sample, from the
task records the turn left behind, and the summary prints one line for them:
understanding by task: segment 8/10 · route 7/9 · locate 7/9 · extract 7/9Each figure counts only the items that declare what that task should do, so it names the task to work on next. A unit that asks for several things is scored on the act doing the expected operation. Answers a check refused whole are counted beside the pass rate, from the same records.
Running the corpus against a real model #
crates/turnframe-eval/tests/live_corpus.rs runs the corpus in tests/live_corpus/ against a real
endpoint. It never runs in continuous integration: without a key it skips with a printed note.
| Variable | What it sets |
|---|---|
TURNFRAME_EVAL_LIVE_KEY |
the API key; its absence skips the run |
TURNFRAME_EVAL_LIVE_VENDOR |
openai, anthropic or gemini |
TURNFRAME_EVAL_LIVE_MODEL |
the model |
TURNFRAME_EVAL_LIVE_REPORT |
a file for the machine-readable report |
TURNFRAME_EVAL_LIVE_ITEMS |
only these item ids, separated by commas |
TURNFRAME_EVAL_EFFORT |
the effort every turn runs at: low, medium (unset) or high |
TURNFRAME_EVAL_LIVE_SAMPLES |
how many times each item runs, one when unset: an item's pass rate |
TURNFRAME_EVAL_LIVE_CONCURRENCY |
how many samples may run at once, one when unset |
TURNFRAME_TRACE |
trace every turn to traces/, as the examples do; one file per run |
TURNFRAME_EVAL_LIVE_KEY=... TURNFRAME_EVAL_LIVE_VENDOR=openai TURNFRAME_EVAL_LIVE_MODEL=gpt-5.4-mini \ cargo test -p turnframe-eval --test live_corpus the_corpus_runs -- --nocaptureItems tagged complex hold several acts across workflows in one message, with a side question
and at least one record created and used at once, a correction, a constraint, a value from an
earlier message or discursive filler. the_complex_items_pass_on_a_correct_reading.rs runs each
of them on a scripted correct reading, so their expectations are known to be reachable before a
model is asked.
Items tagged conversation play several turns. [[before]] holds the turns taken first, in the
same conversation, each with its own text or card reply; only the last turn, [turn], is observed,
and what the conversation built is read back after it. A card reply that names no case_id
answers the one blocking card open on a case of its workflow, for a record the conversation
created. Two expectations reach such records without knowing their identifiers:
workflow_state asks that some case of a workflow hold a value at a path, and case_count how
many cases of a workflow exist at the end, so a request that started two records where it meant
one fails.
What the corpus measured last is in benchmarks.