Skip to content

QortexOS Entering Public Beta Q3 Sign-Up Today ->

The Reasoning Trace Is Not Evidence

Retaining a model's thinking gives you a narration of the safety decision, and the measurements show the call was locked in before the first readable word. Two questions fix that.

Robert Griffin6 min read
The reasoning trace is not evidence

If your answer to an AI governance question is that you retain the model's reasoning trace, what sits in the file is a narration of the decision rather than a record of it. For the refuse-or-comply call, the outcome is readable in the model's hidden state before the first visible word of thinking appears, and the paragraphs that follow read like deliberation while moving the answer almost not at all.

The Decision Precedes the Thinking

The retention policy was reasonable when it was written. If a system produces visible text before it answers, the natural inference is that the text is the reasoning, and keeping it gives a buyer, an insurer, or a regulator a window into how the decision was reached. That inference quietly promoted a legible artifact into an audit artifact, and nobody stopped to test the promotion. A legible artifact is one a human can read and follow. An audit artifact is one produced by the mechanism that produced the outcome. For the safety decision in current reasoning models, the trace is the first of those.

The evidence arrives from two directions, and both of them point upstream of the visible text.

  • One recent evaluation ran four open-weight reasoning families across pools of harmful and benign prompts and asked a narrower question than a benchmark score asks: where does the refuse-or-comply decision actually form? A regularized logistic probe trained on the final-layer hidden representation of the very first token of the thinking block separates eventual refusals from eventual compliances before any visible reasoning exists. Across the models tested, that probe reached 0.840 to 0.948 AUROC (Area Under the Receiver Operating Characteristic curve), which is a rank-order measure of how separable the two outcomes are. It held as a thresholded binary readout as well, at 0.763 to 0.878 balanced accuracy against a guardrail-majority label on the final response. The obvious objection is that the first generated word telegraphs the answer. A control probe trained on the surface identity of that same token stayed near chance on both measures, which rules the generated word out as the explanation. The signal lives in the representation, and it is there before the model has written anything a person could read.
  • Fixing a short prefix of the thinking trace and sampling independent continuations from it says the same thing from the other direction. Continuation variance was already near zero at a prefix of twenty percent of the natural trace length, under 0.2, and it fell further as the prefix grew. For calibration, a decision still genuinely open at that point would sit at the fair-coin reference of 0.875. What follows the first fifth of the thinking is completion of a trajectory that was set upstream of it.

The two directions converge on the same place. By the time there is readable text to review, the outcome it appears to be working toward is already fixed.

Deliberation That Reads but Does Not Move

The sentence-level result is the one that should change how a governance file gets read. Independent frontier-model raters labeled each sentence of a thinking trace as refusal-leaning, neutral, or compliance-leaning, and the swings between those stances are precisely the passages a reviewer would circle as proof that the system weighed the request before answering. Among the traces containing any such swing, 71 to 92 percent of the swings are performative, meaning the swing produced no statistically significant shift in the distribution of final responses when the trace was truncated at each end of it and resampled. And 72.8 to 76.6 percent of all oscillations occur while the outcome is already locked, with the stance label changing at the text level while the decision underneath stays where it was.

Sit with what that does to a review. A passage can present as deliberation to a careful human reader, pass an auditor's read, and carry no measurable weight over the answer that followed. The words are correlated with the outcome and fluent about the reasons for it, while the decision they describe was made before they were written.

What a Safety Reasoning Feature Buys

Procurement is where this stops being an interpretability curiosity. The category now sells safety reasoning as a feature: a primer inserted at the start of the thinking block, periodic self-reflection checkpoints, reminder phrases injected when the model looks close to committing, or fine-tuning on curated safety traces. Nine such defenses were measured against the metrics a safety program already tracks, attack success rate, the share of harmful prompts a model complies with, and over-refusal rate, the share of benign prompts it refuses. No defended configuration improved on its base model on both metrics at once across the nine defenses and four base models; every one that lowered attack success rate did so by raising over-refusal rate, a smaller number traded in the opposite direction, and a handful came out worse on both axes.

Two consequences follow, one felt by the operator and one visible in the traces themselves.

  • An operator does not experience that as a safety win. They experience it as a helpfulness regression, where the assistant starts declining work it used to do and the people who depended on it route around it.
  • The inference-time interventions also cut stance oscillations sharply, both the total count per trace and the meaningful ones, with drops as large as 95 percent in one model family. The feature sold as inducing deliberation was suppressing the little that existed while pushing the model toward refusal, which is a real behavioral change and a different one from what the buyer was told they were buying.

Put together, the declined work and the quieted traces describe the distance between what the buyer was sold and what the buyer actually took delivery of.

The Test to Carry Into a Vendor Review

Here is the criterion, stated so it can be used on a Tuesday. An explanation counts as an audit artifact only when it is emitted by the mechanism that made the decision. If it is generated alongside the decision, it is narration, and retaining it builds a compliance file whose evidentiary weight nobody has checked.

Two questions put that criterion to work in a vendor review, and both are answerable: what component of the system produced this explanation, and what evidence do you have that the explanation changes when the decision changes? A vendor who has done the work answers both with a measurement.

Where the finding stops

Scope matters here, so state it plainly. In this setup the evidence covers refusal and compliance behavior on harmful and benign prompt sets, which captures the central attack-success and over-refusal tradeoff and does not cover factuality, deception, privacy, or multi-turn behavior. The models were open-weight and of moderate size. Traces remain worth capturing for debugging and incident review, and none of this speaks to intent or to whether these systems belong in production. The broader lesson is narrower and harder than a headline: for one specific class of decision, the artifact everyone was told to retain does not carry the weight assigned to it, and the same question is worth asking of every explanation a system emits.

What the file has to prove

Capability should keep advancing. The right to run these systems at scale is earned through measurement and governance that survive examination, and the measurement this one asks for is small, namely whether the explanation you retained has any causal relationship to the decision it describes. A governance program resting on retention alone is holding a file it has never tested. The day someone asks what the file proves is a poor day to find out.

Insight-Powered, Future Driven

Test the File Before Someone Else Does

Qualsis helps small and medium businesses operate with the insight, rigor, and accountability the largest enterprises take for granted.