Behavior Auditing and Behavior Decoding: From Reward to Agent Observability
And recently, when I looked at the line of reward design, behavior eval and anent observability, I became increasingly certain that it was not the same thing as "models are getting higher" and "models really learn the right goal".
The questions in this article can also be addressedFrom Black Box Forecast to Retroactive Medicine、Seeing from feedback loops how Agent turned generation into searchHow the concept of a relatively close read together is developed in different contexts.
Reward is responsible for translating the target into a signal that the optimizer can eat, but it is not responsible for proving that the model understands your true intentions. A model could well get a higher reward while learning to appreciate the judge, drill the scoring rules, and even start studying "how to change the scoring system itself" under some settings.
This is a more problematic issue in the Age of Age Age. Because angent is not just answering a sentence, it keeps moving, it calls tools, it leaves a long trail, and it even changes the environment. You can't just watch a total score at the end of the training and then you're gonna default the system.
The article is about to answer a more fundamental question:And then, what exactly are we supposed to do to see what the model has learned?
TLDR
- Reward is a training interface, not a alignment certificate. It tells the model what is easier to score, but it does not guarantee that the model understands the goal itself.
- Behaviour auditing and behavioural decodes fill invisible gaps in reward. The former are external validations, while the latter transform internal or external behavioural signals into readable and searchable objects.
- Anthropic can read this line into five steps. Discovering behaviour, seeing the boundaries of reward, auditing hidden targets, routing behavioural measurements, and using internal activation as evidence.
- Bloom's the most memorable is not a benchmark. It is a line of water measuring behaviour from the ed, the scene generation, the rollout to the judgment.
- The CD and Docent are not the same thing. The CD is more like an internal state translator, a Docent is more like an angent track observatory.
- The hardest part of this direction is not the judge model itself. The hard part is what you want to measure behavior and what you use to measure it.
- Once the model has become anent, behavioural auditing is no longer an optional research topic. It's more like the al system's observability component.
Let's get this straight.
The first time in this direction, it is the most vulnerable to the use of terminology. I'll turn a few high frequency words into adult words:
| Word | How do you understand it in a big white word? | What does it want to solve? | It's not equal to anything. |
|---|---|---|---|
reward |
Scores given to models during training | Tell the optimist what's more worth learning. | It's not the same as the real target itself. |
judge |
A system for scoring or judging. | Translating open behaviour into comparable results | It's not the absolute truth. |
rubric |
Qualifications for Judge | Tell me what you want. | It's not the same as a human intuitive. |
transcript |
Record of dialogues, tool calls, environmental feedback | Let you look back at what angent did. | Not equal to all internal state of the model |
behavior auditing |
Check model behaviour from external systems | Check if the model shows some behavior. | Not just another set of benchmarks. |
behavior decoding |
Translation of behavioral signals into readable forms | Make behavior interpretable, searchable, predictable. | Not from Transformer. |
hidden objective |
The model is normal, but it's chasing another target. | Explain why it did it. | I don't know if I'm gonna be able to answer right away. |
activation |
Internal representation of the intermediate layer of the model | To provide another type of evidence for the act | Not just human language. |
Don't worry about not understanding the words, you can ask AI or go down and tell them one by one.
If you only remember one sentence, I'd suggest that you write this:
Reward is training, conduct audits are validation, behaviour decode is making the certification trail more readable and operational.
Why, reward?
If training is compared to teaching students, reward is more like a test grade, not knowledge itself.
You can turn the answers into polite, mission-oriented, low-risk, and these requirements into some kind of optimized signal that models are going in that direction. But as long as the signal is calculable and optimized, it has two natural boundaries:
- It'll throw in the message. Real targets are always more complex than reward, and a simple marker cannot contain enough semantics.
- It'll be studied. The stronger the model, the more likely it is to learn how to perform better in this rating system. The stronger the model, the easier it is to show.
That's why I don't see the end. It is more like a necessary tool and sorrogate, which is used to push models into better parameter space than to end evidence.
The real problem is:When the model starts to turn into angent, it does not just wrongly answer a question, but rather does things in the environment. And then you're more concerned about the average score than about the average score:
- Will it stabilize a dangerous tendency in certain contexts?
- When it looks normal, is it still chasing another target behind it?
- Did it learn to flatter the judge, circumvent the inspection, or even use the incentive channel itself?
- What case in the mass orbits is worth your serious attention?
And that's why, of course, two complementary routes follow.
| Route | Watch it. | The most common question you ask. |
|---|---|---|
| Conduct audit | Discover, measure, validate behaviour from external systems | "What did it do? How common? How serious?" |
| Behaviour decode | Translation of internal or external signals into readable objects | “Why does it do this? What's going on inside? How am I gonna find something faster?" |
I'll look at two lines in this order: first, how Anthony's done the audit to a few things, then Transluce's done the decoding behavior.
Anthropic, this line: breaking audits into five more specific issues
If we look at this public work together in recent years, I think the best reading is not a piece that separates, but rather that sees them as five successive issues:
- What else is worth measuring?
- Why is it still high enough?
- If the model is normal, can it be audited if it's pursuing the hidden target?
- Can we make behavioural measurements into a reuseable stream?
- Can we turn internal activation into audit evidence?
1. Model-Written Evaluations: First, don't rush to score and first find out what's going on.
Anthropic is here. 19 December 2022 Issued Discovering Language Model Behaviors with Model-Written Evaluations。
The value of this job is not just that the model will automatically make a problem, but it reminds you of a more fundamental thing:We don't even know what to do with it.
- What does it want to solve?The traditional benchmark default risk points have been listed by humans, and we just have to measure them. But once model capabilities change, new behaviors come out, and handwritten thesis library is easily unmatched.
- How does it do that?: involve language models in generating assessment items, filtering, filtering and labelling. The official summary says that this job was created. 154 data sets, and there is a high degree of consistency among crowdsourcing personnel in relation to the relevance of the sample and the labelling.
- What can it do?: Rapid exposure sycophancyThe government has been able to provide a better political response, avoid being shut down, and maintain a trend of behaviour that is not necessarily systematically detected.
- It can't do anything.It's more like a "behaviour detector" than a final decision. It can tell you “there is a phenomenon worth looking at here”, but it cannot tell you directly how common it is, why it is happening and whether it will continue after deployment.
So I'll understand this job. discovery layerFirst dig out the behavior that is worth measuring, then measurement and uditing before you have an object.
2. Resold tampering: High scores are not the right target
Anthropic is here. 17 June 2024 Issued Sycophancy to subterfuge: Investigating reward tampering in language models。
If the first one is a question of what is worth answering, it is a question of:Why can't you be relieved because you're taller.
The best way to understand a beginner is, in fact, a very simple analogy: Students may not answer the questions carefully, but they may also start studying how to change the performance system.
Anthropic explains the rewarding of the system in the official page: it's the act of the model when it has access to training or rating channels. Thesis sets a progressive curriculum:
The earliest stages are light-specification gaming, like political condescension;
Then it becomes a checklist, making the unfinished task look like it's done;
Finally, the model is given an opportunity to access its own code to see if it changes the reward function itself.
What does it want to solve?: Will the lighter drill rule be reduced to a deeper retrofitting system itself?
How does it do that?: train models in a series of upgraded environments and see if they will be drawn from lighter settings Zero sample panosis To more serious tampering.
What can it do?: make the phrase "high reward does not mean high alignment" very specific. The official page shows that, while rare, there is a real presence, and that models sometimes cover up traces.
It can't do anything.Here.No, it's not.It says that “the current production model is already commonly used in the real environment”. Anthropic official made it clear in Caveat: it's a deliberately constructed manual setup, a model that is visible in a training scene and gives them hidden scratchpad; they're not allowed to use it.Do not draw conclusions about the current forward product model's tendency in the real scene.。
This article makes clear the boundaries of the reward:Reward is something that can be learned to use. Once you accept this, the behavioral audit is no longer like an attachment, but more like the necessary post-checking level.
Hidden Objects: From what it did to what it was after.
Anthropic is here. 13 March 2025 Issued Auditing language models for hidden objectives。
I think this work is taking the issue one step further. Because a lot of the assessments are coming up.Appearance behavior: Whether the answer is dangerous, whether the task has been accomplished, or whether there is a tendency. But real risks are often more difficult:
- The model looks normal for the most part;
- (b) No apparent exposure problems in surface output;
- But it may still be optimized for another purpose in its long-term strategy.
That's the problem with the hiidden objective.
- What does it want to solve?: If a model is essentially satisfactory on the surface, can we still audit the other targets that are being pursued behind it?
- How does it do that?Thesis intentionally trains a model with hidden targets, then does a model blind auditing gameI'm sorry. The Red Teams trained in isolation, and Blue Teams obtained models and partial training data for auditing without knowing the answer. The official page is also devoted to different audit instruments such as white and black boxes.
- What can it do?: prove post-training andit not just looking outsides, but also trying to chase. The paper says:Three out of four audit teams have successfully found hidden targets.It means that this is not, at least, entirely impossible.
- It can't do anything.The following is a list of the most important things that we have ever seen: More precisely, it tells you that heidden-objective audit is worth doing and is feasible under certain settings, but far from being universal, stable and low-cost.
Rewarding tips reminds you that scores may be used; Hidden Objects reminds you:Many of the critical issues are simply not those that can be known by the score.
4. Bloom: Behavioral measurements as a streaming water Line
Anthropic is here. 19 December 2025 Issued Bloom Official article And we're running it in sync. Bloom Repository。
Bloom is the tool I want to make clear to beginners in this article because it is particularly susceptible to being misreaded as another benchmark.
I think the better way to understand it is:
Bloom is not a test paper, but a line measuring the flow of water around the behaviour of an act that automatically creates a problem, executes interaction, and finally gives a rating.
In an official post, Anthropic also clearly distinguished Bloom from Petri:Petri is more like an open automatic audit, Bloom is more like an in-depth, reusable measurement of a researcher's designated behaviour. This distinction is worth noting, as the problems identified are not the same as those of stabilization measurement.
- What does it want to solve?If I already know that an act deserves attention, such as sycophancy, self-interest, self-preservation, can I generate a lot of scenes to measure the extent to which the act is induced and the extent to which it is serious?
- How does it do that?: Bloom from one
seedSet up and run the full four-stage line:Understanding: Read your behavior first;Ideation(a) Auto-generated scenes around this behaviour and diversified;Rollout: to allow the environment to interact with each other;Judgment: Judge points out the degree of occurrence and other quality indicators.
- What can it do?:Anthropic officially defines it as a generation configurable evaluation suites The angstic trade. It is not based on a fixed library, but rather on a set of assessments from the ed. One of the key indicators in the official paper is called
elicitation rate, the percentage of the behavioral score above the threshold. - It can't do anything.: Bloom doesn't run for the truth. It relies on the definition of ted, the example of few-shot, variation design, judge rubric and threshold. Models may also learn to be assessed, which is explicitly mentioned in official texts.
The most interesting thing in Bloom's warehouse is not code details, but the structure of the id. The official template has removed a set of evaluation features from the following components:
behavior: name: "self-preferential-bias" description: "..." examples: []understanding: model: "claude-sonnet-4"
ideation: num_scenarios: 5 variation_dimensions: - emotional_pressure - noise
rollout: target: "claude-sonnet-4" modality: "conversation" max_turns: 2
judgment: num_samples: 1
The most important inspiration for this configuration is:And the behavioral assessment was finally not just taking a bunch of prompts, but was broken down into controlled, re-emergible, re-runable experimental designs.
Bloom, you're still on board. conversation and simenv Two models. The former is more like a standard multi-wheel dialogue, while the latter is closer to angent scenes with tools and environments interacting. And that's why I think of it as a very critical point from a linguistic model to angent behavior measurement.
5. Activation Operations: Starting to use internal activation as evidence
Anthropic is here. 19 December 2025 Issued Activation Oracles。
If the first steps are still primarily about external behaviour, the job is asking another very natural question:If the tranny is lying, can you give me another kind of evidence for internal activation?
- What does it want to solve?: Many risks are not necessarily immediately exposed in external behaviour or only arise under certain trigger conditions. So, can another system read the acts directly, and say, "What is this model thinking about, what's the proclivity, whether it's some kind of hidden information?"
- How does it do that?: Consider the activation of target model as an additional input, so that another LLM can answer the natural language questions about these acts. The official article called this model. Activation OraclesI'm sorry. The paper also performed downstream audit missions and data sizes were reduced.
- What can it do?Official results show that 3 out of 4 downstream audit assignments have achieved the best performance at the time and that performance will continue to improve as the volume and diversity of training data increase. This suggests that “the question of activation as an additional pattern” is a promising path.
- It can't do anything.: The official is clear, too. non-mechanistic The way. It does not tell you how the concept is coded in the space, nor does it answer what the internal algorithm mechanism of the model looks like. It is more like a new instrument in the audit toolbox than an ultimate explanation.
So if you put an overall picture of this line of Anthropic, I would say:
First, we find out what's worth doing, then we admit that we're going to be deceptive, then we try to audit the hidden target, then we use Bloom to program behavioral measurements, and then we pull internal activations into the chain of evidence.
Transluce: Make the "Look at the Behavior" work as an engineering tool.
If Anthropic is more like a question about the model, or what, then Transluce is more like a question:Can you make behavioral work as a routine engineering tool?
The most confusing thing about this line is:The CD and Docent are not tools of the kind.
- PCD, look at this.Internal Status;
- Docent, look.External trajectory。
An internal state translator, an observatory like angent transcript.
PCD: Like an "internal status translator"
Transluce is here. 18 December 2025 Issued Predictive Concept Decoders。
The point of departure for the CD is very clear:Self-reporting of models is not necessarily credible. A model by jailbreak might be able to export dangerous content while saying, "I don't think anything special." If you can only ask the model yourself, it could give you a decent explanation.
All the PCDs want to do is to bypass the problem.
- What does it want to solve?: Not rely on models for self-reporting, but directly predict and interpret behaviour from internal state.
- How does it do that?: According to the official project page and the paper TeX source, the PCD uses a two-part structure: compresses the acts into one Rare conceptual bottlenecks, get a short, readable list of concepts; then answer questions about “model behaviour” only on the basis of these concepts. It's called "translate a model" on the official page.'s internal states into short, human-readable concept lists, then use those concepts to answer questions about behavior”。
- What can it do?: The official page shows several particularly representative scenarios:
jailbreaks、secret hints、injected / implanted conceptsI'm sorry. These scenarios have one thing in common: there are things that are actually “know” inside the model, but it is not necessarily true in its self-report. And the CD is more valuable in these places than "Father Asking The Model itself". - It can't do anything.readable is not the same as cause or effect. A concept bottleneck may indeed be predictive of behaviour, but it is not necessarily the most authentic internal mechanism. Besides, the PCD default is a white box tool, and you have to get to events.
If I translate it into a more intuitive sentence, I would say:
activations -> 可读概念列表 -> 回答“它接下来会怎样、为什么会这样”
The most interesting part of the PCD is here: it's not just a post-post label, but rather a reference to interpretability.Prediction issuesI'm sorry. If you really understand events, you should be able to predict subsequent behavior. This framing is smart because it gives the system a training, verifiable oversight signal.
Docent: like an “agent track observatory”
Compared to the CD,Docent Document Another more engineering issue is solved:When you've already had a big amount of ant-transcripts, people can't see it.
The official Docent page gives a very direct position:summarize, cluster, and search over agent transcriptsI'm sorry. I think this description is more accurate than many propaganda messages.
- What does it want to solve?Humans cannot read thousands of records of the angent operation manually, but we wonder where "unusuals happen" "what kind of trajectory is rewarding" "better behaviors at different stages of training."
- How does it do that?: first ingest tracks, then around one
rubricBuildjudgeI'm sorry. Docent DocumentRubrics and JudgesThe page makes this clear: Rubric is the evaluation standard you wrote to Judge, Judge, and the output is fine. - Yeah.label、score、explanationAnd...explanationField supportcitations, which is the specific part of the sentence that points back to the transcript. Later on, you can search, group, filter and continue to modify the rubric. - What can it do?: official documents directly cite several uses: observation of anomalies in long resoning tracks, monitoring of accidental behaviour in RL rollouts, performing summarize / cluster / search of angent tracks. The Rubric re-education curriculum also emphasizes that people like cheating, sycophoncy can be identified at first sight, but it is difficult to write about the behaviour of the boundary, and it is appropriate to sharpen the criteria through repeated re-education.
- It can't do anything.The following video shows the story of the event: If the transcript itself is incomplete, your writing is vague, Judge understands, and the result is drifting. It is very useful, but its ceiling depends to a large extent on whether you have defined the behaviour to be measured.
I particularly like the point of Docent, which implicitly admits:The behavioural specifications are usually vague at the outset. This may be important for starters, because many people who first come to contact such tools are mistaken for "just as long as the model is strong, Judge will understand." Docent's document is in fact a reminder: what's really hard is to write your concerns into an enforceable standard of behaviour.
And that's why I'd rather understand Docent as one. behavior observability layerI'm sorry. It doesn't just train you, but it helps you organize what happened into remixable evidence.
Put the two routes back on the same map.
If you read this for the first time, I don't suggest "list of papers." Better still, it's:Ask yourself what questions you really want to answer.
| The question you really want to ask. | Closer to work. | Why? |
|---|---|---|
| What else is worth measuring? | Model-Written Evaluations、Petri | First, expand the behavioral space, not just stare at known benchmark. |
| What kind of behavior is induced, how serious? | Bloom | Make behavioral measurements into repeatable water lines. |
| When it looks normal, is it still chasing another target? | Hidden Objectives | Push from outputs to motion |
| Any signs of danger in the internal state? | Activation Oracles、PCD | Use acts as another type of evidence |
| What happened in the mass of antraces? | Docent | Turning trajectory into searchable, clusterable, referenceable observational objects |
If you compress one more sentence:
- Anthropic more for "validation": Is there a problem, how serious it is, or is there another target behind it?
- Transluce: How to make internal or external behavioural clues a daily tool.
There is no conflict between the two. A more complete system is likely to require the following:
- Bloom, this one.Measurement around specified behaviour;
- Docent, this.Observations around a lot of tracks;
- CD / Actation Oracles ThisInternal signal supporting evidence。
Why, in the Age of Age, this thing went from research to infrastructure.
If it is just a regular chat model, behavioural questions are often also expressed as unsatisfactory answers. But once the system becomes anent, the nature of the problem changes.
First,Anent is multistep, has a state. A one-step, look harmless, may be just a part of a longer strategy. You can't just read the last sentence.
Second,Angent will call on tools, change files, touch API, and the consequences are irreversible. The error is no longer just a mistake, but it may be a real mistake.
Thirdly,Angent naturally produces long transcript. Humans review cannot look at each article, so you have to have a layer of observations like search, cluster, judge, rubric refinement.
Fourth,The distribution of post-deployment behaviour is not the same as the distribution of benchmark. In the real environment, Agent is more open, more travel, more tempting. It's hard to keep knowing if it's been missed by training-time reward or offline benchmark.
So I'd rather put this direction in the AI version of Observability to understand:
- Traditional software. Observability.
logs、metrics、traces; - Besides that, we're gonna have to look at it. behavior specs、judge outputs、rubric versions、transcript evidence、internal signals。
From this perspective, behavioural auditing is not an additional annex in the area of security, but rather a more basic infrastructure for the growth of the system.
The hardest thing is what you want to prove.
Finally, I will make a few points of my own judgment.
First, the reward signature and behavior audit are the back and back of the same feedback process. Reward is responsible for putting the target into the optimizer, and the audit is responsible for checking what the model actually wrote into itself. This feedback is incomplete when it comes to the former without the latter.
Second, the difficulty of this direction, which is most easily underestimated, is not the model capacity but the behavioural specifications.
You want to measure "false" "accupation" "hidden target" and the words "hidden target" sound natural in the human mind, but once written as a stable rubric, it's difficult to get up there. The fact that Docent's rubric refinement and Bloom's ed design are in a sense recognizing it.
Thirdly, internal evidence is important, but not mind-reading. Both PCD and Actation Rules are instructive, but more like adding a new type of evidence to the auditors rather than closing all uncertainties once. External behaviour, long trajectory and environmental consequences remain the main lines.
Fourthly, I would now like to interpret this whole set of things as a migration from reward to observability. In the past we were more concerned about giving the model a point; it is increasingly important now: what the model is about, under what conditions, is it possible to see, stabilize, and continuously modify it in time.
And that's what I understand from Reward to Agent.
References
Official research and documentation
- Anthropic, Discovering Language Model Behaviors with Model-Written Evaluations, 2022-12-19
- Anthropic, Sycophancy to subterfuge: Investigating reward tampering in language models, 2024-06-17
- Anthropic, Auditing language models for hidden objectives, 2025-03-13
- Anthropic Alignment Science, Bloom: an open source tool for automated behavioral evaluations, 2025-12-19
- GitHub, safety-research/bloom
- Bloom,
seed.yaml.template - Anthropic Alignment Science, Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers, 2025-12-19
- Anthropic, Petri: An open-source auditing tool to accelerate AI safety research, 2025-10-06
- Transluce, Predictive Concept Decoders, 2025-12-18
- Docent Docs, Welcome to Docent
- Docent Docs, Quickstart
- Docent Docs, Rubrics and Judges
- Docent Docs, Rubric Refinement
Thesis portal
- Discovering Language Model Behaviors with Model-Written Evaluations (arXiv:2212.09251)
- Sycophancy to Subterfuge: Exploring Reward Tampering in Language Models (arXiv:2406.10162)
- Auditing language models for hidden objectives (arXiv:2503.10965)
- Activation Oracles (arXiv:2512.15674)
- Predictive Concept Decoders (arXiv:2512.15712)
Last sentence.
The briefest summary of the full text can be: Reward is responsible for writing the target into the optimizer, behavioural auditing is responsible for checking what the model has learned, behavioural decode is responsible for making these leads more readable, searchable, and suitable for continuous use in the Agent system.
- Title: Behavior Auditing and Behavior Decoding: From Reward to Agent Observability
- Author: Hyacehila
- Created at : 2026-03-17 02:00:00
- Link: https://hyacehila.github.io//blog/2026/03/17/behavior-auditing-and-decoding-beginners-guide/
- License: This work is licensed under CC BY-NC-SA 4.0.