From RL Agents to LLM Agents: Paradigm Shift and Uncertainty Modeling After The Second Half
The AGI is achieved by AI being able to obtain money from human society (if the money still exists), and if that is possible, to raise the ROI. Until it gets 99% of the world's wealth.
Make money from human --> Create value for humanity
The questions in this article can also be addressedAEnvirron: Agent Dev Why do you need an interactive environment layer?、From memory formation to memory governance: the panorama of Agent MemooryHow the concept of a relatively close read together is developed in different contexts.
Model shift and uncertainty modelling after RL Agent to LLM Agent: The Second Half
Human interaction and software engineering developments have been addressing human uncertaintyI'm sorry. From DOS interface to GUI, then to voice and action input, technology allows computers to understand more vague intentions and to serve people accordingly.
The birth of Large Language Model gave machines the ability to understand natural languages, process visual information and integrate information thinking. The human interaction has been accompanied by a new kind of entry: The old operating systems translate vague human instructions into precise computer instructions, while natural languages may become the operational systems language of the new age. LLM Agen is like the early prototype of this operating system.
Before the advent of the language model intelligence, Agent was not a strange concept; in the field of enhanced learning, the training of feedback through environmental incentives, such as AlphaGo, AlphaZero and OpenAI Five, were all part of the context of enhanced learning intelligence. However, experiments have shown that Agent genericization, which relies on environmental incentives alone, is limited. As the complexity of the environment increases and action space increases, traditional RL enhanced learning algorithms become increasingly difficult to absorb.
Language model Agent changed the problem portal. LLM Agent and traditional intelligence are different input spaces and Action Space; the former are more space and less likely to be directly constricted. But the a priori knowledge that natural language and image pre-training provide that Agent does not have to rely entirely on blind testing in the environment, and that uncertainty can be described in more detail.
From RL Agent to LLM Agent, not only " old models are replaced by new models " . More important changes are the migration of learning signals from numerical to linguistic space, and the shift in the mechanisms of broad-based testing from blind testing in the environment to pre-trained a priori, reasoning, feedback and memory.
The First Half Concerned about "How to get a Panorama", the Second Half asked, "How to get this panorama right in the real world, right in use, and sink into a continuously evolving smart system."
Why did RL Agent crash into the wall in the open world?
Enhanced learning is not weak. On the contrary, it has shown extraordinary power in closed environments where rules are clear, well-received and where the state of transfer is stable. Whether DeepBrue, AlphaGo, AlphaZero, or OpenAi Five, they state that RL can be approached or even exceeded human strategy levels in an environment that can be accurately simulated, can be sampled in big quantities, and has relatively clear incentive mechanisms.
The problem is that the open world is not a chess game, nor is it a clean enough arena. In an environment such as Minecraft, intellectual decency is not a limited and highly regulated state space, but rather a world where processes are generated, long-term, task-dependent and rewards are extremely thin. A seemingly simple target — such as “composed iron —” — often contains dozens of pre-conditions and thousands of steps atoms. The ultimate reward only appears when the mission is completed, and in almost all the previous actions the environment is unable to tell the intelligent “Is that the one you just took is in the right direction”? This is the problem of credit assignment in a long-range task.
At the same time, the space for action to open the world is too large. The state is no longer a limited board, nor is the action a discrete object, but rather a hybrid space for continuous perspective control, discrete operation combinations, use of articles, path selection, resource dependence and tool chain construction. Pure RL can easily be degraded here into a cost-effective random walk: it may learn to “jump” or “dig” local behavior, but it is difficult to automatically form an abstract causal chain of “pre-made tools before acquiring resources”.
That is why tiered and intensive learning has long been a high-end aspiration, and it has always been difficult to break the rules in an open world. The problem is not just that it's not deep enough, but it's also thatPure RL systems lack original symbolic abstraction and semantic planning capabilitiesI'm sorry. It is difficult to derive directly from environmental feedback structured knowledge such as “If resources B are to be available, tools A” must be in place. Thus, the open world reveals the boundaries of RL Agent: it is not unable to learn, but is not enough to rely on numerical feedback in the environment when the mission requires common sense, abstract planning and semantic compression.
The Second Half
This section discusses the creation of LLM Agent, from the language model to the language model smart body, based mainly on The Second Half of Shunyu Yao, and more topical issues in the new era.
Why did you enter the Seconda Half -- from research on Model to research on Evaluation? The underlying reason is that algorithms, models and RLs have all shown some kind of generalization. We need to study Model before we solve the problem of generality; now, because the former has acquired the capability to be generalized, we need to move to the study of Dataset and Evaluation to address the real problems in the field.
Before we go back to the second half of the discussion, we will briefly review what we are doing in the first half of the session and why we have entered the second half. On the surface, Pretraining, Scaling Up, Ressoning, Reacting makes it possible to construct language model smarts and they actually work to solve problems. But the reason is RL's Generalization.
Only RL can reach the AGI, achieve a possible complete panorama model -- a long-standing judgement. From DeepBrue at the beginning of the 21st century to AlphaGo / AlphaZero, to OpenAI Five, these were attempts at RL-AGI and did work well. But even in OpenAI's attempt to build a perfect Envirronment to get RL to learn, RL can solve a lot of problems, but it does not work.
The final generalization of the GPT series models explains the puzzle that we lack -- Pretraining. In fact, the environment and algorithms are not the most important part, and Priors, which Pretraining brings, is the core of the generalization. It has nothing to do with RL, but it is the key to building AGI. The only thing that can be achieved is a massive Priors program that will not be able to achieve the final action. Human capabilities are derived from Planning, or Thinking; and test-time compute ultimately brings better capabilities to pre-trained language models and combines with Agent, giving it the possibility of wider spatialization.
Thinking about language space is not feasible for traditional RLs. The vast Action Space leaves the RL algorithm undecided. But now we have Priors, a language that makes pre-decision reasoning possible and effective. The First Half concludes here with a broad model (pre-training language model) and knows how to strengthen its capabilities (RL) through algorithms and apply it to the real world (React). The foundation of the Front Token Protection to Language Mode Agen is laid here, and the next step is to keep it general until all issues are resolved.
In The First Half, we created better algorithms and benchmarking tests, and we kept cycling over them: solving the benchmark tests, and then creating more difficult benchmark tests. Now, RL's Generalization will destroy this cycle.
In The Second Half, it became more necessary to re-evaluate and address real problems. AI has become so powerful over the past few years: it has achieved near-full scores in various calibration tests, won a champion in competitions representing the highest level of human beings, and we have received a near-absolute expert (at a very low cost compared to the human itself). But human society has not changed dramatically. This is the question of how to use AI, which is now the most important issue; and it is rooted in the mismatch between existing assessment techniques and the real world, and in the misunderstanding of how to use AI itself.
We'll give you two very simple examples.
- Now, evaluation techniques are automated, but the real scene needs to be a multi-player conversation, not a response that is long-thinking, not even a response after a multi-cycle decision. LMArena and Tau-Bench have somewhat eased the problem, but, like the hallucination study, the construction of the associated Benchmark does not solve the problem, with the emphasis on replacing all Benchmark.
- Now the evaluation technique believes I.I.D. - an ancient assumption that is already considered a statute. Mandates are carried out independently and then average results, but not in the real world, and it is normal to align closely.
In First Half, these assessments work well and they do improve the model ' s intelligence. But in The Second Half, new assessments — that is, assessments oriented towards real-world problems — must be presented; new common approaches must also be built around these assessments, so that they can again enter a cycle of acquiring higher levels of intelligence to solve real-world problems.
We will move from problem solving to definitional issues: The broadization of algorithms means that we do not need completely new structures and methods to acquire intelligence; from improving indicators to converting the complex problems of the real world into indicators; and from improving algorithms on fixed data sets, such as ImageNet, to upgrading modelling capabilities in real missions. And real-world interaction is the biggest assessment, and the world model will replace the data that are available from the real world and that will lead to the ultimate Generalization of data and models from the world model.
And that's why the modeling of the Minecraft game is not the game itself, like the Dota game, but it provides a testing ground that is sufficiently complex, long enough to be close to the real mission structure: If a smart system can prove here that it no longer relies on purely environmental blind exploration, but instead begins to rely on a priori knowledge, planning, text feedback and external memory, then paradigm shifts already occur.
Another perspective: Era of Exiperity versus the route
The Second Half explains why the paradigm is being shifted from the point of view of RL generalization and evaluation. Almost the same time David Silver and Richard Sutton were in the Welcome to the Era of Experience The first of these is the following:The core issue of AI is moving from "how to learn more from human data" to "how to get angent to act in the world and learn from its consequences".
Their core argument is that the incremental benefits of high-quality human data are becoming smaller, and many of the truly important new capabilities — the mathematical, scientific, complex planning of superhuman levels — are by definition not included in the available data. The next stage of the core data source is Angent's own experience of action, observation, error, feedback in the environment (experity)I'm sorry. They have broken the new phase into four core changes:
- Streams:Agent is no longer a question-and-answer chat model, but is living in a continuous stream of experience, building up knowledge and revising strategies over months and even years.
- Grounded Actions / Observations: Inputs are no longer limited to text, but are actually in the environment — web pages, code implementers, API, robotic sensors.
- Grounded Rewards: The reward is no longer just"Humans think that's a good answer."It is the result of environmental consequences — whether the code runs, whether the test results are better, and whether the task is completed.
- Planning Beyond Human TracesIt is not always human-written choin-of-thought; it is possible for angent to develop internal calculations and planning that differ from human expressions.
This framework is consistent with the goal of the Second Half -To define new and valuable questions and give them value- But the path is different. There is a need to make a clear judgement:This short article is more like a study of the Declaration than a proven law. The best reading is to consider it as a new general outline of how it affects the research we should do.
2024-2026: From perspective to reality
These directional judgments have begun to become a reality in a number of systems:
| System | Time | What does it mean? |
|---|---|---|
| OpenAI o1 | 2024 | The ability to reason is increasingly dependent on RL and reasoning, rather than continuing to stack pre-training data |
| DeepSeek-R1 | 2025 | RL, pure, comes out of reasoning, proves experience can grow new capabilities. |
| AlphaProof | 2025 | Formalized certification provides high-quality validation signals and is the ideal scenario for environmental feedback-driven learning |
| AlphaEvolve | 2025 | Search with AutoAssessor driver/calculator, angent in"Test-evaluation-change"Evolution in the cycle |
| Operator | 2025 | Angent directly performs tasks in GUI, moving from answer to action |
| Gemini Robotics | 2025 | Multimodular model into physical environment, visual-action closed loop |
Of which DeepSeek-R1 is particularly noteworthy: R1-Zero Do a large RL, not a SFT cold start.The pure RL does emerge self-verification, reflection, long COT. But there are also questions about endless remission, or rather, the problem of the choice. This just suggests that:Experience can develop new capabilities, but it does not automatically bring good use and stability. The age of experience is not a substitute for all human data, but a reduction in the necessity of human data — pre-training remains indispensable.
AlphaProof received silver medals on IMO in 2024, including the most difficult topics by only 5 players. The core of its success is not how strong RL is, but how strong it is.Once quality environmental feedback exists, learning from experience can quickly go beyond pure human data.I'm sorry. Formalized certificate provides an enforceable, repeatable, scalable authentication signal - - That's the comparison."People think you're a good answer."Much stronger.
Route: There's no consensus on this.
Around"What's the smallpox that fills human data? Board", there are at least three competitive routes:
- Silver / SuttonRL + experience + standard reward. The core conviction is... Reward is Enough—All objectives can be expressed as maximizing the award of cumulative target amounts. Sutton even called LLM in 2025"♪ The world's instant obsession ♪"(a passing fad)。
- LeCun: World Model + JEPA. Builds a world model of coding physics, causality and temporal evolution, which in abstract terms predicts the future in space. By the end of 2025, LeCun created AMI Labs to bet on the path, and the core argument is that the human brain can understand the world with little data and that architecture is more important than data.
- Hassabis: a pragmatic compromise. DeepMind is used in practice in combination - AlphaProof with Gemini + AlphaZero, Gemini Robotics' Mixing Large Models and Control Strategies.
The answer may be a dialectic.The pure RL route underestimates the value of pre-training -- that is, Priors, brought by Pretraining, has made possible the generalization of RL in linguistic space, as has been demonstrated before this paper. But Silver and Sutton are absolutely right about the core cycle of anent moving in the environment, getting feedback, using environmental consequences to reward itself.LLM provides a priori and semantic planning capability, and RL provides continuous optimization driven by environmental feedback - the integration of the two is the possible intelligent route. This is also the direction that the mixed structures of Plan4MC and Voyager, which are to be discussed below, are already being tested.
Whatever route wins, one thing is consensus:The route of relying solely on the available human data is approaching the ceiling. The difference is just what to make up for. And the Second Half and Era of Exchange are the same answers from different angles: defining questions is more important than studying the way they are solved -- - As long as you have a perfect Rewarder that can be close to human needs, RL can take the model to that position.
Paradigm evidence in an open world: Plan4MC / GITM / Voyager
In the search wave using LLM-enabled intelligence, three distinct but inspiring technological routes have evolved in academia. These three routes represent how to combine LLM a priori with different levels of action space: the "Mixed Structure" (RL Bottom-Dipple + LLM Skills Spectrograph) represented by Plan4MC, the "Wext Action Closed Ring + Memory Retrieving" represented by GITM, and the "Code-as-Policy" + Autocourse" represented by Voyager.
Plan4MC: LLM for top planning, RL for bottom execution
Plan4MC is a typical representative of the combination of LLM ' s senior planning capacity with RL ' s bottom-level continuous control capacity. It does not simply declare RL obsolete, but rather makes a very important division of responsibilities:LLM is responsible for dismantling macro-targets and RL for strengthening and implementing skills at the bottom.
To avoid delays and logical illusions associated with the online call of LLM in complex environments, Plan4MC, through LLM, produces skills-dependent Skill Graph at the offline stage, solidifies atomic skills and their pre-positioning into a graphic structure; mission planning is done through a mapping search when online implementation takes place. The design is essentially to say that the truly difficult part is no longer “how to learn from zero all the actions”, but “how to transform world knowledge into an enforceable hierarchy”.
More crucial is the issue of incentives. Plan4MC uses pre-trained visual-linguistic models to shape intensive awards, so that RL no longer has to wait for a thin reward signal at the finish. This suggests that even where RL is still needed, the system has begun to use pre-training to alleviate its most fundamental dilemmas. In other words, Plan4MC didn't deny RL, but proved it. RL needs to be wrapped up in a stronger afore-mentioned and clearer stratification to work effectively in an open world.
GISM: Text interface closed loops and external memory
If Plan4MC retains RL as the core of the bottom nerve controller, then Ghost in the Minecraft (GITM) represents a more radical shift: It almost completely abandons the bottom gradient update and constructs a pure text interactive closed loop.
The key approach of GITM is two things. First, it maps the bottom key mouse operation as a limited structured natural language action, allowing the environment to be re-expressed from pixels and control signals as semantic interfaces. Second, it allows LLM to take direct advantage of the Internet knowledge and mission dependency, to gradually dismantle ambitious targets such as “access to diamonds” into structural sub-target trees and to re-strategize continuously through feedback.
The real turning point is the way to learn. GITM does not accumulate experience through parameter updates, but rather summarizes successful action sequences as long-term text memories, and is included in external databases, and is re-used in future missions by retrieving the generation (RAG). The blogger adds:It replaced the parameter gradient with text memory and retrained with empirical retrieval. This is one of the most critical changes from RL Agent to LLM Agent: Learning begins to move from inside model weights to readable, searchable, groupable external memory.
Voyager: Code is strategy and auto-course
Voyager shows a stronger form of the current LLM intelligence: action space is neither continuous control nor restricted pure text action, but directly generates well-developed performance codes for Turing.
Compared to the GITM, which each time allows the model to generate a step-by-step text action, Voyager allows GPT-4 to directly write JavaScript Snippets to call the API environment to control the intelligence. Codes naturally support complex control streams, for example while Looping,if-else Branches and comboable functions call. A code can run for several minutes without the need for a model to intervene repeatedly, which means that the strategy is no longer a one-step exercise but can be packed into a reusable, combustible, long-term accumulation program unit.
Further, Voyager has built up automatic curriculum mechanisms: It presents new tasks on the edge of “comfort zones” on its own, based on current resource status, terrain and historical experience; model writing codes, executive codes, reading environmental feedback and interpreter error reporting, and self-correction. Once the code is validated successfully, it is written into the external skills pool and will be reused in the future by zero samples through vector retrieval. At this point, learning is no longer primarily a re-enactment of parameters, but rather a reflection of the fact that it is not a new one.Horizontal expansion of the skills pool, accumulation of code strategies and non-parametric deposition of experience.
Synopsis: What does the open world prove?
Plan4MC, GITM and Voyager are not three isolated papers, which constitute a very clear direction for evolution: From allowing RL to continue working in an open world to translate the environment into an understandable semantic space of LLM, to directly write strategies into enforceable and reused codes. Together, they testify to a larger judgement:The main paradigm of Agent is moving from a numerical error in the environment to a hierarchical system of language, code, feedback and memory.
Returning from open world to language space: REAC / TOT / Reflexion
Open world research has demonstrated the need for paradigm shifts, but it has not yet fully explained how they occurred in linguistic space. What really made this clear was the work of Rect, Trey of Thoughts and Reflexion. They are not subsidiary techniques, but rather key nodes of the Agent core mechanism in the linguistic space.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao. arXiv:2210.03629, 2022.
Core InsightThe uniqueness of human intelligence is that it can seamlessly combine mission-oriented action with linguistic reasoning — a close synergy that enables human beings to learn quickly about new tasks and to make sound decisions without the environment or information that they have not seen before. However, before React, the reasoning capabilities of LLM (e.g. COT) and operational capabilities were studied separately. COT can carry complex reasoning but can easily create a cumulative sense of illusion and error because it is completely enclosed in the language model and lacks interaction with the outside world to verify facts; and a purely operational approach, while interacting with the environment, lacks reasoning to develop and adjust plans. The core contribution of Rect is to put both together.Interstapoment integration- Make LLM alternately generate Thought and Mission Action (Action) in the same trajectory, creating a cycle of Thought-Action-Observation.
Why is this the beginning of Agent?:React created for the first time a full sense of the language intelligence -- thinking -- closed circle --
- Delineation-driven action (Resoning to Act): Logic trails help model action plans, track current status, and deal with anomalies. The model is no longer blindly implementing actions, but thinking before acting, if you want, and then thinking about the reasoning itself as action and feedback in linguistic space, so that a simple set of ideas can be used to understand the common bottom of reasoning and action.
- Action to Reason• To introduce observations into the reasoning process by accessing real information through interaction with external environments (e.g., Wikipedia API, web page, compiler) and to correct the language model’s own intellectual illusions.
- Explanatorys and humansThe following are some of the reasons for this: The trail of reasoning makes the entire decision-making process humans readable, diagnostic and trustworthy. This is also the premise of Human in the Loop -- one must understand what Agent is thinking to effectively view and edit in the middle.
The reason for the REACH is that LLM has a strong a priori knowledge of language. Otherwise, in the reasoning of introducing external knowledge in multiple steps, it is extremely difficult to find the right next step. It is because of pre-training-induced Priors that models can show common, flexible and efficient performance with minimal context-learning samples.
Limitations and EvolutionThere is a typical error pattern in Rect - models generate previous ideas and actions repeatedly, and fall into a chain of reasoning that cannot be taken out. This reveals a lack of pure react in the depth of reasoning. One effective improvement is to combine React with pure COT: when ReAct does not find an answer within the prescribed step, back to COT to allow the model to be reasoned with internal knowledge many times; and to use ReAct to get information from outside when CT cannot produce a stable answer. This complementary strategy leads to a deeper insight -If reasoning and action were not mistaken, the use of external information should always be better than reliance on internal knowledge alone. In addition, for complex missions with large action spaces, REAct needs more demonstration samples to learn, which can easily exceed the length of the context - and this is one of the driving forces behind the Memoory study.
The Nature of the Age Perspective:React the core problem to be solvedHow to make language models from generator to actorI'm sorry. The pure COT is thinking but not acting in the language space, and the traditional RL Agent is acting without thinking (at least not in the language space) and relies on a large amount of expensive manual feedback data. REAC learns strategies in a lower-cost way - because the decision-making process requires only the language of the reasoning process. It harmonizes the two: thinking in linguistic space, acting in real environments, calibrating thinking with environmental feedback, and then thinking to guide the way forward. As the paper foresees, language will play an increasingly critical role in interaction and decision-making as LLM develops as a basic cognitive mechanism. This thought-action-OMS cycle is the most basic structure of a language intelligence - LLM provides intelligence, Thinking provides reasoning, Action provides interface with the world, Feedback closes the whole circuit.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan. arXiv:2305.10601, 2023.
Core Insight: The standard LLM level of token self-regression is essentially System 1 - fast, linear, intuitive decision-making. The core inspiration for TOT comes from the two-process theory of cognitive science and the classic AI search paradigm: Introduction of System 2 Care for LLM, allowing models to beExplore multiple lines of reasoning, self-assessment and, if necessary, retreat。
From Delineation Policy to Agent ThoughtToT is ostensibly an escalation of reasoning (from the linear chain of the COT to the tree-like search), but it has the whole Agent idea behind it -
- Planning: Breaking the problem into a number of thinking steps (Thought Steps), each of which is a semantic intermediate rather than a single token. The particle size of the thinking is determined by the task — it can be a word, a calculation, or an entire paragraph written.
- EvaluationLLM acts as its own inspirational function, which values each intermediate state. The inspiration in traditional searches is either handwritten or trained, and tot for the first time allowed LLM to do this by language reasoning.
- Search and Back (Search) & Backtracking): Implement BFS or DFS on the think tree, allowing models to retreat when they find the current path to be unworkable, breaking the limit of self-regression to create a "path to black".
The Nature of the Age PerspectiveTot and SC-CoT share the same thing: they are not just reasoning techniques, but they are transposing the classic AI search-assessment-decision Agent loop into language space. COT is linear action in linguistic space, SC-CoT is parallel sampling in linguistic space, while TTT achieves it in linguistic space.Planning, assessment and retreat- A complete decision-making loop. This, together with the ability to reflect, represents two core dimensions of the language intelligence from generator to problem solver:TT focuses on planning and searching during generation, and Reflexion focuses on reflection and improvement after the results are generated. Together, it is the full capacity that the Agent system must have — planning, action, reflection and re-action. TTT can also be considered as rect in linguistic space.
Reflexion: Language Agents with Verbal Reinforcement Learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao. arXiv:2303.11366, 2023.
Core InsightRL: The traditional RL is learning by updating the weight of gradients, but it is expensive and inefficient for language intelligence. Reflexion suggests a fundamental paradigm shift -Update with a natural language reflection alternative parameter, converts the gradient signal in the RL to a semantic signal. Agent no longer develops experience through a strategy of optimization through reverse dissemination, but through linguistic self-reflection.
Learning from parameters to memory learning: The ReformAgent structure consists of three modules:
- Actor: LLM-based strategy model, which is responsible for interacting with the environment to generate action tracks (such as COT or REACH).
- Evaluator: Assess the quality of the trajectory, output-target incentives - which correspond to the traditional RL incentive function.
- Self-Reflection: Core innovation. It magnifies the thin target incentive into a specific, operational natural language reflection, answering where there was a mistake, why, next time what to do, and also reflections on what to put into scenario memory (Episodic Memoory).
Three.Action, evaluation, reflection, memory, improvement, actionThe anecdotal closed ring. The policy of Agent is not coding in the weight of the model, but in the cumulative reflection text - this achieves a tactical optimization that does not require fine-tuning.
The Nature of the Age Perspective:Reflexion solves the problem of language intelligenceLearning issues: How to get Agent to gain transferable experience from failure without updating the parameters. It moved the traditional RL credit allocation problem from numerical to linguistic space — no longer what parameters should be increased, but what actions should result in failure and what should be done next. This is the direct expression of pre-training that makes learning in linguistic space possible. At the same time, Reflexion has also expanded the source of Feedback - feedback is no longer limited to external environmental incentives, and model self-assessment is in itself an effective feedback signal.
From method to main line
If three routes in the open world tell us that the traditional RLAgent paradigm has been recasted by layers, text interfaces, code strategies and external memory, then React, ToT and Reflexion further illustrate: This rewriting is not an accidental engineering technique, but a technicality.Rebuilding the minimum closed loop, search and learning mechanisms of Agent in the language spaceI'm sorry. REACT Solving the closed circle, ToT Resolution Planning and Backwards, Reflexion Solutions Language Learning; Together, they make LLM Agent after The Second Half a truly sustainable evolution system, not a model that only talks.
From RL Agent to LLM Agent: Reordering responsibilities and hybrid architecture
Now.
What does a full linguistic intelligence contain? From the perspective of the CoALA structure, it may consist of the following components:
- LLM(intelligence): LLM for Next Token Predation, which is widely pre-trained.
- Thinking(Reasoning)The reasoning before action is the key to human resolution of complex problems and another important factor in the evolution of LLM Agent towards universality.
- Action(Feedback)Implementation, feedback and further reflection on the next multiple rounds of action are ways in which humankind can accomplish complex tasks.
- Memory(Learning)• Long-term memory and learning skills to address complex issues over a long-term period.
Now we can think about what the current smart structure is doing:
- ReActAs a pioneering work in Agent technology, it combines LLM, Thinking, and Action and its counterpart Feedback, acting after reflection, responding on the basis of feedback from the results, and making a very good generalization of complex issues.
- CoTTime Scaling, CoT greatly enhances the capabilities of the model, which come from Thinking and LLM on the one hand, and Action and Feedback on the other.
- RL Agent: Agent technology before the language model appears, which lacks LLM, Thinking and Memoory, and relies on environmental feedback for Reward learning, only with Action and Feedback.
- LLM: Remove the LLM after the CTT, it's a source of intelligence, but not everything (at least not now).
- Human in the loop: It's hard to be used in traditional RLs, but the language model Agent can handle this problem very well, and manually can easily find where the reasoning went wrong and modify it.
- Reflexion: from modifying parameters to modifying hints, this is an attempt at a long-range memoory and a reflection on the source of feedback. Feedback can be from outside or from the language model of self-reflection.
- Tree of ThoughtsToT's role is to allow structures that would have been one-way reasoning to be thought-backed, both in terms of reasoning tactics and in terms of intelligent construction. ToT can be understood as the expansion of the CTT, an Action and Feedback in language space.
In this sense, today ' s language intelligence is no longer lacking most of the core parts. LLM, Shining, Action and Feedback have been fully explored and validated, and even Learning has begun to appear in non-parametric forms such as Reflexion, RAG, Skill Library. What is really pending is not the existence of these components, but how to allow them to work together in a long and stable way in the real world.
Future
From the discussions that have taken place earlier, it can be observed that, with the exception of Memoory, the core components of the language intelligence - LLM, Shining, Action and Feedback - have been fully explored and validated. Existing generation models and their ancillary components are already sufficiently powerful to be able to solve complex problems. The core issues of longer-term research are therefore twofold:
- Better Model: At The Second Half, algorithms and models are strong enough. The next step is to use RL to extend existing technology to issues that really work for humanity and build a better and Useful Model. This requires a rethinking of the environment and assessment.
- Dynamic Memory: Construct a layered dynamic memory system and learning mechanisms that will allow Context to always contain only the most critical information and which is an essential additional component of existing technical constraints.
Environment and Evaluation
Better Model is built on the premise that you find the right environment. After RL generalization, Evaluation is more important than algorithms. This environment requires two conditions at the same time — useful and scalable. Real physical and human environments are too expensive to access scale data; virtual worlds and game environments, though cheap, are difficult to migrate to real scenes and intelligence is limited to closed environments. Therefore, interaction with the digital world (Internet, code, software) is a better option — it is both real enough and affordable in scale. In the longer term, the world model would be the environment for concluding all environmental studies.
With the environment, the next question is how to assess. There are three pathways to the assessment: manual assessment (most accurate but expensive), machine assessment (cheap but of poor quality) and manual design-based assessment (between them). In the case of Collie and Webshop, they generate rules for modelling the Self-Evaluation Prompt and also provide rules-based external evaluation - which is in line with the thinking of Rect and Reflexion. But well-designed rules themselves require a great deal of knowledge in the field, and that is an inescapable price.
And, to go further, we can use the intelligent itself as an environment -- to study the interaction between the intelligent. But at this point the assessment becomes more difficult: both social models and social simulations are difficult to measure with a single indicator. One possible direction is to use smarts that are bound by a lot of rules as assessors (from Rule-based Envirronment to Rule-based Agent), perhaps to assess what we thought was only humans can judge. But it still requires a great deal of knowledge in the field to design the rules.
This constitutes a circular iterative structure: environment generation data, assessment-driven improvements, and improved models require better environment and assessment. One of the fundamental difficulties in generating type of task is that it is difficult to determine strictly the end state, and we do not always know when it is done. But as long as the model is getting better and solving the problem we want to solve, it's worth it.
Memory
Memoory is the most unsolved component of a language intelligence. How to construct a truly effective long-range memory system — allowing Agent to accumulate and retrieve experience across missions, across time scales — remains an open and difficult issue. We chose not to discuss it here, leaving it for more specific follow-up exploration.
Therefore, it is not the disappearance of RLs that is more likely to occur, but rather the reordering of duties: RLs are returned to lower layers of control, continuous action and local optimization, LLM leads high-level planning, semantic reasoning, tool/code generation and memory synergy. After the Second Half started, the cap on Agent will increasingly depend on how the environment, assessment, feedback engineering and external memory are organized into the same system.
Appendix: Agent Dev Industry Practice Guide
Whether it is Plan4MC treatment of multi-modular incentives or Voyager ' s use of code as a strategic space, they leave several reusable clues to the project ' s fall: action space is layered, environmental feedback is readable, and memory libraries cannot simply build up logs. These elements are not the main lines of this paper, but are likely to be the first problems encountered by the industrial scene.
Action Space Decoupling
In the landing of complex intelligent systems, it is inefficient and high-risk to force a single LLM to complete the whole stack mission from the intent to understand to pixel-level control. The action space decoupling is the most sophisticated design principle of the day:
- Cognitive and logical isolation: The intent recognition of the task, complex reliance on dismantling and long-range planning is given to LLM or a multi-intelligent collaboration framework; the bottom-prone command sequence does not allow LLM to be exported directly.
- Code / Script Processing Determination Tasks: If the industrial environment already has a stable API, such as the cloud control table, Web Dom, terminal CLI, then LLM is allowed to generate Python or script to carry circulation, unusual capture and control streams.
- RL Treatment of unstructured physical controls•: Enhanced learning strategy networks are placed in bottom control only in settings where there is a lack of accurate API and where feedback from high frequency sensors is required. LLM is better suited for high-level semantic commands.
Textual and Feedback Project for Environment Watch (Observation Textualization) & Feedback Engineering)
Since the LLM's original mode is natural, the design quality of the environmental observation interface directly determines the intellectual limit of Agent.
- High-dimensional state requires semantic downgradation: Do not throw unprocessed arrays or disorder logs directly to LLM, but use the parser to translate the remaining resources, node health, pedal structure, etc. into structured JSON or refined text.
- Close-ring feedback is good enough.: When environmental execution is reported as wrong, the bottom implementer should not simply return
False, and returns the background of the stock track, missing dependency, resource margin. The ability of the model to modify the logic depends largely on the availability of these feedback texts.
The ongoing evolution of memory-based architecture (Memoory-based Continuous Learning)
A fine-tuning of the costs could also trigger catastrophic oblivion. For most engineering scenarios, the installation of non-parametric memory systems for intelligent bodies is a continuous learning path that is easier to maintain.
- Experience externalization and RAG retrieval: Systematically document successful ideas for addressing specific business issues, effective API calls templates and validated code clips, which then are organized into searchable fields of skills.
- Self-validation and Sedimentation Mechanisms: The data written in the memory library must be validated first. This workflow is worth sinking only if the implementation logic has truly changed the state of the environment and achieved the established sub-targets.
- Skills set: When planning a new mission, the model should proactively search the memory library and combine historical skills as with the standard library function. The more business experience is accumulated, the more opportunities the system has to move from “mechanical implementation” to “re-use old experience to solve new problems”.
In summary, the a priori knowledge and non-parametric memory of the large language model provides a more realistic engineering route for the long-distance, open-world mission. Rather than replacing the traditional intensive learning package, it disengages senior planning, environmental feedback, tool implementation and experience deposition, allowing each layer to use a more appropriate approach.
- Title: From RL Agents to LLM Agents: Paradigm Shift and Uncertainty Modeling After The Second Half
- Author: Hyacehila
- Created at : 2026-03-09 04:00:00
- Link: https://hyacehila.github.io//blog/2026/03/09/from-rl-agent-to-language-agent-v2/
- License: This work is licensed under CC BY-NC-SA 4.0.