What Problems Are We Really Solving When We Build AI Agents?

What Problems Are We Really Solving When We Build AI Agents?

Hyacehila

It's written in front.

It's gonna be a long blog. I'll talk about this heading in a more random order. There are many links to the text: examples and extensions of views. They are not responsible for drawing conclusions for me, but give the reader more material to judge for himself.

This article is more oriented towards technically-born readers and does not interpret the underlying concept too much, and I have some background information that has been reduced to a single article that has been briefly spoken. It will not be a mere opposable article, with many examples and arguments. Better leave some time before you're ready to read it.

This is really a hard article to start. I have a lot of things I want to write, but it's not very good. I've been writing for a long time to find a little cut. It doesn't sound so logical, but I've tried.

I believe this piece will be very useful to readers who are doing AI Agent, but it cannot begin with a clear idea of what the harvest will be. It's a piece of personal experience, not a lesson.

Let's start with the story.

A big leap.

This post is part of our special coverage of the World Summit for Children, held in Beijing in July and August. In the last six months, the AI Agent field has undergone a major leap forward. If you have to draw a specific time node, it can start with OpenClaw's blast, until now.

The leap here has been a bit of history, and they are somewhat similar in their journey. AI Agent has had some hard-to-reach quality changes over the past six months, but in the Internet itself, and in the ToC and ToB products, numerous products and technologies, Demo, have been introduced like dumplings. The e-mails have been running numerous AI Agent products and dropping out of the air in a few months; Ali’s vast number of functional units are being reorganized, crucified by an internal change that eventually leads to no action; and byte even combines soybean bags and flying books, it is hard to imagine that one day the two software will come together. Almost every Internet company, even non-Internet companies, has its own lobster, which is being raised by users from different backgrounds. 60 years ago, you were practicing steel, and now you're practicing Agent.

Why? We've got new wonders?

There's a strange thing about technology, but is OpenClaw the strange thing? Claude Code was opened to all users in 2025, and Codex was one step late in October 2025, and the CA, Cline came far earlier than OpenClaw. The Agent, who makes the decision to help users modify and test the code, is not unusual for programmers, and the world has not changed. OpenClaw didn't break any technological bottlenecks, humans didn't enter the world after the singularity, but only a new product from a water-to-tunnel with enhanced silicon-based intelligence.

If there are no wonders, why did we make this big leap?

Because it's burning up.

Ordinary users were actually left behind in the past few years of the AI wave; they saw the AI nuclear bomb explode again every day, but they never understood what the AI round was, and what it would bring; they were anxious, but they didn't have an exit, or an antidote, until OpenClaw appeared.

What happens after the fire?

In the pre-OpenClaw era, almost all AI Agent Dev was monopolized by Researcher. Only Researcher and VC have a little idea of the principles of AI Agent and what they might be able to do. A number of programmers working on the front line are also trying to reach AI Agent, a group that likes to embrace new technologies. For those of untechnical origin and management who leave technology first line early, access to AI may be limited to bean bags and reports from those under hand, and reporting must be done by reporting. Despite the rapid technological changes, the company has maintained its years of silence.

Now OpenClaw has broken into these management's vision, and for the first time they know that AI is already able to do so much. Without much technical detail, it might be felt that it could solve everything now. A beautiful illusion came out. The pressure from the top comes, big AI projects are set to compete for new user access, internal AI-efficient KPI/OKRs are issued, employees from all backgrounds start learning AI Agent to fill the relevant staff gap, and the entire Internet industry appears to have been lit.

An Ai big leap came.

Now, six months later, new user accesses have stabilized, a large number of internal efficiency tools in various industries have largely failed after various attempts, and the development of the AI capability boundary has become clearer. The leap ended in silence.

AI Agent has spent six months walking through a lot of industries for five years.

If you want to know why OpenClaw is a fire, reference"BettaFish, Mirofish, OpenClaw and Agent's Trust Border"

Chapter I

This article is for everyone who did AI Agent during the Big Jump, whether you're interested in it or under KPI pressure.

I'll try to answer three questions:

  • When you're building an AI Agent, what are you doing?
  • When you want to solve a problem in a real scene with AI Agent, you have to solve your own problem?
  • Why did an AI Agent project succeed and why failed?

Many people can't answer, not because of their talent and ability. Every person will have external and internal pressure during a period of great leap forward, and a great deal of energy will be devoted to achieving it faster than to thinking about these issues. These questions have also been covered by three or four different answers over the past two years, each of which was correct at the time.

There is a reason not to be on the individual side. The real lesson of a big leap is not that steel is not made, but that indicators replace targets. The stove was lit, the production was reported, and as for the calibrated, no one asked. The efficiency projects I have seen on the flash drive have access rates, call numbers, coverage, and very few “who does this job now, how much faster, how wrong it is to be found”. The indicators are well achieved and the tools are not used. When the indicators themselves replace real goals, no one will ask why.

Now we have time to think.

These are the questions that cannot be seen just by looking at the six months. The six months of virulent activity are more like waves on the surface, and the direction and changes of currents need to be considered at a longer time scale. Before answering the three questions before us, let's go back to 2024 and see the way AI Agent came.

Time to come.

This chapter was originally written here, and was then opened in a single article:"AI Agent's Timeline: Where did we think it was a bottleneck?"I'm sorry. It has been walking over the past two years in chronological order, and has been divided into seven layers: Prompt, RAG, WorkFlow, Tools, Context, Models and New Words, and Evaluations. Each layer is a speculation about where the bottlenecks are, and a set of tools, terms and projects have emerged; some of them are deposited into today ' s infrastructure, others are investing a lot of resources and finally the road is not working.

And then the next part will repeat what's in that chapter. If you're good at Agent, you don't look and you don't have to read, you can just look at the article's subheadings; if you want to know where these judgments come from, you can go over the road and come back.

Chapter II

Before starting to present a clear view, there is a second chapter.

It seems that the talk of so many past projects has been somewhat detached. So many engineering practices ahead are actually encountering a variety of landing bottlenecks, and we're trying to solve the problem and improve the final Agent capability. In terms of results, we have made great strides and have sunk many useful programmes. They all answer these questions, and it should be clear how each of the technologies that have been used to answer them is a little bit.

Try to think, no matter what you are building.

  • When you're building an AI Agent, what are you doing?
  • When you want to solve a problem in a real scene with AI Agent, you have to solve your own problem?
  • Why did an AI Agent project succeed and why failed?

Some of the pens were left behind in the introduction. Prompt, is it really over? There are still many plugs and Skylls on the market about building a better Prompt. What are they solving? What do we weigh in engineering? Context Engineering is covering a much more than condensed range. What are we going to do in different scenarios? Evals is very important, and the article talks about Evals' techniques, but it's not so clean in the works, how do we decide what to do and how to change it?

More about the following is a few personal experiences and a summary, and I stepped on a lot of pits and thought about it after. And of course I'm learning the practice of others. Let us start with some personal experiences and studies, by answering three questions I have raised, and by talking about what we should do when we go to an Agent in the future.

Some thoughts and judgment.

♪ To think of things that have gone before ♪

A lot of the ideas that follow are based on personal experiences, and here is a brief chat about what I used to do and what I saved.

I started working on safety at the end of 2025, almost in early March 2026, and then I saw some friends trying to try and get results in this field. My discussions will all build on those experiences and not on the more distant past and future.

I did the safe Agent is quite simple, and the goal is to build a one that can help us find a loophole in the code library. The central purpose of the holes is to build a high-quality data set for digging, to consider the Training model and to enhance its ability to dig holes. The universal intelligence in 2025 is getting stronger, but it's still weak on this special mission, choosing to be trainning at that time was a non-mistakeful choice.

Of course, if we look at the future, we are doing a study that is very easily replaced by universal modelling capabilities. The team of Startup should think about whether the next upgrade of the model will eliminate or add value to its own product, and doing research is sometimes a small Startup, but we're doing our own VC.

The specific work is not extensive, and a multi-intelligence system containing gap information collection, source code static analysis, sludge stream modelling and self-checking, and certification of the CodeQL engine. It's only been developed for two weeks, and the rest of the time is changing the small issues. The main focus is on how to express a state of detour and to design a set of good tools to achieve external interaction; to do some static analysis, validation and self-reflection; and to configure a CodeQL as the ultimate validation machine to give a True or False answer. Collecting a bunch of data, washing, running, beating tag, doing lessons, eventually getting a little higher. This is a standard Post-Training Pipeline.

A lot of experience comes from l3yx, and at almost the same time he's studying AI Agent for auto-defense and has done a good job at the TCH Smart Infiltration Challenge. He first developed a very complex Multi-Agent system using LangGraph, making various angles of human penetration an independent Agent and tool, and then adding a lot of penetration SOP and tool descriptions. This is a common agent that brings together a great deal of expert experience, but expert experience is limited to indicative words; it is a human being with a big tool, but a description of the tool and feedback can lead to explosions. So he proposed an Agent Framewok, a Dynamic Workflow, to bring expert experience and process to light; and another, a similar call optimization for Programme Tool Calling, to find a balance between old MCP and pure Skill Script.

The story is only halfway down here.

TCH, I3yx, take it out in the second. Cairn, a programme that is completely different from the previous generation of multi-intellectual bodies. A Dispatcher + shared Blackboard architecture, the smallest dispatch unit being a Codex Agent, used to make and validate assumptions. The movement control system is structured around the final target, and the entire system is a direct search in an unlimited space. Blackboard only contains Fact, Intent and Hint, which allow Agent to record the discovery, future targets and human hints that are available to achieve communication.

No process is fixed, all generated dynamically by Agent based on questions and simple initial tips; multiple intelligences injected by functional and hint processes are cancelled and all tasks are created only when running; all forward knowledge RG and Skills are removed, believing the model has been understood after so many years.

This is a very simple structure, but sometimes it is more difficult to design it than a complex one. A model that generates the emergence of, and matches, a structure that happens at this moment.

Is it a front-to-back infusion of knowledge or a back-to-back feedback? What information should the model get?

In that section of the RAG, I mentioned one point:RAG is designed to be one-way.I'm sorry. We spent two years doing the "how to get knowledge in" very finely, but few people did the "how to get the results back." It's an asymmetrical. In the section of the Function Calling, we talked about SWE-agent and ACI,A well-designed tool and feedback signal came up with amazing power after combining it with Loop.I'm sorry. The advance of model capabilities and autonomy of Agen at the C end of 2026 is the result of these stories.

Recall two examples of safe Agent. I've done code-digging Agent, extracting forward-drive knowledge from external information and some hard-coded processes and tips, obtaining back-dip feedback from independent static self-checking and certification of the CodeQL engine; the first generation of l3yx systems based on LangGraph, which relies on hard-coded processes and knowledge, and on the results of the operation of the penetration codes; and Cain, which abandons knowledge and processes and relies on knowledge within the model and penetration results to accomplish its mission.

After the AI giants had finished pre-training on the scene where the holes were being dug, I made the Training thing that was worth it. A universal model, which relies on a universal intelligence body, now achieves far better results than the Harness I did. The internal knowledge of models is fully adequate in a growing number of scenarios, and it is worth weighing whether it is spent on training to expand internal knowledge or whether it is supplemented by external injections. (p.s.) So think twice before you start training the small models.

In six months, the evolution of models has eliminated the space before Harness and knowledge were injected into existence. A very simple structure, coupled with a broadly comparable model, brings final performances that go well beyond the type of knowledge injected and process bound.

It's clear that I'm here. The progress of the underlying model will not stop, and the progress of the model will certainly eliminate some of the process constraints and the space in which knowledge is injected. Models are becoming more and more knowledgeable and less knowledge of the world is worth being injected into.When you're trying to plan the process, inject the rules, you might want to see if you have a good feedback signal for the system.I'm sorry. This may be the information that the current model needs more.

As the core of the quality of feedback, the ACI is important, and we have mentioned it once before. If feedback is the one in front of the zero, then research the ACI of each tool and how these tools should be called, the zeros in the back. **Tools is never a list of tools, but AI Agent can see and influence the world. ** Put every MCP Tool on the ground, but absolutely not enough. The context costs of a large number of calls also need to be addressed, as MCP patches, and we have a Programme Tool Calling; if we are willing to believe in model capabilities, then we might leave space for models to write their own scripts (Skill Scripts). I3yx used them in two generations, and they did good.Go to sharpen each Tool description, adjust the cover of a Tools, plan the feedback that the system can give and send back to Agent, and think about how to call Tools. Function Calling was right, Scripts is right now, but the future is not yet coming.

The feedback loop is an interesting topic,"Looking back from feedback, Agent How to Turn Genesis into Search."And when you talk about LLM's candidate generation, you can see how the system can search, validate, cut and choose.

Is a multi-smart body natural? How human division of labour became the Agent division of labour

In the Workflow section, we talked about the success of the workstream framework and the failure of the multi-intellectual framework, and in the Harness, we talked about the great success of the Single Age programme, and about the Subagent's re-use after he had lost his body.The benefits of multi-smarts come from isolation, not from human role-playing.I'm sorry. Should we be doing a multi-mixer? If so, what should we do?

This is a structured decision-making exercise, which will briefly review what our classic architecture is before we begin our discussions.

  • Pure workflow system: n8n, Coze; a low-code structure
  • Pure multi-intelligence systems: MetaGPT, AutoGen; communication systems are at the core
  • Multi-tip smart body bound by the workstream: LangGraph; no freely communicated multi-twistle body but with a separate workPrompt and Node flow, one medium State
  • Autonomous intelligence with appropriate structural constraints: Cain, Dynamic Workflow; process becomes free, constraints become loose, but minor constraints bring about better communication, at least not multiple roles.
  • Full-owned smart bodies and sub-agents: Claude Code, Codex, PI; Master Agent and Subagent

It is simply a matter of understanding the differences in thought that the Agent structure is difficult to discern. The core of understanding the differences between them may be what people inject into the system in the process of development, and what they leave to AI to make its own decisions.

Humans need to divide their roles because of some of the natural limitations of humans as carbon-based intelligence. People ' s cognitive bandwidth is limited, and it is impossible to know the full range of things and know the details of the various areas; they are also constrained in their energy and their ability to work efficiently for long periods of time. So we have done complex tasks in the form of multi-person collaboration, which brings together the sum of your abilities and the cost of information and awareness exchange.

But the same capacity boundaries and energy ceilings do not exist for silicon-based intelligence. Putting a human intelligence different from human intelligence into an organizational structure that has grown because of human limitations is tantamount to bringing together human communication losses. That's probably where MetaGPT and AutoGen failed, and the success of the CCC doesn't mean it's the right way.

From the perspective of the Startup, there's a very realistic bill. Dismantling a mission into Planner, Researcher, Coder, Reviewer, and putting in a layer of communication protocol, does not create new capabilities; in many cases, it simply dismantles the work that a strong model could have done more expensive, slower and more difficult to assess. The context, reasoning and tool capabilities of the model are upgraded a little, and these people are likely to be the first to be eaten. A product that relies on “More Agent Collaboration” for storytelling should ask itself: what is left of the product after removing the role and replacing it with a stronger Single Agent, with enough Token, tools and feedback?

However, purely multi-smart bodies are not worthless. The key is not whether Agent is imitating the human sector, but whether identity, local information and independence are part of the mandate. Social modelling, market behaviour, group collaboration, and dissemination of information are issues that are studied in terms of how multiple subjects interact with each other with incomplete information: different subjects see different worlds with their own goals, memories and consequences. At this point, all roles are placed in a single Agent, and they are usually only produced in a text that is like a multi-party dialogue, which is difficult to really maintain an independent and evolving state of affairs. The core value of multi-intellectual bodies may not be in common office collaboration, but rather in modelling this less than complete information world.

The title of this section is in fact finished with the answers, and here we will make a few more suggestions on the selection.

Let's get out of the way, and think about the nature of the problem you're doing from the first principle: what kind of input you want, what kind of output you want to give, what principles you want to follow. Ask yourself why it is being used, what it is trying to solve, and never answer “xx is doing it”. And never let the structures of human collaboration kidnap AI, always based on the problem itself and on the limitations of the future's silica intelligence.

For the vast majority of internal process tools and C end-products, a workflow system that has access to various sectors, modules and which allows input and validation is efficient and practical. A little LLM is completely enough, and there's no need to be disgruntled because it's not high enough. Many tasks are waterlines themselves, and it is the proper waterline that matters most. Of course, the Workflow system does not mean that you should choose LangGraph directly. There are processes in human collaboration because of the division of human functions, but AI does not understand anything, and it does not necessarily need to be divided by the job itself, not by the function. The pipelines in the development of the game are a good example.

And the only choice that is worth mentioning is that of Agent, which we have divided into two categories. When the objectives are clear, but the processes are vague, it should be chosen. The two types of multi-stellar body are actually the same: a context-segregated and parallel accelerated exploration. For the implementation scenario that is ahead of the mission, the different intelligent bodies are themselves the same, but they are assigned to different tasks.Agent's pouncer structure is only mission-free, no role.

Cain is an example. The last generation of it made the various angles of human penetration independent, each one of which had its own hints, its own tools, its own location; in this generation, all these roles disappeared, and the smallest of them were returned to an unidentified Codex Agent, Dispatcher, just to send a mission, and everyone would do the same. The division of labour is gone, parallel and faster than the previous generation.It's a parallel, it's a profit. It's not a division of labour.We have often considered these two matters as one thing in the past two years.

The difference between autonomous intelligence is that the structure of communication is different from the degree of restraint.Move Tradeoff to a position, and make more.

This section has answered more fully the question of multi-smart bodies and the technical selection of Agent. For an executive mandate, how to design the segregation and exchange of information is more useful than defining 100 Codes; only when identity, local information and independence are themselves the subjects to be studied will the role cease to be packaging, but rather models. Workflow will always coexist with autonomous intelligence, with a wide range of tasks in the world, with AI having changed capacity boundaries and naturally flexible structure choices. The architecture itself already contains the corresponding Context management method: Workflow constraints write Context into code, multi-smarts separate multiple Context, and then combine normal compression and recall, which is all about Context Engineering.

I'm here."Enterprise AI Why is the master card piloting?"It talks more about this from the perspective of Enterprise and about where the unsuccessful projects at Agent come from, and it's interesting to look at it.

The system's ability depends on the human ability to express? It's more than knowledge.

We have just spoken about focusing on feedback, not on research into the injection of knowledge, and now come to the face. Harness is going to inject us a lot of things: basic Prompt, system status, mission objectives and constraints, MCP and Skills tools. A lot of injections are necessary, and what difference does it make from the injections we talked about?

Coding Agent has two features worth talking about. One is Plan Mode, which used to be community-based typologies and is now an Agent frame, and the other is the community-based Grill-me Skill, which is extremely hot. Their core hints are the same: to find the shortcomings I have said and to ask questions until we reach consensus. Is this the place where knowledge is injected? Nope. This is a alignment of norms and hidden habits that are not coded and not hinted into.

No RAG's thinking is limited, and many tasks require much external knowledge simply to describe the mission's objectives and boundaries themselves. And humans are not AI. We are not. /list The full list of related matters can be automatically identified. When AI Agent became so almighty, the ability to own the Agent system began to depend on the ability of humans to express themselves. It is difficult to express a clear and complete intention for a highly complex system.

As the model gets better, Plan actually absorbs the Grill-me, but absorbs need to find that balance.

We talked about Prompt, we talked about language; we found out it wasn't necessary; and now it's a problem with the definition of the target, and it's back again. The model's autonomy has been increasing, and Loop is looking at a goal that is not clearly defined, and there's no point in it, and people don't know what the goal is, and Agent is just lucky.

Let's see how safe Agent is in dealing with this. I3yx, the first generation of systems, whose goals and constraints are all written phrases, is indeed an expert, but the ability to express is ultimately limited and it is impossible for him to write them all at once. Cain is only concerned with infiltration of this ultimate goal, and only simple hints can be inserted into people, all inents are created and modified by the AI system itself. That is the value of Cain: the ultimate goal of penetration can be easily defined, the hint is random, does not create faulty constraints, incent is generated by system integrity, and is not limited by the ability of people to express themselves.

Cain has at least one set of tasks, and it's been able to exchange all the information, so it's working. But it's not even a matter of time, Agent, much.

Is your Harness supplementing the capacity gap, clarifying the mission itself, or is it adding to the trust deficit? If the model really doesn't understand, it's necessary. But if it's just because it doesn't do what you think, try to trust it once.The project is to release the model capability, not to press down the model ceiling.

For a Coding Agent, the knowledge worth inflecting includes technology warehouse, infrastructure, regulation, experience, architecture, description and indexing of the code warehouse itself, business habits, personal style. And of course, it can't be done blindly, how to fill it in, how to make sure it works, more than piles of stuff.

The angle of the question description is not the same in the WorkFlow as in Agent. The restructuring of the system, the modification of internal hints or the enhancement of the hints are all part of the work, and it is difficult to make a few recommendations that must be made, and it is generally useful to think about this in the context of your own structure.

Even if AI becomes stronger, people can never be removed from the system. In complex systems, people need to help AI filter information, find the right, project-compliant, personal taste information. Pure freedom will run on the road to the shit mountain like the GPT browser and Claude's C compiler.

Evals, what are you gonna do? Why would you say this is better?

The last floor of the road I came to say, "Don't skip Evals, leave it to the Anthropic blog." The chapter II also left a question: how do we decide what to say and how to say it? This section is to repay the bill.

When we build an Agent, what we do must be a system engineering, and it requires systematic thinking. The assessment of the system must be an essential part of the picture, of every complex system around you, of Evaluation and Trace, which is in some way or another. Of course you can leave it because sometimes Demo First, but leave the interface and remember that you didn't do it before.

Evals is the core driver of the next phase of development in the era of AI Agent, even earlier, the existence of ImageNet, which has allowed every algorithm progress of CV to be based, and the constant introduction of LLM Benchmark, which has been updated to fill new models, and Evals, which has provided direction for technology development. As long as Evals is itself of the right stage, we can continue to be back-overs at low cost; Evals can also locate problems, attribute errors, compare different versions of complex systems with A/B test, and make it easier to fix them; and Evals can have back-tests, so that we do not have back-overs. That's the value of Evals.

What should we modify with the evaluation? Not every module is worth changing, but only controlled small step changes, A/B test and an iterative approach can lead to a better future. We have to fix a lot of things, and there is no standard answer to this question.Every person who's been in the field of human rights should have their own Belief, and then make a bold judgment, and then make a real correction.I'm sorry. I can talk about a little experience.

Workflow is often easier to trace than Agent Loop, which sometimes gets a result from A/B test and a mountain-mounted log that can be analyzed and tracked by trace. If we only have final results, they don't seem to make any difference, but if we try to locate the problem and the attribution, Tracing is valuable. Thinking about Workflow after understanding the question is a good answer, or do Demo and some tests. What kind of system expansion is often the most important thing to be judged boldly, which requires you to understand the business itself.

The design of the forward-injection knowledge, the back-in-the-back feedback design, how the hints are designed for each part of the system are at the heart of our optimization. Feedback and hints are generally based on A/B test, and the assessment of the forward-injection knowledge is worth thinking. Whether every article in the knowledge system is of sufficient quality, whether all articles cover the knowledge we need, whether the retrieval system works, and what results and operational results are given. Knowledge systems themselves need to be optimized over time, not developed once, and who is to optimize and who is to be held accountable. The failure of retrieval is sometimes simply a lack of access, which is easy to locate, and who will judge whether the search itself is correct? This is a more worthy issue, and different scenarios have different problems.

Agent Benchmark is being environmentalized over the years, from a topic to an environment where it can run, where it is both an implementation and a sentencing place. SWE-bench run test, tau-bench look at the database, the rating is not hanging out, it's the environment's own property. And this is the same thing that ACI said before: a well-designed environment, a cleaner feedback signal, easier to evaluate, two things actually involved the same project input.Can you comment on that? It was decided when you designed the environment.

As for the choice of the project, my experience is probably as follows:

  • What? Only things you'll decide on. Whether to change the model, whether this change in Prompt is better or worse, whether the tool description is useful or not.
  • How much? Small and steady is too much better and more complete. 20 examples of how to run every day are much more useful than 500 for six months.
  • How often are they evaluated: after changes that affect judgement. Models are subject to evaluation, because many of your conclusions from the last edition may have been invalidated.
  • When should it be possible to leave the evaluation: the period of exploration and the prototype period could be avoided, but it was important to know that the debt was being owed and that the account would sooner or later be repaid.
  • Indicators are the proxy for the goal, not the goal itself: Indicators themselves also need to be assessed, and it is increasingly a common problem for Reward Hacking to turn the pursuit of indicators into the goal itself.

Finally, two indicators are mentioned. The code scene is common. pass@k, run k at least once; and tau-bench pass^k The question is whether you can run k every time. The former is about capacity, while the latter is about reliability. This difference is particularly important when you're doing Autonomy: how much autonomy you dare give depends on how much you know about reliability. One. pass@5 It's beautiful.pass^5 A system of terrible sights is not fit to let it run. Capability determines whether it can be done, reliability determines how much you dare not look at it.

The specific assessment methods can be rereaded.《Demystifying evals for AI agents》pass^k Design this line with a verifier, which can be consulted"How Reward and Training Closed in Real Age"

It's written at the end.

And here, the three questions I'm asking myself are not really standard answers. If you have to put it in one sentence, it's probably: to build an AI Agent, to decide what to give to the model, what to leave in the system, and to believe that it's right.

These earlier judgments will be outdated. Every upgrade to the model will eliminate a shipment of Harnesses, and some of the conclusions that have been set up today will make it ridiculous. Everyone has their own Belief, and everyone's belief is not static, and maybe I'll have a new perspective in an hour, and the reader should have his own.

There is one more thing to admit. Multiple twilights are inevitable, not because Agent is stupid, but because the first time someone says the need is wrong. Do Agent is like this, and write this article.

It's been a half month, and this is the end of the story, and the next one is about to talk about some of the recent experiences and make a little summary.

A little comment on the idea of a DGG

And as this blog evolved, it actually happened in the AI circle that was worth talking about. ** A mathematical scientist found a reverse example of the DGG/Goemans cost assumption by repeatedly prompt GPT-5.6 Pro, and thus perjured it. ** The total number of tips is only about 60 words, and AI Auxiliary Mathology is no longer 95% human and has made AI additional, it is a human being who proposes initial goals and simple directions, leaving all of them to AI to explore.

We talked about it in the front."The progress of the underlying model will certainly eliminate some of the space where process constraints and knowledge are present ... It may be useful to see if there is a sufficiently good feedback signal for the system." DGG's guess is that the search for a counter-examples is a similar generation of variable searches. The code validates whether the LLM proposed counter-case is correct, and the existence of feedback turns one-off generation into an inspirational search. And the way the DGGs assume that the way they prove it can mean that we've been overestimating knowledge and underestimating feedback.

Last year, if we think about building an AI mathematician, we might do a few different-made Agents, analysis, criticism, optimization, proof, etc. Then a suitable communication system is designed. But now the basic model has become stronger."Strong model + long enough search budget + high quality feedback" It may be enough, complex Agent's value itself is questionable.

Feedback is different in different areas, if you need a strategic consultation, Agent, and the feedback is very difficult to give, advertising is much more realistic, further codes have more cheap and adequate feedback signals, and mathematics has the cleanest feedback. AI was the first to emerge in areas that are probably not those where humans feel the greatest need for creativity, or those where there is a strong, extremely cheap, extremely automated verifier. Now code, future chip design, formalization algorithm validation.

What am I doing when I'm building an AI Agent? What we did with AI Agent and Harness is whether it brings irreplaceable power to models or fills a temporary capacity gap.

  • Title: What Problems Are We Really Solving When We Build AI Agents?
  • Author: Hyacehila
  • Created at : 2026-07-28 12:00:00
  • Link: https://hyacehila.github.io//blog/2026/07/28/what-problems-are-we-really-solving-when-building-ai-agents/
  • License: This work is licensed under CC BY-NC-SA 4.0.
Comments