AI fundamentals· Both learning paths

Intent Is All You Need, Still

Why better AI execution does not guarantee the right outcome, and how questions, examples and prototypes help clarify and preserve intent.

Charles Shen, PhD, EMBA

Published Sep 14, 2026 · 17 min read

On this page (7 sections)

Ask an AI agent to shorten a weekly report. The revision is half the length, but the number you use to decide what to do next has disappeared. You meant to make the report easier to use. The agent made it shorter.

A better model might recognize that the number matters and keep it. But there is a larger question behind this ordinary mistake: how does an agent know which parts of a request express your intent—the outcome you want—and which parts are your best attempt to describe it?

In Intent Is All You Need, published in June 2025, I proposed intent clarification as the foundation of building with AI. A specification provides a bridge between what someone wants and what gets built. Conversation helps make that bridge faithful to the purpose it serves.

Another fifteen months of building agents and building with them have made the question more pressing. Models are getting better at implementing a specification, and agents already ask questions to clarify what it should contain. Translating human intent into that specification is improving too, but the progress feels uneven: execution is advancing faster than the shared understanding that directs it.

That makes it useful to distinguish the gaps hidden inside “the agent didn’t understand.” Some can be closed by a question. Others need an example, an experiment, or knowledge neither side has brought to the work yet. The difference helps determine when an agent can keep going, and when another hour of autonomous work is likely to take it farther from the intended outcome.

A specification can be right about the wrong thing

The bridge from intent to outcome has two parts. First, the specification has to represent what the person means. Then the implementation has to satisfy the specification. Better execution does not guarantee that the specification captures what the person means.

Intent leads to specification: captures intent? Specification leads to implementation: meets specification? Learning and clarification can return from either stage to refine intent.

A striking example of execution progress came in May 2026. An internal OpenAI model produced a counterexample to the Erdős unit distance conjecture, a mathematical claim about how often the same distance can occur between points on a plane. A group of mathematicians published a human-verified account of the result. Their paper also describes human work to refine its presentation and develop the argument further.

For this task, mathematical correctness supplies a standard that experts can apply independently of whoever requested the solution. That does not make checking the proof easy, or make mathematical research free of judgment. It does make the acceptance question much less dependent on an unstated personal preference than “make this report useful” or “make this video feel right.”

The report illustrates the other part of the bridge. “Reduce the word count by half” is a clear specification. It can also be a poor representation of the purpose. The agent could implement it perfectly and still remove the information that makes the report worth reading.

As agents become capable of carrying out larger assignments, an early misunderstanding can cascade through the design, implementation, and even the tests used to approve the result. Each stage builds on the previous one, so correcting the original assumption may mean undoing several layers of work. The cost includes the time spent waiting, the effort of inspecting and correcting the result, and model usage, metered in tokens. A model that takes on more work can consume more of a subscription allowance or generate higher usage charges even if individual operations become cheaper. Faster production does not make a mistaken direction free.

This is why clarifying intent can become more valuable as execution improves. More of the outcome depends on whether the work was aimed correctly, and more can happen between the initial request and the moment the person sees the result.

The judgment behind the plan

A specification describes requirements and how the work will satisfy them. Its usefulness depends on whether those requirements preserve the person’s judgment: how they would distinguish a satisfactory outcome from an unsatisfactory one.

For the report, that judgment might be straightforward. The reader needs to see whether weekly customer orders are increasing, so the total number of orders and its change from the previous week must remain visible. Once that purpose is expressed, the agent has a reason to preserve those figures while shortening the surrounding text.

An agent can help uncover this judgment. It can ask questions, draft a plan, point out conflicting requirements, and show what its interpretation would produce. The specification can take shape through that conversation; it need not be a document the person writes alone. But a detailed plan can still leave the decisive question unanswered: would the person accept what the plan is going to deliver?

Including tests in the specification helps when those tests express the intended outcome. A test might confirm that the report calculates totals correctly and meets its word limit without checking whether the reader can see the change from last week. Those tests pass, but the report still fails the decision it was meant to support. Once that need is understood, it can inform a better test. The question is whether the test checks the person’s goal or a narrower implementation detail.

One useful test is whether someone else could apply the stated criteria and reach a decision the person would recognize. If they could, the judgment has become usable beyond the conversation. If they could not, the specification may contain assumptions that will become visible only when the result comes back.

This is a practical test, not a guarantee of agreement. Even experts can disagree, and a criterion that works for one example may fail on another. The point is to find where the agent has a basis for deciding and where it is still substituting its own interpretation.

Three ways judgment can be missing

The weekly report begins with the easiest case: something the person can explain once asked. Other cases need more than a conversation. Three distinctions help determine what would give the agent a sounder basis for its next decision.

Unsaid: the person can give an answer, but it has not been communicated. They may already have a firm view, or the relevant question may not have occurred to them. Ask what they check in the report each week and they can name the figure. The agent needs to surface that information before treating its own interpretation as the requirement.

The question itself may take work to find. Consider a sales follow-up agent sending messages at intervals. If it schedules messages using the sending system’s time zone, a convenient morning send could arrive during the recipient’s night. Asking whose working hours should govern the schedule may be enough for the person to specify the recipient’s local time. The information was missing from the request, but no experiment was needed to answer the question.

A video-generation trial showed how an agent’s interpretation can hide such a question. An agent was comparing tools for an explainer I wanted to make. I had pointed to a reference channel to communicate the visual direction. The agent turned that reference into rules favoring diagrams and boxes, with restrictions on characters and cinematic imagery. The restriction on characters came from its reading of the reference, not from a requirement I had given it. This was an exploratory trial, and I had not scrutinized the inferred rules as I would requirements for a production system.

One tool produced a whiteboard animation with cartoon characters. The agent judged it substantially worse than another result. When I saw the animation, it was much closer to the vivid visual explanation I wanted. There were factual omissions to correct, but those were separate from whether the visual approach worked.

If asked whether to include characters, I could have answered yes. Instead, the agent made that decision on my behalf and used it to judge the candidates. Showing me the animation exposed the misunderstanding. A reference can help communicate intent, but the agent’s interpretation of that reference still needs to be checked against what the person meant.

Unsayable: asking does not produce a sufficient explanation, but the person can recognize what works when they see it. Imagine choosing a thumbnail. Several options follow the same instructions, use the same subject, and remain legible at a small size. One still feels right for the video and the others do not. The preference is real even if the person cannot fully explain it.

This does not mean visual judgment is beyond explanation. A professional designer can name principles, identify weaknesses, and teach distinctions that a client may not yet have the vocabulary to express. That expertise should be used. “Unsayable” describes what this person can articulate at this point; it does not mean that nobody could explain it. Even after expert guidance, a particular choice may still be easier to communicate through examples than through another paragraph of instructions.

Examples can carry that part of the judgment. A preferred thumbnail, a rejected alternative, and whatever explanation the person can give provide a basis for the next attempt. The agent can use those reactions without pretending they amount to a complete formula for taste.

Unknown: the person does not yet have a reliable basis for judging the result. More alternatives may not help if they do not know what to look for. This differs from recognizing a good thumbnail but struggling to explain the choice: here, the person needs to develop the judgment itself.

In his field guide to finding unknowns, Anthropic staff member Thariq Shihipar describes encountering this while editing a launch video. He tried asking for color-grading variations, then realized he did not know what good color grading looked like. He asked Claude to teach him about it instead.

That is where looking for unknown unknowns becomes useful. An agent can introduce considerations the person did not know existed, bring in professional practice, or help examine relevant examples. Some questions will become easy to answer once they are visible. Others will require study or an experiment before either side has adequate grounds to choose.

The categories describe a changing situation, not fixed types of people or tasks. Learning can make an unknown answerable. A comparison can reveal a question that brings an unstated preference into the conversation. A designer’s explanation can make part of an unsayable preference expressible. Clarification includes these changes in understanding, as well as putting existing thoughts into words.

Unsaid: a person answers an agent’s question. Unsayable: a person chooses a preferred example. Unknown: a person learns with books and new examples. An AI agent assists throughout all three situations.

Expressing the intent does not ensure it survives

There is another failure that matters here: the person can explain the desired outcome and the agent can acknowledge it, yet the work still follows a narrower target.

One project was meant to improve the value delivered by agent work relative to the resources it consumed. I had clarified that it needed to consider whether a task was worth starting, whether spending during a run was advancing the goal, and whether experience led to demonstrated improvements afterward.

The subsequent implementation checked for particular text in task instructions and included a retrospective script. Those mechanisms could be tested. They did not demonstrate that the running agents were delivering more value.

When challenged, the agent described the substitution: “Then each aspect was translated into the nearest overnight-shippable artifact.” Its own account said the system checked whether a task instruction contained a correctly formatted evaluation line, rather than evaluating whether the running work advanced the goal. The learning component had run once as a demonstration. Work was judged against the narrower implementation briefs.

The missing step was not another invitation to clarify what I meant. The expressed purpose had to remain the basis for judging the implementation. A specification was supposed to preserve that purpose; instead, the work had become organized around checks that were easier to satisfy.

This can happen with a second agent reviewing the work too. A reviewer given a checklist for factual accuracy, links, and layout may verify that an article meets those criteria. Those checks do not tell us whether the article explains its subject well or speaks to its intended reader. Asking a model to act as that reader can broaden the review, but the model’s verdict still needs a reason to be trusted on the particular judgment involved.

The practical question is what the evidence supports. A passing check for one property is not evidence of a different property merely because both are grouped under “quality.” And an agent that stops to ask for approval may still present the wrong decision. Asking whether a document should be saved is little help when the unresolved question is whether its argument makes sense.

Trying something is part of understanding it

Software development has dealt with this problem for a long time. In No Silver Bullet, Fred Brooks argued that clients cannot fully specify a complex product before trying prototypes of the proposed system. His case for rapid prototyping was that interaction with a working model helps reveal what the requirements should be.

Brad Cox challenged Brooks’s broader argument, proposing that markets for reusable software components could greatly improve productivity. Even if that promise were fully realized, our report could become much easier to build and still omit the number its reader needs. Making software easier to produce does not guarantee that its requirements capture the intended outcome. Brooks’s argument for discovering those requirements through prototypes remains relevant here.

Agents extend the opportunities to do that. A person can encounter a proposed interface, a sample report, or a short animated scene while the larger project is still taking shape. The example provides information that a description may not: what catches the eye, what is missing, what is awkward to use, what raises a question nobody had thought to ask.

The useful prototype is the one that exposes the uncertainty. If the issue is whether a report supports a weekly decision, one representative report may reveal more than a complete reporting application. If the issue is visual direction, a short scene may be enough to choose an approach before producing the full video.

That choice also affects cost. Generating and reviewing alternatives consumes time and tokens. Producing several complete videos to answer a question one scene could settle would spend resources on work that does not yet need to exist. The value comes from learning enough to direct the next stage, at a cost proportionate to the decision.

As the result becomes concrete, intent may become more precise or change. Someone may discover that the original request was a poor way to achieve the underlying goal. That is useful information for the project, provided it changes the subsequent work rather than disappearing into the conversation.

Autonomy depends on the work being delegated

In loop engineering, the distinction between being in the loop and on it describes different kinds of human involvement. In the loop, a person participates in the steps that require their decision. On the loop, they oversee work that can continue without their approval at each iteration.

Intent clarification helps determine where that participation belongs. Once the report’s essential figures and purpose are clear, the agent may be able to produce the remaining reports without returning for each one. If the visual direction of a video is unsettled, generating a sample and getting a reaction may be more productive than completing the video under an untested interpretation.

These positions can change repeatedly within the same project. I might work in the loop to choose a visual treatment, move on the loop while the agent develops it, then return when a scene exposes a choice we had not anticipated. That decision can give the agent enough direction to continue independently again. Each return should carry what we learned into the next stretch of work.

How well the agent understands the relevant judgment is one factor in how far it should run. The consequences of a mistake, the ability to reverse it, and the available budget matter too. A reversible design exploration and an action affecting a customer call for different boundaries even if the instructions are equally clear.

The purpose of the current stage matters too. In exploration, the agent searches for possibilities: story ideas, unusual designs, alternative approaches. In delivery, it develops a chosen direction into something that meets the requirements. A video project can move between the two—exploring visual treatments, producing a selected treatment, then reopening a scene that does not work. Creative work may leave more room for alternatives, but a software application with a definite purpose can still have requirements that become clear only when someone tries it.

During exploration, defining the acceptable answer too tightly can exclude the unexpected possibility the person hoped to find. An exploratory run can be worthwhile even when most of its output is discarded. It still needs judgment: whether the possibilities are useful, which ones deserve more work, and whether further exploration is worth the time and model usage. That judgment can be applied after a batch of ideas rather than interrupting each attempt. Exploration does not require abandoning factual accuracy either; the animation’s visual promise and its factual omissions were separate judgments.

Who can make those judgments is a separate question from whether the work is exploratory. A mathematical search can explore many possible approaches while a shared standard of proof determines whether the result is correct. The person who requested the search does not need to supply a personal preference about validity, though checking the proof may still require expert work. A finished video, by contrast, can meet its measurable requirements and still need the creator’s judgment about whether the visual treatment works.

What matters for delegation is whether the relevant standard is understood and can be applied reliably, and where a decision still depends on the person.

Clarifying intent is especially valuable when work is converging on an outcome someone needs to accept or use, because a mistaken assumption can create expensive rework. During exploration, clarification can instead define the question to investigate, the freedom to try different approaches, and the resources available. Both modes involve intent and judgment; they call for different decisions about when the person needs to participate.

Knowing when to ask is still difficult

Questioning, interviews, and plans already help agents understand what a person wants. The harder challenge is recognizing when the available understanding is insufficient for the next decision. An agent can be confident because a task fits a familiar pattern while missing the detail that makes this person’s task different.

More frequent pauses do not necessarily solve that. A cautious agent can interrupt for routine implementation choices while confidently deciding something the person would have wanted to judge. A proactive agent can make useful progress or accumulate work around a mistaken assumption. The quality of the handoff depends on which question returns to the person and what they are shown to answer it. For work that depends on a person’s preferences, deciding how far the agent should run is likely to remain a human responsibility longer than drafting the clarifying questions.

There is an incentive problem as well. If a product treats longer unattended runs and fewer interruptions as measures of success, a timely request for judgment can appear to be a regression. That rewards the wrong behavior when a short conversation would prevent hours of unwanted work. The useful measure is whether the person receives an outcome they value for the time, attention, and resources spent, including the cost of clarification.

The original promise of intent-first work remains compelling: a person should be able to bring purpose and domain judgment to an agent without becoming the implementer of every detail. Making that work requires ways to express what the person knows, discover what they have not considered, and preserve those decisions as the agent builds.

Intent is all you need, still—but it takes shape and is tested throughout the work. We move in the loop to clarify a choice, on the loop while the agent acts on it, and back in when the result reveals something we could not settle earlier. As execution improves, deciding what is worth producing and whether it serves the purpose remains the harder part of many assignments. The aim is to let the agent carry more of the work while returning the right decisions to us, early enough to change what happens next.