A sequel to AI Agent Engineering: Building Agentic Systems for Enduring Value
In December 2024, Anthropic published Building Effective Agents, and I wrote down what I took from it. The lessons were plain. Keep things simple. Treat workflows and AI agents as a spectrum, with predefined steps at one end and a model directing its own work at the other. Move toward the AI agent end gradually, so that each move is “justified by a demonstrated need.” A few months later, AI Agent Engineering carried the same advice into its best practices: start with the simplest architecture that solves the problem.
Those principles still stand. The landscape around them has moved. Anthropic’s article now opens with a note that much of the tooling it described has changed since December 2024. Two years ago, building an AI agent was a project: choose a framework, wire up tools, write the loop, and only then learn whether the idea worked. Today a capable AI agent is something you can easily download and install, or simply open. Off the shelf it reads files, runs commands, browses and searches, and once connected to the tools you already use, it does a great deal more before you have built anything.
More and more people now start directly with an AI agent, and it is a sensible place to start. If you want an AI agent to deliver a particular result, for yourself, your team or your customers, you are building, whether or not you would call yourself a builder. That is the argument of Intent Is All You Need: Why Everyone Is Already a Programmer. Here is what the shift means, in one sentence: when everyone starts with an AI agent, what sets yours apart is how you engineer it to be your own.
Trying an idea with an AI agent is now easy
Software has long had an answer for not knowing your requirements: build something small and try it. Fred Brooks argued decades ago that clients cannot fully specify a complex product before working with a prototype, a point I returned to in Intent Is All You Need, Still, where trying something is part of understanding it. App building absorbed that lesson when tools appeared that turn a description into a working app in minutes. The prototype became a commodity.
The same shift has reached AI agents. When you have a task or a workflow you would like handled, you can hand it to an existing AI agent this afternoon and see how far it gets. Your first trial can happen in the AI agent you already use every day. That trial is the prototype. It does not suit everything: a fixed, well-understood process is still better served by fixed steps. But the easiest thing to try is now often an existing AI agent, so exploration can begin at the AI agent end of the spectrum. Whether the trial achieves what you want still governs what you keep and what you build.
Every starting point comes with engineering inside
A trial like that shows what is possible. It also shows what does not fit, and to see why, it helps to be precise about what an AI agent is.
An AI agent is more than its model. In the ABC framework, it has three parts: Action is what it can do, Brain is the model that reasons, and Context is the information the model is given. Something has to run those three together. That is the harness: the machinery that assembles what the model sees, runs the tools the model asks for, feeds the results back, and operates the whole thing in a loop within boundaries set by humans.
As of October 2026 there are many AI agents to start from, and each arrives with its harness already built. OpenClaw, which began in late 2025, and Hermes Agent showed what an always-on personal AI agent could do. The large AI companies followed with their own: Claude Cowork, ChatGPT Work and dots, Grok Bot, and Meta’s Muse. Meta’s chief executive, Mark Zuckerberg, called OpenClaw “a very exciting glimpse of what types of things should be possible,” and OpenAI hired its creator “to drive the next generation of personal agents.” For builders there are coding agents such as Claude Code and Codex, and harnesses meant to be built on: Pi, which calls itself “a minimal agent harness,” and Prime Agent, a harness designed to improve itself. Both leading labs now also run a harness for you. Anthropic describes Managed Agents as a “pre-built, configurable agent harness,” and OpenAI’s Agents API offers “access to the Codex harness.”
These starting points differ in how much has been decided for you. Pi intentionally decides little and leaves room to build. The products from the large AI companies carry a great deal of opinionated engineering, formed for the purpose each was built for, and it serves that purpose well. Any of them can act as a capable general tutor or support desk. The difference appears when the AI agent has to be yours: your material, your policies, your customers, your way of working. Whether to build from scratch, which AI agent to build on, and how far to customize it are each becoming critical decisions. One observer at a recent Y Combinator Demo Day put the pattern this way, in a post the YC president shared: “Everyone is just basically building a domain-specific harness.”
When I wrote about AI agent engineering in April 2025, I described this layer as orchestration: how tasks are decomposed, how AI agents collaborate and how errors are contained. The vocabulary has settled since, to the point where the labs sell a harness under that name. Shaping one for a domain is now where much of AI agent engineering happens. It is not the whole of the discipline, since models, knowledge and tools all still need choosing and building. The harness is where they come together.
What is available, and what is required
Suppose many of us start from the same AI agent. How does one become different from another?
Go back to Action and Context in the ABC framework. For each, there is what is available to the AI agent, which the model may use at its discretion, and what is required, which the harness does whether or not the model asks.
| Available, at the model’s discretion | Required, by the harness | |
|---|---|---|
| Context | what it can draw on | what it is always given, and what is kept from it |
| Action | what it can do | what always happens, and what is refused |
Most of what people add to an AI agent goes in the left column: instructions, documents, memory, connections to other tools. Agent Skills belong there too. A skill is a folder of instructions, often with reference material and sometimes with scripts, that the AI agent picks up when the model judges a task calls for it. Many people already reach for a skill to control how an AI agent behaves, and a skill does guide it. But the model can follow a skill closely, follow it loosely, or leave it unread, and a script inside a skill runs only when something calls it. A skill guides. The harness enforces. As I put it when building a support desk on Grok Bot, a sentence in a skill does not enforce the same boundary as code that refuses the action.
The right column is code the harness runs on its own. Take a tutoring AI agent: it teaches from your material, asks questions, remembers how you answered and brings each topic back for review. When several learners use it, one learner’s history must stay out of another’s lesson. The code that ensures this is nothing exotic. It is a database lookup for the signed-in learner. What makes it part of the harness is where it sits: it runs at the start of every session, before the model sees anything, and the model cannot skip it.
Who decides what is required? Whoever designs the harness. When the designer is a human, the chain stops there. When the designer is an AI agent, whether a coding agent writing the harness or a harness revising itself, a human has set what that agent may decide. What is mandatory for the model is a choice for the designer.
How much to leave to the model is a design choice with no single right answer. A model that searches a learner’s history and picks the next topic is how a lesson becomes personal, and the more capable the model, the more of that you can leave to it. What stays required is what the product has decided whatever the model can do, such as whose records may be read. I drew these same distinctions for Claude Code in the Agent Capability Lifecycle, as active versus reactive and deterministic versus not. They hold for any AI agent.
One path, different depths
Those starting points sit along one path, and they differ in how far along it they let you go.
Nearly all of them let you add to what is available. In ChatGPT, Claude, Grok Bot or Muse, that is most of what you can do: skills and connectors, usually packaged as a plugin. It is also how you put something in front of the people who already use those AI agents. Fewer let you add what is required. Claude Code and Codex have hooks, and Pi has extensions, which are code that runs inside the harness and can block an action before it happens. A managed harness lets you configure more while the lab runs it. Deepest is building and running a harness you control. A plugin for a consumer AI agent and an extension for Pi are the same kind of work at different depths.
Depth buys control, and it has a price. On a consumer platform the company runs the AI agent and brings the users, and you control little. OpenAI, for example, says developers “can directly launch new native experiences to our collective 1.2B weekly users.” On a managed harness, such as Anthropic’s Managed Agents, the lab runs it, and reaching customers is your job. When you run an open harness yourself, you control its implementation and take responsibility for hosting and distribution. No depth is right in general. It is chosen for the outcome and for the customers who have to be reached, which makes it as much a business decision as an engineering one.
Engineering toward the outcome
You start with the requirements you know, and trials surface the ones you did not. That does not conflict with defining success first. Before each round of building, you state what success means as you now understand it, and you revise that statement openly when a trial teaches you something. What success must not be is written afterward, by whatever the build happens to do.
I find it useful to think about building an AI agent in six stages: Outcome, Requirements, Design, Foundation, Build, Evaluate. For a tutoring AI agent, the Outcome is that the learner understands the material and remembers it. Requirements turn that into behavior someone can observe. Design settles what is available and what is required. Foundation is the choice of starting point and depth. The stages feed each other, and a clear specification can still be right about the wrong thing, which is why the requirements come from the outcome and are checked against it.
Three loops organize what the resulting AI agent does. In the first, the model chooses an action, a tool carries it out, and the result returns to the model. The second belongs to the application, meaning the AI agent as a person uses it: it takes input, runs the model as needed and decides when the work is finished. The harness runs both. The third loop is improvement, where experience and evaluation change how the AI agent operates next time. That loop is shared. People conduct it, a coding agent can conduct parts of it, and a harness built to improve itself can conduct parts of it too, within what a human allows it to change. This is where designing the loop meets what AI Agent Engineering said about lasting value: models keep changing, and the evaluation that tells you whether a change helped is what keeps its worth.
The same holds when an AI agent writes the code, which is where AI agent engineering meets agentic engineering, the practice of building software with AI agents. One small example from my own work: two implementations of the same tutoring AI agent, both written with AI from one description, passed the same tests and then behaved differently when a saved review date was malformed. One stopped with an error. The other quietly dropped the topic from practice. The requirements had not said which was wanted. Whoever writes the code, the outcome has to be specified and the result checked against it.
A new engine, and still software
The idea that AI agents will replace apps keeps gaining ground. Bill Gates wrote in 2023 that with an agent “you won’t have to use different apps for different tasks.” I think the idea is broadly right. To be more precise, however, an AI agent is itself a software application.
Picture a flashcard app: cards on a phone, a swipe, a button. Picture a tutoring AI agent, and most people see a chat window. Those look like two different things, and what differs is the interface. The interface is not the essence. The same flashcard screens could sit on top of an AI agent, with a conversation beside them. What changed is underneath: the engine that drives the application. It is the same car with a next-generation engine. This is what I meant by AI-native in 2024: AI at the core as a foundation, where AI-augmented adds it as an enhancement.
A new engine does change what the interface can be. Interfaces sit on their own spectrum, from open conversation at one end to fixed screens at the other, and the same principle applies as with workflows and AI agents: choose the point that serves the person. Sometimes a proven interface stays. Sometimes the engine makes possible one that could not exist before. Chat becomes part of the experience, and so does voice. Around the engine, the durable software still has to be built: a saved review date needs a job that runs on that date, a reminder that reaches the learner, and records that survive a crash.
What sets yours apart

The landscape changed, and the principles held. Starting is now easy, and a first trial shows what is needed. Simplicity still decides what you build from there: add only what the outcome needs. But everyone who starts from the same AI agent starts with the same defaults: choices about how the AI agent works, made by someone else for a general purpose. Turning those defaults into yours takes engineering: defining the outcome you want, deciding what your AI agent has available and what is required of it, choosing how much of it to own, and checking the result. Models will be replaced, and the AI agent you started from may be too. That engineering carries over, which is why AI agent engineering has moved to the center.
When everyone starts with an AI agent, what sets yours apart is how you engineer it to be your own.
