AI engineering· Agentic engineering path

Build Your Own AI Agent Harness with Pi and Prime — Hands-on Course

Build a tutor agent with Pi and Prime Agent. Customize its harness through extensions and an SDK, test its behavior, and explore refinement and scheduled review.

Charles Shen, PhD, EMBA

Published Sep 29, 2026 · 32 min read · intermediate

On this page (34 sections)

An AI coding agent can already explain a subject, ask questions and write code. What does it take to turn those capabilities into a tutor that remembers practice, selects the next review and follows a consistent review policy?

We’ll build Academy Tutor on Pi, inspect how its harness works and verify its review behavior. That gives us a working tutor we can use and extend. Two continuations then explore owning the application through Pi’s SDK, and using Prime Agent to adapt future instructions and start a review while the chat interface is closed.

We work with AI throughout design, implementation and testing. Both the supplied reference and alternative implementations were developed with AI; we compare their behavior and design choices. Reading their code lets us understand what the model decides, what the program enforces and what our tests establish.

This course continues How AI Agents Actually Work and AI Agent Engineering: Building Agentic Systems for Enduring Value. You don’t need to read them first. We’ll recap the concepts we need, then apply them to a working project.

Before you begin

You should be comfortable running terminal commands and reading some JavaScript. You’ll need Node.js 24, npm, a terminal, an editor and a supported model connection. Node’s download page provides the runtime. Prime’s later exercises also need uv; installation is covered when we reach them. Any editor works. My Mac setup walkthrough covers the optional Nix and LazyVim environment.

Get the course resources before starting the build: they contain the cards, complete implementations and checks used below. The written steps stand on their own; the companion video shows the live build and troubleshooting.

Choose how far you want to go:

Milestone Sections What you finish with
Build and check the Pi tutor 1–5 A working extension, saved practice and tested review rules
Own the application 6 A standalone SDK tutor and a verified due-review cycle
Explore adaptation and background review 7–9 A Prime tutor, inspected refinement changes and a review that starts while detached

The first milestone is a complete project. Continue to the SDK, Prime, or both when those capabilities are useful to you. Prime starts from the Pi extension and does not require the SDK application.

1. Understand the harness and the design playbook

An agent is the whole system pursuing a goal. The model is one component of that system. In the Action–Brain–Context framework, Action is what the agent can do through tools; Brain is the model that reasons and chooses its next step; Context is the information the model receives, including instructions, the task, documents, messages and selected memories.

The harness operates those components. It prepares context, calls the model, executes permitted tool requests and returns their results. It runs this loop within boundaries set by humans.

For example, a model can request that a file be read. The harness executes the read operation and supplies the result to the model. The next model response can use information that was absent from its earlier request. The model chooses the next action; the harness provides the mechanics and controls that make the action possible.

The harness prepares context, calls the model, executes permitted tool requests and returns their results.

Harness design is therefore part of agent design. We need to understand the experience we want before deciding what to change in the harness.

Three loops to keep in view

There are three useful levels at which to examine an agent system:

Three loops: model and tool interaction, application coordination, and system improvement.

Loop What repeats Tutor example
Model/tool loop Model response, requested tool execution, updated context Assess an answer, save the assessment, continue teaching
Application loop Prepare work, interact, continue or finish Select eligible concepts, run practice, end the session
Engineering improvement loop Evaluate behavior, revise the design, try again Find a failure in review selection and change the implementation

The harness participates in the first two. It is not synonymous with the application loop: it also executes the mechanics of the model/tool loop. Tools are reusable capabilities; their place in a particular sequence depends on who decides to invoke them.

These levels can interact. A self-adapting application can run a review process that changes future instructions. That is an improvement mechanism inside the application. We still need to evaluate whether the change helps the application achieve its goal.

Guidance and enforcement

A skill can package instructions, examples and supporting scripts for the model to use. That can be an excellent way to teach an existing agent a procedure. But an instruction to use a script and a program that always invokes the script are different arrangements.

Consider three controls across ABC:

  • Action: asking the model to request approval guides its behavior. Checking approval before executing the action enforces the requirement.
  • Brain: asking the model to stop after a budget guides it. Refusing another model request after the budget is exhausted enforces the limit.
  • Context: asking the model to retrieve relevant material helps it select information. Preparing or filtering the input before a model request controls what that request receives.

A model can certainly curate context using search, file-reading or Python tools. The design question is whether that choice should remain with the model or whether a particular operation must happen regardless of its choice. Skills and harness controls can work together.

Across Action, Brain and Context, instructions guide the model while harness controls check approval and enforce request and context limits.

A playbook for designing the whole agent

Start with the desired outcome. Translate it into requirements, including how you will judge success. Design the model, context, tools and harness together. Choose a foundation: build the underlying loop yourself or reuse an existing agent. Then build and evaluate.

Outcome, requirements, design, foundation, build and evaluate, with continuous improvement revisiting decisions at any stage.

Evaluation comes after a working implementation, but its criteria begin with the requirements. And the sequence is iterative: a new requirement may change the design and our choice of foundation. We’ll do exactly that when we reach Prime.

2. Design Academy Tutor

Our outcome is simple: help a learner understand and remember course concepts.

We’ll use eight cards drawn from Agenteer Academy. Each card has an ID, a question, a reference answer and source links. Changing this material gives the tutor another subject to teach; the harness itself does not depend on these particular concepts.

The project requirements define the behavior we want:

  • Explain a new concept, then ask a question.
  • During review, ask before giving help so we can test recall.
  • Save the concept, assessment, whether the learner received help, and evidence from the answer. Preserve earlier attempts.
  • Review a correct, unassisted answer in three days; review other assessed answers in one day.
  • Select due concepts before unseen concepts, using the latest attempt for each concept. Exclude future reviews.
  • If nothing is eligible, start no lesson and make no model call.

The explicit one-day/three-day rule makes the review policy easy to inspect and test.

Course cards and saved attempts feed eligible practice. The model teaches and assesses; record_attempt validates and appends assessments.

There are two decisions that are easy to conflate. The model decides whether an answer is correct, partial or incorrect and when to submit that assessment. The record_attempt tool validates the submission, calculates the review date and writes the record. Deterministic recording does not make the model’s assessment deterministic.

Responsibility Our design decision
Select the practice material The command or application prepares the queue before teaching starts
Explain, ask and assess The model decides how to conduct the conversation
Submit an assessment The model requests record_attempt
Validate, schedule and save The tool implementation applies the rules
Execute model and tool requests Pi supplies the underlying loop
Start, accept input and exit Pi’s interface for the extension; our application for the SDK version

The tool by itself is not the reason to customize a harness. A skill could guide a model to call the same tool. Here we also choose to prepare eligible practice before the conversation and give the application an explicit start and finish. Those are decisions about how the agent runs.

3. Set up and explore Pi

Pi is an agent foundation with model connections, tools, session management, a terminal interface and extension APIs. We can reuse those capabilities while adding the tutor’s behavior.

Building directly on model APIs is another useful route when we want to understand or control the lower-level implementation. An existing SDK is also a foundation we can reuse: the Market Mind SDK chapter uses OpenAI’s Agents SDK. Here, we’ll begin by seeing what Pi already provides.

Create the project

The course resources contain the starting cards and tracer, complete reference implementations, requirements and comparison files. The additional implementations preserve the alternative extension, SDK application and Prime adaptation developed during the course, with instructions for running them. Place the resource folder at ~/projects/academy-tutor-course. Its initial layout is:

academy-tutor-course/
  README.md
  requirements.md
  starter/
    cards.json
    tracer.ts
  reference/
    tutor.js
    tutor.mjs
  comparison/
    generation-prompt.md
    behavior-checks.mjs
  examples/
    examples.md
    pi-generated/tutor.js
    sdk-generated/tutor.mjs
    prime-tutor/index.ts

Run these commands at your shell, one line at a time:

cd ~/projects/academy-tutor-course
pwd
ls
node --version
npm --version
mkdir academy-tutor
cp starter/cards.json starter/tracer.ts academy-tutor/
cd academy-tutor
npm init -y
npm install --save-exact @earendil-works/pi-coding-agent@0.87.1
npx pi --version
npx pi

Use a new course folder if you already have work under academy-tutor; don’t copy over an earlier project. If Node or npm is missing, install it before continuing.

npm init -y creates package.json using default answers. The install command adds the exact Pi version to that project and creates package-lock.json, which records the resolved dependency tree. npx pi uses the installed project executable. We install Pi locally because the SDK application will import the same package later. Pi’s homepage also offers a global installation for general use.

Expect version 0.87.1. Inside Pi, use /login to connect your provider and /model to select a model. The demonstration uses GPT-5.5 through an OpenAI Codex connection. A local package installation does not clear existing user authentication, settings or extensions in ~/.pi/agent.

You can check the project executable from Pi’s input with !!npx pi --version. Pi’s double exclamation form keeps the command output out of model context; a single ! includes it. See the Pi’s terminal-command documentation.

Look at the material before asking Pi to teach it

Open cards.json in your editor. The first card asks for the three parts of ABC and the role of each. Its answer and sources sit beside the question, so we can understand what we are asking the tutor to use.

Now ask Pi:

Read the first card in cards.json and ask me its question without showing the answer.

Answer it yourself. We already have a conversational tutor because the underlying agent can read material and explain it. What we have not yet installed is our specific record format or review policy. If ordinary Pi chooses to save progress in a file such as learner.json, that is its chosen implementation. The extension we build next will use attempts.jsonl under explicit rules.

Inspect what happened with a tracer

Before adding tutor behavior, let’s observe Pi’s existing loop. Exit Pi with /quit, then launch it from the same academy-tutor folder with the supplied tracer:

npx pi -e tracer.ts

The -e option loads an extension. Our tracer subscribes to events and saves observations. It leaves Pi’s prompt, tools and provider payload unchanged. This is one use of the extension API: an extension can observe the agent before it changes anything.

Send the same first-card prompt. Then open trace/latest/trace.log beside tracer.ts in your editor. Also open system-prompt-1.txt and the request-*.json files in that directory. Each launch creates its own trace folder; latest points to the most recent one.

The short log gives us the sequence: prompt, model request, tool request, tool result and subsequent response. The request files let us inspect the actual payload at Pi’s provider-request hook.

Here is the tracer’s request hook:

pi.on("before_provider_request", event => {
  const file = `request-${++requestNumber}.json`;
  writeFileSync(join(dir, file), JSON.stringify(event.payload, null, 2));
  note(`REQUEST SNAPSHOT: ${file}`);
});

Follow the first-card exchange through those files. The model receives instructions, the user request and tool definitions. It can request a file read. Pi executes that request and includes the result in a subsequent model input. The model can then ask the question from the card.

The system prompt alone is not the entire context. Nor does naming a file automatically put its contents into context: inspect the read result and the later request. These snapshots show what Pi prepared at the hook; their number should not be treated as a count of successful network calls.

We now have a concrete view of the foundation we will extend.

4. Build the Pi tutor extension

The complete extension is short enough to read in three parts: the assessment contract, the record_attempt tool and the /academy command. Copy the complete file into your project, then inspect those parts alongside their comments. The excerpts below explain the implementation; they are not separate files to paste and run.

Exit Pi with /quit. At the shell:

cd ~/projects/academy-tutor-course/academy-tutor
cp ../reference/tutor.js .

Define what an assessment contains

The tool accepts four fields:

const attemptSchema = Type.Object(
  {
    concept: Type.String(),
    correctness: StringEnum(["correct", "partial", "incorrect"]),
    assisted: Type.Boolean(),
    evidence: Type.String({ minLength: 1 }),
  },
  { additionalProperties: false },
);

The model supplies a concept ID, its assessment, whether help preceded the answer, and evidence. The schema defines the shape of that request. It does not decide whether the assessment is educationally sound.

Save one attempt and calculate the next review

pi.registerTool makes record_attempt available for model calls. Its implementation first checks that the concept remains in the selected queue. It then applies the interval policy:

const days = attempt.correctness === "correct" && !attempt.assisted ? 3 : 1;
const reviewedAt = new Date();
const dueAt = new Date(reviewedAt.getTime() + days * 24 * 60 * 60 * 1000);
const record = {
  ...attempt,
  reviewedAt: reviewedAt.toISOString(),
  dueAt: dueAt.toISOString(),
};
appendFileSync(join(ctx.cwd, "attempts.jsonl"), JSON.stringify(record) + "\n");
queue = queue.filter(card => card.id !== attempt.concept);

A JSONL file stores one JSON object per line. appendFileSync adds a new line without replacing previous attempts. Removing the concept from the in-memory queue means it has been completed for this practice session. The tool returns both the saved record and the concepts still remaining.

Prepare practice before the model starts

The /academy command reads the cards and saved attempts. A map retains the latest attempt for each concept in memory; it does not overwrite the history on disk. The command separates due reviews from unseen concepts, joins them in that order and excludes future reviews.

It then selects the tutor’s model-facing tools:

pi.setActiveTools(["read", "record_attempt"]);

This selects registered tools for the model. It is not a restriction on what the extension’s JavaScript can do, and it does not turn the surrounding process into a sandbox. Our command still reads files directly with Node’s filesystem API. Other extensions can also change the active set later, which matters when we compare implementations.

When there are eligible concepts, pi.sendUserMessage starts practice with the selected cards and teaching instructions. For a new concept, explain before asking. For review, ask before explaining. After assessment, request the record_attempt tool and continue using its returned list.

There is no JavaScript loop around every question in this extension. Pi runs the model/tool loop, while the model follows the tutoring instructions and decides what to say next. The command owns preparation; the tool owns validation and saving. If the queue is empty, the command returns without sending a model message.

Use the tutor and inspect its record

Launch it:

npx pi -e tutor.js

Enter /academy. Answer a question, then open attempts.jsonl in your editor. Find the concept, assessment, assistance flag, evidence and two timestamps. Our first learning exchange produced a partial assessment and a one-day review interval.

Notice that new learning is assisted practice: the explanation came before the question. Even a correct answer after that explanation gets a one-day review under our policy. A later correct answer without help is the case that earns three days.

The slash command gives this extension an explicit entry point. In the SDK application, launching the program will start practice. Both can use the same preparation logic; the interface determines how the learner initiates it.

5. Evaluate and compare implementations

A tutor that appears to work in one conversation can still select the wrong cards or overwrite history. We need to check the program rules as well as use the tutor ourselves.

The behavior checker was written for this course from the requirements. It loads the real extension with Pi’s loader, creates temporary course data and supplies synthetic assessments. It does not ask a model to grade answers.

For example, to check the one-day policy, the test supplies three assessments: correct with help, partial without help, and incorrect without help. It reads the saved record and checks that dueAt - reviewedAt is one day. This tests the tool’s arithmetic and persistence, not the quality of teaching.

Ask Pi to explain and run the checks

Exit the tutoring session with /quit, then start ordinary Pi in the same folder:

npx pi

Ask:

We’ve built the tutor. The requirements are in ../requirements.md, and the tests are in ../comparison/behavior-checks.mjs. Explain one test so I can see how it checks a requirement, then run them on tutor.js and show me the results.

Starting ordinary Pi gives the coding assistant its normal tools. The tutor session may have selected only its practice tools. If you prefer to run the checker directly, exit Pi and use:

node ../comparison/behavior-checks.mjs tutor.js

The checker writes behavior-results.json. Its seven contract checks cover registration, queue selection, recording and duplicate rejection, review intervals, invalid assessment inputs, and sessions with no eligible work. It also has a separate invalid-date comparison probe. Read the names and results rather than treating a total count as sufficient evidence.

Generate another implementation from the brief

We can also ask Pi to build the tutor from the same requirements. This is useful because AI can produce different implementations of the same design. We want to compare their behavior, not their line count.

Open generation-prompt.md before using it. It is the project brief: what the tutor should do and what files and APIs it should use.

At the shell, create a sibling project:

cd ~/projects/academy-tutor-course
mkdir academy-tutor-generated
cp academy-tutor/cards.json academy-tutor/package.json academy-tutor/package-lock.json academy-tutor-generated/
cp comparison/generation-prompt.md academy-tutor-generated/
cd academy-tutor-generated
npm ci
mkdir docs
cp node_modules/@earendil-works/pi-coding-agent/docs/extensions.md docs/extensions.md
npx pi

npm ci installs the dependency tree recorded in the copied lockfile. We want the same Pi version when comparing the implementations.

Ask Pi:

Build the tutor described in generation-prompt.md. The course is in cards.json, and the installed Pi extension documentation is in docs/extensions.md.

Then open the generated tutor.js beside the conversation and ask:

Walk me through this tutor. How does it choose the next concept, save an answer and decide when I should review it?

A separate folder makes the comparison easier to understand and keeps histories separate. It does not prevent a coding agent from reading neighboring folders. Inspect which files it actually used; don’t describe this arrangement as access isolation.

Exit the coding conversation with /quit. At the shell in academy-tutor-generated, run npx pi -e tutor.js, then enter /academy and answer a question. Ask where it saved your attempt, then open that exact file. A record in another implementation’s folder does not establish that this tutor saved anything.

Exit and start ordinary Pi again. Ask it to run the same checker on this implementation. The direct command is still:

node ../comparison/behavior-checks.mjs tutor.js

The alternative implementation and supplied reference both passed the seven requested contract checks. The malformed-date probe exposed a difference: the reference reported an invalid review date, while the generated version silently left that concept out of the queue. This was a useful edge case to discuss separately from the requested contracts.

There was also an interface difference. The generated extension restored the previous active tools inside record_attempt as soon as the remaining queue became empty. That explained why ordinary coding tools could reappear after tutoring. The code explicitly changed the active set again when practice finished.

These checks give us evidence about selection, persistence and scheduling. Direct interaction gives us evidence about the learning experience. We need both, and a production tutor would need a broader evaluation of assessment and learning quality.

Milestone: a working Pi tutor

You can finish here with a tutor extension you understand: /academy selects practice, an assessed answer produces a saved attempt, and the checks verify the selection and review rules. Inspect your own saved record and test output before moving on. The malformed-date comparison also shows why passing the requested checks leaves room for further investigation.

Continue below if you want your own application interface. To explore Prime instead, go to section 7 with the original academy-tutor project.

6. Build and test an SDK application

Start here with: the academy-tutor project from sections 3–5, its working tutor.js, cards.json, installed Pi package and model connection. The alternative sibling implementation is optional for this continuation.

An extension lives inside Pi’s interface. An SDK application uses Pi’s capabilities inside an interface and lifecycle that we own.

An extension uses Pi’s interface and lifecycle; an SDK application supplies its own. Both reuse Pi’s runtime and add domain rules.

The Pi SDK documentation explains how to create an agent session, provide tools and resources, send prompts and subscribe to events. Our complete reference application shows the tutor’s policy with a terminal input loop.

Read it in three parts: selecting practice and defining the record_attempt tool; creating and observing the Pi session; and accepting learner input until practice ends. We can then ask Pi to make an application from the extension we already understand.

Generate the application beside its inputs

Exit Pi. From the original project:

cd ~/projects/academy-tutor-course/academy-tutor
mkdir sdk-app
mkdir sdk-app/docs
cp tutor.js cards.json sdk-app/
cp node_modules/@earendil-works/pi-coding-agent/docs/sdk.md sdk-app/docs/sdk.md
cd sdk-app
npx pi

sdk-app is a child of academy-tutor, alongside that project’s node_modules folder. Node can resolve the package installed in the parent. The copied tutor.js is an intentional input for this conversion.

Ask:

Turn tutor.js into a standalone terminal tutor using Pi’s SDK. Use docs/sdk.md and keep the same learning and review rules. I want to run it with node tutor.mjs, answer questions in the terminal and quit when I’m finished.

Open the resulting tutor.mjs and ask Pi to explain how it starts a session and handles an answer. We can do that in the generation conversation; there is no need to change projects merely to inspect the code.

Our SDK application reused the tutor as an inline extension:

const resourceLoader = new DefaultResourceLoader({
  cwd,
  agentDir: getAgentDir(),
  extensionFactories: [makeTutorExtension()],
});
await resourceLoader.reload();
const { session } = await createAgentSession({
  cwd,
  resourceLoader,
  sessionManager: SessionManager.inMemory(cwd),
  tools: ["read", "record_attempt"],
});

makeTutorExtension() packages the tutor behavior for the resource loader. The application creates and owns the Pi session through the SDK. Our reference application uses the SDK’s custom-tool arrangement instead. Both are ways to build on Pi.

The application sends its initial practice prompt with session.prompt(startPrompt()). It then reads an answer from the terminal and sends it through session.prompt(answer), repeating while concepts remain. Its cleanup closes the terminal input, unsubscribes from events and disposes the session. Those are application responsibilities around Pi’s model/tool loop.

There are differences worth reading: the generated app uses the default resource loader, whereas the reference explicitly disables several automatically loaded resources and supplies its own system prompt. That changes which existing customizations can enter the application.

Check the application, then use it

The extension checker cannot simply be pointed at an SDK application. Its test fixture expects an extension export and command handler. Ask Pi to adapt the checks while keeping the same requirements:

Check this application against ../tutor.js’s behavior and ../../requirements.md. The extension checks are in ../../comparison/behavior-checks.mjs. Adapt the relevant checks for an SDK app and compare it with ../../reference/tutor.mjs. Use temporary course data. Explain a test and show me the results.

Check that it tests more than syntax. Empty-course and future-only histories should cause the app to finish without a model call. Invalid dates should be handled visibly rather than quietly losing a concept. The course initially produced only a few SDK checks; we asked for fuller coverage of the relevant rules. The meaning of a check matters more than matching the extension test count.

Now exit Pi and launch the actual application:

node tutor.mjs

Practice starts immediately. Answer a question yourself. Inspect attempts.jsonl, then type /quit at the application’s learner prompt when you want to stop. An in-memory Pi conversation does not mean the learning record disappears: the record_attempt tool writes that file independently.

Test due review without waiting a day

Start ordinary Pi from sdk-app and ask it to prepare a separate demonstration:

Make a review-demo folder with a copy of this app, its course and learning history. Keep just the concept I practiced, keep a backup of the copied history, and make its copied review date overdue. Show me the card and the before-and-after date so I can check the setup.

Inspect the copied card and overdue date. Exit Pi, then run:

cd review-demo
node tutor.mjs

Because the one concept is due, the tutor should ask before explaining. Our answer describing the roles of Action, Brain and Context was accepted as correct and unassisted. The new attempt was appended with a review date three days later.

Run node tutor.mjs again. There should now be no eligible concept: the only card has a future review date, so the app exits without starting another lesson. The one-card setup is what makes this result unambiguous; a larger course could still contain unseen concepts.

SDK milestone complete: you can launch your own tutor application, distinguish learning from recall, and verify that a saved review moves out of the eligible queue. You can now use it with your own material or continue to Prime.

7. Adapt the tutor to Prime Agent

Advanced continuation, sections 7–9. Start with the original academy-tutor/tutor.js and cards.json from the Pi build. The SDK app is not required. We’ll add Prime separately, keep its learning history separate, and inspect refinement and background execution. These exercises use model calls through your connected account.

Now add a requirement: the tutor should use experience and feedback to adapt its future teaching instructions. We return to the playbook. The outcome remains learning; the new requirement leads us to revisit the design, foundation, implementation and evaluation.

Prime Agent began as a hard fork of Pi and changed important parts of its architecture. It retains familiar extension concepts and internal package names, but we must check Prime’s actual capabilities instead of assuming every Pi tool remains available.

Prime provides a persistent Python execution environment, saved harness state and refinement mechanisms, and background workers with scheduling. We’ll inspect refinement and scheduled execution directly. Our short tutor exercise does not establish a large-context performance benefit from the persistent runtime.

Install and connect Prime

At the shell, check uv --version. If it is missing, follow uv’s installation guide. Then use the official Prime installer with the course version:

curl --proto '=https' --proto-redir '=https' -fsSL https://app.primeintellect.ai/prime-agent/install.sh | PRIME_AGENT_VERSION=0.9.6 sh
prime-agent --version
cd ~/projects/academy-tutor-course/academy-tutor
prime-agent

Expect 0.9.6. Inside Prime, /login connects a provider and /model selects the model. We used ChatGPT authentication and GPT-5.5, matching the Pi exercises. If the initial Prime-hosted login is not the connection you want, use the provider login flow for your own account.

Ask Prime to adapt the tutor

Ask Prime itself to perform the port:

I want to use our Academy Tutor in Prime. Check tutor.js against Prime’s extension documentation, then make an adapted copy at .prime/agent/extensions/prime-tutor/index.ts with cards.json beside it and its own learning history. Explain what changed and why.

Give it Prime’s versioned extension documentation. The installed package in our project is Pi, so its local documentation is not sufficient evidence of Prime compatibility.

Our adapted extension lives at:

academy-tutor/
  .prime/agent/extensions/prime-tutor/
    index.ts
    cards.json
    attempts.jsonl    ← created after the first saved assessment

Prime discovers this project extension. Use /reload after creating or changing it, then /academy to start practice. Open index.ts and ask Prime to walk through the changes. If your generated layout differs, establish the actual entry file and data paths before launching it.

Understand the tool difference

The first port retained read in its active-tool list. Prime initially claimed that was fine. We investigated the registered tools rather than accepting the claim.

Ask Prime to check its actual tool registry and show the result: “Which tools are registered in this Prime session? Check the extension API and show me the output, including whether read exists.” In our investigation, a temporary diagnostic extension used pi.getAllTools() to print the registered tools.

The tool registry showed ipython, plus our custom record_attempt when the extension was loaded. The corrected selection was:

pi.setActiveTools(["ipython", "record_attempt"]);

pi is the name this extension gives its Prime extension API parameter; it reflects the shared API ancestry.

Why could practice start before this correction? The extension read cards.json using JavaScript, selected the queue and supplied it to the model. That did not require the model to call a read tool. Listing a missing tool name did not create the tool.

In Prime’s default model-facing configuration, Python provides the execution interface for file operations and other programmatic work. A persistent kernel can retain working data between calls. The model uses that interface; the learner does not need to create Python variables manually. This architecture is described in Prime’s RLM documentation.

The model sends code to a persistent Python workspace and receives selected results; working data remains in the workspace.

Our custom recording tool still matters. Registering it defines its schema and implementation; selecting it makes it available alongside ipython. Python availability does not itself implement our review policy.

Pi exposes separate built-in tools; Prime uses ipython as its default execution interface. Our recording tool can be added to either.

The diagram compares the foundations’ built-in interfaces plus our custom tool. Our Pi tutor selects only read and record_attempt from that larger set.

Run /academy, answer a question and inspect the saved record. Unlike the Pi reference’s working-directory paths, this adapted extension stores data beside its own entry file:

const extensionDir = dirname(fileURLToPath(import.meta.url));
const cardsPath = join(extensionDir, "cards.json");
const attemptsPath = join(extensionDir, "attempts.jsonl");

That is why we inspect the extension’s actual directory when looking for its history.

8. Inspect manual and automatic refinement

A model producing a different answer is weak evidence that a refinement mechanism worked. We want to see what changed, where it was saved and what entered the subsequent conversation.

Prime’s harness-state documentation describes persisted prompt notes, memories and refinement events. /refine reviews experience and changes supplemental state; it does not rewrite the base system prompt. Global entries persist under ~/.prime/agent/harness/; session-local entries belong to that session’s artifacts.

A model reviews experience, keeps existing state or saves instructions and memory, which can enter later context.

Save a teaching preference after observing the need

First use the tutor. In our practice session, a partial answer was recorded and the conversation moved toward the next concept. We wanted an example and a follow-up before moving on. That gave the refinement a concrete purpose.

Enter this in Prime:

/refine --global When tutoring me in Academy, if my answer is partial, give a concrete example and ask a follow-up to check my understanding before moving on to the next concept.

Ask what changed and inspect the saved entry. This global instruction is a reusable teaching preference, rather than a record of a particular learner answer.

Start a new conversation with /new, then /academy. Try an incomplete answer. In our new session, Prime gave a travel-booking example to distinguish the agent, model and harness, then asked a follow-up question before proceeding.

Now ask it to investigate the evidence:

Pause tutoring. Find our saved teaching instruction and show me where it appeared in this session’s log. Show the original entry as well as the full saved instruction.

Read the file output it uses to answer. Our session log contained a compact harness_digest with a shortened version of the instruction. The full instruction existed in the global harness-state file; its complete text appeared in the conversation later, when the agent read it during our investigation.

That distinction matters. We can show persisted state, a digest entering the new conversation and the subsequent teaching behavior. We should not say the full instruction was present from the beginning when the log shows only the digest.

Observe an automatic refinement

For a short experiment, ask Prime to make a separate copy:

Let’s try automatic refinement in a separate adaptation-demo folder beside this project. Copy .prime/agent/extensions/prime-tutor/index.ts as tutor.ts with its cards beside it and no learning history. Set this project’s autoRefine.turnInterval to 3, keep my global settings unchanged, and show me the setup before I launch it.

The temporary project setting is:

{"autoRefine":{"turnInterval":3}}

It belongs in that demonstration project’s .prime/agent/settings.json. Ask Prime to show the resulting setting and copied files before proceeding.

Wait for Prime to finish, clear its input, then press Ctrl+D inside Prime to return to the shell. This detaches the interface; it does not stop the background worker. We’ll examine that distinction and the stop command in the scheduling section. At the shell:

cd ~/projects/academy-tutor-course/adaptation-demo
prime-agent -e tutor.ts

Enter /academy and practice normally. The interval counts assistant/model turns, not questions answered. A question can involve a model response, a recording-tool request and another model response; automatic refinement can therefore occur during the first learning question.

After practicing, ask:

While we were practicing, did Prime refine anything automatically? Show me the original event from before this question, what triggered it and what changed.

We found an event with source auto that created a session-local progress memory about the ABC follow-up. This was distinct from the earlier explicit, global /refine operation.

The result also revealed a limitation. After the learner answered the follow-up and moved on, the progress memory had become stale. The agent updated it when we investigated. An automatic change had occurred, but that particular memory was not strong evidence of better teaching.

We also saw the model try to record the follow-up answer again. The tool rejected it because the concept had already been removed from the remaining queue after the first assessment. The teaching instruction had adapted; the deterministic recording rule had not changed. That rejection helps us see the division of responsibilities in operation.

Ask Prime to remove the temporary project interval when you finish:

Remove the temporary turn-interval setting from this demonstration project and show me what remains.

The lesson is to separate adaptation evidence from learning improvement. A saved change entering later context demonstrates adaptation. Establishing that the learner understands or remembers more requires an evaluation of learning outcomes.

9. Schedule a review after closing the interface

Our tutor calculates due dates, but writing dueAt into a file does not automatically create a scheduled job. Let’s separately test whether Prime can deliver a review prompt while its terminal interface is closed.

A human schedules the demonstration worker. Connecting saved due dates to jobs and notifying the learner remain unbuilt.

Prime’s background-agent documentation distinguishes the interface from the running worker. The terminal interface is where we interact. The worker is the running agent process. The session is its conversation and associated state. Detaching closes our connection to the worker; stopping ends the worker.

Prepare one due concept

After Prime finishes removing the temporary setting, clear its input and press Ctrl+D inside Prime to return to the shell. Then return to the original project and start Prime:

cd ~/projects/academy-tutor-course/academy-tutor
prime-agent

Ask:

Help me test whether our tutor can start a review while its chat interface is closed. Make a new scheduled-review folder beside this project. Copy our Prime tutor from .prime/agent/extensions/prime-tutor/index.ts as tutor.ts, and use one course card with a clearly labeled made-up attempt that is already due. Show me the card and its due date.

Inspect the card and synthetic attempt. Keeping one due card makes the expected result clear and leaves the original learning history alone.

After Prime finishes and its input is empty, press Ctrl+D inside Prime to return to the shell. Then run:

cd ~/projects/academy-tutor-course/scheduled-review
prime-agent -e tutor.ts

Inside Prime, name the running agent:

/name academy-review-today

Use a different name if an earlier worker still has this name, and replace academy-review-today with your chosen name in every prompt and command below. Naming it identifies the target for the scheduled prompt; it does not start practice.

Ask:

Use Prime’s one-time scheduler to send /academy to academy-review-today one minute from now. I’ll close the interface while we wait. Show me the job and when it will run.

Check the target, prompt, scheduled time and job identifier in the result. Prime also supports recurring schedules and heartbeats, but a one-time job is enough for this experiment. A schedule delivers specified work at a specified time; a heartbeat is a periodic opportunity for an agent to check whether attention is needed.

Detach, then reconnect to the same worker

After Prime finishes and its input is empty, press Ctrl+D inside Prime. You should return to the shell. This key means something different if your editor or shell receives it, so confirm you are interacting with Prime before pressing it.

At the shell:

prime-agent list

Find academy-review-today. Its client count should be zero. Note the message count, wait until after the scheduled time, then list it again. We observed the message count increase while no interface was attached.

Reconnect:

prime-agent attach academy-review-today

The review question should already be waiting. We are reconnecting to the running worker, rather than merely reopening an old conversation in a new process.

Ask:

Show me that this review came from the job we scheduled earlier. Compare its job ID, scheduled time and completion record. Did it save a learning attempt yet?

The job ID and completion record matched the original scheduled job. The tutor had asked its review question, but no learner had answered, so it had not saved a new attempt. A background prompt can start the interaction; it cannot supply the learner’s response.

Finish by detaching again, then stopping this worker:

prime-agent stop academy-review-today
prime-agent list

Earlier exercises also left detached workers running. Run prime-agent list again and identify those belonging to your course projects from their project/session information. For each one you have finished using, run prime-agent stop <agent>, replacing <agent> with that worker’s listed identifier or unique name. Leave unrelated workers alone; if a listing is unclear, use prime-agent attach <agent> to inspect the session before deciding. Detach with Ctrl+D to return to the shell before stopping it.

This experiment demonstrates execution with the interface detached on a running machine. It does not yet connect every saved due date to a scheduler or deliver a notification to a learner. Those are separate application requirements. We have established the underlying behavior we would need before making those connections.

10. From Academy Tutor to your own agent

Academy Tutor brings us back to the question we started with: how do we turn a general agent’s abilities into a tutor that remembers practice and follows a consistent review policy? We gave the model course material and room to teach, while the surrounding harness selected eligible practice, validated assessment submissions, preserved learning history and applied the review policy. Designing these responsibilities together gave us a tutor whose behavior we could understand and check.

Pi let us build that design inside an existing agent through an extension, then take ownership of the interface and lifecycle through its SDK. Prime provided mechanisms for saved experience to influence future instructions and for reviews to begin without an open chat interface. Inspecting those mechanisms also showed an important distinction: instructions can adapt while application rules stay unchanged, and a saved refinement still needs evidence that it helps the learner.

Academy Tutor gives you both a working example and a method you can carry into another domain. You can build on an existing agent, decide what needs to change, and work with AI to implement it—with a clear understanding of how the resulting agent works and how to evaluate it.