⏱️ Reading time: 15 min

Pi 1.0 launched today, October 1, 2026, with a short list of new features and an even shorter rule behind them: nothing gets into the agent harness until it proves it’s worth the complexity it adds. Earendil, the company that maintains it, has spent months applying that rule to decide whether to support MCP, Codemode, or any other trend in the agentic ecosystem.

📑 En este artículo
  1. TL;DR
  2. What is an agent harness?
  3. Why the minimalist philosophy matters
  4. How an agent harness works under the hood
  5. MCP and Codemode: the two ways to equip an agent harness
  6. Practical code examples
  7. Getting started with Pi 1.0
  8. Real-world use cases
  9. Common mistakes and best practices
  10. Comparison: minimalist harness versus full-featured harness
  11. Deeper dive: what Pi 1.0 brings under the hood
  12. Frequently Asked Questions
    1. What exactly is an agent environment?
    2. How is Codemode different from MCP?
    3. When does a minimalist agentic framework make more sense than a full-featured one?
    4. Does Pi 1.0 replace other coding agents like Claude Code or Codex CLI?
    5. What is Pi Durable, and why isn’t it part of Pi’s core?
  13. References

That launch is just the hook. The real question is different: what is an agent harness, and when does it make sense to choose a minimalist one instead of one that comes with everything built in out of the box? That’s what we’re going to break down here.

TL;DR

  • An agent harness is the layer that connects a language model to tools and memory so it can act.
  • MCP connects already-built tools; Codemode lets the model write a multi-step script in a single round.
  • Pi 1.0’s minimalist philosophy waits months before adding a feature to the harness’s core.
  • Deferred tool loading and cache warming show how Pi 1.0 reduces the cost of every model turn.
  • Installing Pi takes a single command: curl on macOS and Linux, PowerShell on Windows.

What is an agent harness?

An agent harness is the software layer that connects a language model to the real world. It receives what the model generates, decides which tool to run, executes it, and returns the result so the model can keep reasoning until the task is solved.

Without that layer, a language model only produces text. With it, the model can act on files, terminal commands, or external APIs. The name comes from a simple metaphor: the model is the engine, and the harness is what connects it to the rest of the machinery.

Claude Code, Codex CLI, Cursor in agent mode, and Pi are examples of the same category: programs that wrap a model and give it hands. What sets a harness apart from a simple chat is the loop: a chat ends when the model responds, a harness keeps iterating until the task is resolved.

Pi 1.0 launched on October 1, 2026, according to Earendil. Foto de Oshan De Silva en Unsplash

Why the minimalist philosophy matters

Every integration a harness adds (a new protocol, a tool format, an execution mode) takes up space in the model’s context, opens a new attack surface, and adds a piece that someone has to maintain when it breaks. Multiplied across dozens of integrations, a “complete” harness ends up carrying features almost nobody uses, but that everyone pays for in latency, tokens, and risk.

Earendil describes the criterion it uses before adding anything to Pi: “We wait until something has proven itself, and only then do we consider adopting it; weighing its true functionality against its inherent added complexity.” This isn’t a stance against innovation: it’s a bet that most novelties in the agentic ecosystem don’t survive more than a few months, and that integrating them right away means maintaining code that will probably get discarded.

Pi itself is the case study: it was already running the latest models from every provider, it was already the daily coding agent for hundreds of thousands of people per week, and it already served as the foundation for building agentic applications before Codemode ever entered the core. Earendil compares the two lists directly: the list of discarded features is much longer than the list of ones that made it to production.

flowchart LR
    A["A new technology appears"] --> B{"Still in use months later?"}
    B -->|"no"| C["Discarded"]
    B -->|"yes"| D{"Does the benefit outweigh the complexity it adds?"}
    D -->|"no"| C
    D -->|"yes"| E["Integrated into the harness core"]

This is, essentially, the filter Codemode went through before reaching Pi 1.0: months of real-world use, weighed against the cost of maintaining it.

How an agent harness works under the hood

Under the hood, almost every agent harness runs the same loop, regardless of which model it uses or what language it’s written in. The diagram below shows that cycle: the harness builds the context, passes it to the model, interprets the response, and if there’s a tool call, executes it and starts over.

flowchart TD
    A["User sends a task"] --> B["Harness builds the context"]
    B --> C["Model generates a response"]
    C --> D{"Is there a tool call?"}
    D -->|"yes"| E["Harness runs the tool"]
    E --> F["Result returns to context"]
    F --> C
    D -->|"no"| G["Harness delivers the final response"]

Every loop iteration consumes tokens: the system prompt, the catalog of available tools, and the full conversation history travel with every call to the model. That’s why a decision like Pi 1.0’s deferred tool loading, loading only the tools the task needs, has a direct impact on the cost and latency of every turn.

What varies from one harness to another is what happens at the tool execution step: how it decides which one to use, with what permissions, and how much context each loop iteration consumes. That’s where MCP and Codemode come in, the two most common ways of handling that step.

MCP and Codemode: the two ways to equip an agent harness

MCP (Model Context Protocol) is an open protocol that Anthropic published in November 2024 to standardize how a model discovers and calls tools exposed by an external server: each server publishes a catalog of functions with their schema, the model picks one, the harness runs it, and the result returns to the context.

Since its publication, MCP has been adopted by desktop clients, IDEs, and agent frameworks from different providers, which explains why Earendil treats it as the default option for connecting third-party tools instead of inventing its own protocol.

Codemode changes that step: instead of requesting one tool at a time, the model writes a complete script that chains several calls together, and the harness runs it end to end in a sandbox. Earendil illustrates it with a concrete case: Pi writes a script that turns a week’s worth of commits into a short summary, without asking for permission tool by tool.

sequenceDiagram
    participant M as Model
    participant H as Harness
    participant T as Tool
    M->>H: requests to list the commits
    H->>T: runs git log
    T-->>H: returns the list of commits
    H-->>M: partial result
    M->>H: requests to summarize each commit
    H->>T: runs the per-commit summary
    T-->>H: returns the summaries
    H-->>M: final result

That sequence with MCP needs two complete round trips through the context. With Codemode, the same work fits into a single round:

sequenceDiagram
    participant M as Model
    participant H as Harness
    M->>H: delivers a script that lists and summarizes the commits
    H->>H: runs the full script in a sandbox
    H-->>M: returns the final summary

For a minimalist agent harness, adding a new protocol isn’t free: it means deciding how to isolate the code it runs, what happens if a tool takes too long, and how to prevent a malicious MCP server from injecting instructions into the context. Pi waited until it had both pieces figured out before integrating them into the core.

⚠️ Heads up: a third-party MCP server can return text designed to manipulate the model (prompt injection via tool output). Running Codemode in a sandbox without network access or full filesystem access reduces that risk, but doesn’t eliminate it.
OptionWhen to use itAdvantageLimitation
MCPConnecting tools already packaged by third parties (databases, APIs, files)Catalog of servers already built; the model discovers functions by schemaEach call takes up a full round of context, cost grows with multi-step tasks
CodemodeTasks with several chained steps over the same toolsA single round of context for the entire scriptRequires the harness to run code safely, not just dispatch calls

Codemode makes it possible to run several tools in a single round. Foto de Verticz Squencz en Unsplash

Practical code examples

The loop from the previous section can be written in just a few lines. This minimal version illustrates the idea without depending on any specific SDK:

def harness_loop(tarea, modelo, herramientas):
    contexto = [{"role": "user", "content": tarea}]
    while True:
        respuesta = modelo.generar(contexto)
        if respuesta.llamada_herramienta is None:
            return respuesta.texto
        resultado = herramientas[respuesta.llamada_herramienta.nombre](
            **respuesta.llamada_herramienta.argumentos
        )
        contexto.append({"role": "tool", "content": resultado})

Each iteration of the while loop is a step in the flowchart: generate, check for a call, execute, append the result. For the task of summarizing commits with MCP, the expected output of a typical run would be text like:

"The week had 14 commits: 9 in the parser, 4 in tests, and 1 CI fix."

The Codemode version of the same work doesn’t request tools one at a time: the model delivers the complete script and the harness runs it end to end.

const commits = await git.log({ since: "7 days ago" });
const resumen = commits
  .map((c) => `${c.hash.slice(0, 7)}: ${c.mensaje}`)
  .join("\n");
print(resumen);

This is, in code, the same example that Earendil describes for Pi 1.0: a script that turns a week’s worth of commits into a short summary in a single run, instead of one call per commit.

Getting started with Pi 1.0

The only real dependency is having a terminal with internet access: Pi is distributed as a self-contained binary, with no runtime to install beforehand. On macOS and Linux, the official install is a single command:

curl -fsSL https://pi.dev/install.sh | sh

On Windows, the equivalent installer runs from PowerShell:

powershell -c "irm https://pi.dev/install.ps1 | iex"

Pi Durable, the experimental package for long-running applications, installs separately and doesn’t touch Pi’s core:

npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord

All three packages and Pi itself are MIT licensed, and the full documentation lives at pi.dev. To confirm the install is active, run pi --version in the terminal: if the command isn’t recognized, check that the folder where the installer copied the binary is in your PATH.

Real-world use cases

Earendil describes Pi working as a daily coding agent for people who use it hundreds of thousands of times a week, and also as a substrate for building custom agentic applications on top of it. With Pi 1.0, that substrate adds support for “virtual models”: extensions that users themselves build by combining several models behind a single interface.

The example Earendil gives is concrete: an extension where planning runs on Claude Opus, implementation runs on GPT, and a router called Jev decides at what point in the task it makes sense to switch from one to the other. The user reloads, starts a new session, picks the router/auto model, and Jev handles detecting when the task has shifted from planning to implementation.

Another use case the announcement highlights is cost observability: Pi’s /session command breaks down how much each model spent within a session and how much was saved through caching, something key when a single task splits work across two or three different models.

For LATAM teams building their own agents, the practical lesson is the same one Earendil applies internally: measure how much each integration costs in tokens and maintenance before adding it to production, instead of copying another team’s full stack from the start.

Common mistakes and best practices

Misunderstood minimalism creates as much risk as maximalism without criteria. These are the most common mistakes when adopting, or avoiding, new protocols in a custom harness:

  • Confusing minimalism with feature poverty: a minimalist harness doesn’t avoid new features, it avoids the ones that haven’t proven their worth. The difference is the entry criteria, not the feature count.
  • Adding MCP from day one without a sandbox: connecting a third-party server without isolating its execution exposes the model’s context to malicious tool outputs.
  • Treating Codemode as a free optimization: running model-generated scripts requires a secure execution environment, not just changing the call format.
  • Not measuring the context cost of each integration: every protocol a harness supports takes up system prompt tokens even when it’s not in use, unless the harness loads tools in a deferred way.
  • Waiting indefinitely: minimalism isn’t inertia. Earendil took months to add Codemode, not years; the criterion is evidence of use, not aversion to change.

There’s no single answer between minimalism and full integration: it depends on how many people maintain the harness and how fast the surrounding ecosystem changes.

ApproachWhen it fitsAdvantageRisk
MinimalistSmall teams, well-defined tasks, high tool turnover in the ecosystemSmaller attack surface, less maintenance, every feature already proved its worthNew features take weeks or months to reach the core
Full (everything integrated)Organizations that need to standardize many integrations from day oneFewer loose pieces to assemble, immediate official supportEvery unused protocol keeps consuming context and maintenance cycles

💭 Key takeaway: a minimalist agent environment isn’t one poor in features, it’s one where every feature already earned its place through real-world use before entering the core.

Deeper dive: what Pi 1.0 brings under the hood

Besides Codemode, Pi 1.0 adds deferred tool loading: tool definitions the model won’t need for a specific task aren’t loaded into the context from the start, which reduces the fixed cost of every turn as the tool catalog grows.

It also adds cache warming for Anthropic models: the harness pre-warms the prompt cache before the user types, so the session’s first response doesn’t pay the full cost of processing the context from scratch. And it adds mid-conversation system messages: prompt or available-tool changes the harness can inject without restarting the session, something that previously required cutting the conversation and losing the history.

The rest of the list is more cosmetic (a new theme for the terminal interface, full-screen mode by default), but it confirms the pattern: Earendil prioritizes polishing what already exists before announcing something new. Pi Durable, on the other hand, stays outside the core precisely because it pushes Pi into different territory (long-running conversations and tasks outside the terminal) that doesn’t fit how most people use the harness today.

Your next step: install Pi with the command above, build a single-tool harness (for example, listing the commits in a local repo), and only then evaluate whether you need to add an MCP server.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What exactly is an agent environment?

It’s the program that connects a language model to real actions: terminal, files, APIs. It receives the model’s output, executes what it requests, and returns the result so the model can keep reasoning until the task is done. The metric that matters isn’t how big the tool catalog is, but how well it decides which one to use at each step.

How is Codemode different from MCP?

MCP exposes tools as individual functions that the model calls one at a time. Codemode lets the model write a script that chains several calls together and runs it end to end in a single round of context. Both solve the same problem with a different context cost per task.

When the team prioritizes a small attack surface and low maintenance over having every integration available from day one. The cost is waiting weeks or months for a new feature to reach the core. The decision usually depends on the size of the available maintenance team, not just the budget.

Does Pi 1.0 replace other coding agents like Claude Code or Codex CLI?

Not necessarily: every harness defines its own criteria for what it adopts and when. Pi 1.0 stands out for waiting until a capability proves itself in real-world use before adding it to the core, instead of integrating it as soon as it appears. The real comparison depends on how often your workflow needs an integration that hasn’t yet reached the core of one or the other.

What is Pi Durable, and why isn’t it part of Pi’s core?

It’s a separate experimental package for long-running agentic applications, designed for conversations and tasks that run outside the terminal. Earendil kept it apart so as not to force Pi into being something different from what it is. Sharing that risk with the core would have meant loading its complexity into every terminal session, even for people who never use it.

References

📱 Do you like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Cong Long Vu en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: Tech NewsTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.