← All projects

Agent Applications

A vendor-neutral reference architecture for building and operating AI agent systems - the products built around persistent, tool-using agents that keep a durable workspace and accumulate real work over time. Published openly as a working draft, for collaboration.

ChatGPT and Claude run agents in widgets. An agent can show up as a WhatsApp or Telegram contact. A coding agent can pick up a pull request and hand the work back there. A business agent can wake on an event with no conversation at all. These look like different products, but they share a shape: an agent returns to work it left unfinished, uses tools, and leaves behind results that outlive the session. Terms like chatbot, copilot, and agent harness each describe a part. This paper names the whole thing an Agent Application - a software system whose primary unit of execution is one or more persistent, tool-using agents operating in durable workspaces to produce or maintain durable artifacts - and works out the architecture that follows from that definition.

It is written as a reference architecture rather than an essay. It identifies the elements a system of this kind tends to have, assigns responsibilities to each, and follows an Agent Project through release, deployment, operation, and change, with drawn views of the system context, the layered stack, the logical structure, the release-and-instance lifecycle, the execution and lifecycle states an instance moves through, and the trust boundaries - including one setting a web application's lifecycle beside an Agent Application's to show where the two stop matching. It is explicitly not a technical standard, a package format, or a reference implementation; a real system can use different names and technologies, and the useful questions are whether it handles these responsibilities, who owns each one, and which boundaries need to interoperate. Because it covers design, development, operation, security, and ecosystem roles, it names the readers it is for - architects, framework authors, platform operators, security teams, buyers, standards contributors - and offers each a shorter path through it.

The argument starts from an analogy to 1994. The web had HTTP, HTML, and browsers before it had an application model; frameworks arrived and settled what people meant by a "web application." Agent developers now have the equivalent building blocks - the Model Context Protocol for tools, portable Agent Skills for instructions, sandboxed compute, and harnesses that run the agent loop - without the application model that turns them into an industry. One consequence the paper draws out: today's agent harnesses are early Agent Application frameworks, and the capabilities they still lack are a roadmap for the people building them. A framework covers the build side - application formats, conventions, tooling, evaluations, connections to outside systems, audit trails. A platform covers the operate side: it packages a tested project into an immutable release and runs long-lived instances from it.

The paper also names the practice it thinks this calls for. Agent Application Programming is the work of decomposing a use case across code, instructions, tools, agents, workflows, state, and policy - deciding what should be written as a deterministic program, what belongs in natural language the agent reads, and what has to be enforced as policy rather than requested.

Much of the paper is about what changes when software stops being identical everywhere it runs. Each instance accumulates its own knowledge, artifacts, and local natural-language instructions, so a fleet diverges in program as well as state - which makes updates, provenance, evaluation, and portability harder than deploying a new version. The paper also argues that the boundary deciding what deserves its own instance is privacy rather than identity: an agent can read its whole workspace, so the workspace has to be drawn where mutual visibility is acceptable. Alongside the definition it maps which existing standards already cover parts of the stack and marks the interfaces - packaging and conformance, artifacts, workspace and lineage portability, delegated authority - where shared contracts would let independent systems interoperate.

Rather than proposing a specification nobody is ready to implement, the paper states a capability model: what an agent project has to be able to express, what a platform has to be able to do with it, and then a direct comparison of how far each of today's agent harnesses already covers that list. The comparison is the more useful half of the argument, because it turns a category claim into something a framework author can check their own product against - and it is falsifiable now, where a normative contract would be waiting on two independent implementations that need the same boundary.

Vendor neutrality here is a constraint on the text, not just on where it is published: every product the paper names is one its author has no stake in. It comes with a Hello World you can actually run - a notebook agent built as a real project in a real agent harness, with the committed files, a recipe, tools, and an evaluation that checks both that its notes survive a restart and that two users never share a workspace, so the mapping from the paper's elements onto working code is something you can inspect rather than take on trust. Alongside it are concept pages, drawn architecture views, and a plain-text edition for machines. The whole thing is a working draft under an MIT license, and the first thing it asks of a reader is small and checkable: take one product you know well and say whether the four defining properties make it an Agent Application, not one, or a boundary case - and if the definition failed you, name what you needed to tell the category apart from chatbots, workflows, and task runners. Counterexamples, prior art, sharper definitions, and implementation experience are all welcome through the public repository.