GuidesAgents on the web

What is AG-UI? The Agent-User Interaction protocol

AG-UI streams an agent's work to an app as typed events over HTTP. What 1.0 changed, how it differs from MCP, A2A and ACP, and whether OpenClaw speaks it.

October 2, 2026The Everpod team
The short answer

AG-UI, the Agent-User Interaction protocol, is an open, MIT-licensed standard from CopilotKit for the connection between an agent running somewhere and the app a person uses to reach it. The app sends one HTTP POST carrying the conversation so far; the agent answers with a stream of typed JSON events over Server-Sent Events: the run starting, the reply arriving token by token, tool calls, reasoning, shared state, the run finishing. Version 1.0, the first with a written specification, shipped on September 17, 2026 and defines 31 event types in eight families.

It sits beside MCP (agent to tools) and A2A (agent to agent), and does for web and phone apps what ACP does for code editors. Frameworks from Microsoft, Google, LangChain, CrewAI and Pydantic support it. As of October 2, 2026, neither OpenClaw nor Hermes speaks it out of the box: OpenClaw needs a community plugin, and Hermes’s adapter is an open pull request.

One request in, a stream of events out

The 1.0 specification puts the protocol in one line: “one request in, one ordered stream of typed events out, carrying everything the user sees of the agent.” The request is a RunAgentInput, posted to the agent’s endpoint. It names the conversation (threadId) and this turn (runId), carries the message history, and can add the app’s own tools, extra context, the state the last run left behind, and answers to anything the last run asked. The reply is a text/event-stream with exactly one JSON event in each frame. The HTTP binding’s own example:

POST /agent HTTP/1.1
Content-Type: application/json
Accept: text/event-stream

{"threadId":"thr-1","runId":"run-1","messages":[…]}

HTTP/1.1 200 OK
Content-Type: text/event-stream

data: {"type":"RUN_STARTED","threadId":"thr-1","runId":"run-1"}

data: {"type":"TEXT_MESSAGE_START","messageId":"msg-1","role":"assistant"}

data: {"type":"TEXT_MESSAGE_CONTENT","messageId":"msg-1","delta":"Hello."}

data: {"type":"TEXT_MESSAGE_END","messageId":"msg-1"}

data: {"type":"RUN_FINISHED","threadId":"thr-1","runId":"run-1"}

Everything else the agent can say fits one of the eight families:

FamilyEventsWhat the app gets
Runs and stepsRUN_STARTED, RUN_FINISHED, RUN_ERROR, STEP_STARTED, STEP_FINISHEDWhen a run begins, ends or fails. The run events are the only ones an agent must emit.
Text messagesTEXT_MESSAGE_* (4)The reply, streamed as it is generated.
Tool callsTOOL_CALL_* (5)A tool the agent calls, its arguments as they stream, and its result.
ReasoningREASONING_* (7)The model’s thinking, plus encrypted provider data the app stores and sends back unread.
StateSTATE_SNAPSHOT, STATE_DELTA, MESSAGES_SNAPSHOTShared state and the conversation itself, replaced whole or amended by JSON Patch.
ActivityACTIVITY_SNAPSHOT, ACTIVITY_DELTAStructured progress the app draws as its own widget.
SubagentsSUBAGENT_* (3)Which delegated agent produced which output.
PassthroughRAW, CUSTOMEscape hatches for anything the protocol does not define.

Most explainers, the protocol’s own GitHub README among them, still say AG-UI has about 16 event types and runs over SSE or WebSockets. That describes the 0.x protocol. In 1.0, HTTP with Server-Sent Events is the binding every HTTP implementation must support, a binary binding (the same POST, answered with length-prefixed protobuf frames) is optional, and WebSockets, message buses and in-process pipes are custom transports that must keep the same ordering and termination rules. 1.0 also added run outcomes, interrupts, subagents and activity events, and renamed thinking to reasoning. A 0.x agent keeps working against a 1.0 client, which translates the retired shapes as they arrive.

What the agent and the app each implement

The agent side is an HTTP endpoint plus a bridge that turns a framework’s native events into protocol events. The spec keeps that bar low on purpose: “The mandatory surface is the run lifecycle and nothing else,” so an agent that emits a start, some text and a finish conforms. Most people never write this part, because the frameworks listed below ship the bridge, and since the protocol carries no framework concepts, one app can face any of them.

The app side carries more. It sends the input, renders the stream, runs the frontend tools it offered (a confirmation card, a map, anything the agent may ask the app to do), keeps the shared state, and, in the architecture page’s words, “decides what requires the user’s consent.” It has to accept every event family even if it ignores some, must not treat an HTTP 200 as success (a failure after the stream opens arrives as RUN_ERROR), and must not report a run as finished when the connection dropped before RUN_FINISHED.

Runs never pause. When an agent needs an approval or a choice mid-task, the run ends with an interrupt naming what it is waiting for, and the app’s next request carries the answer. The interrupt rules say why: “The protocol has no mid-run channel from the consumer, so the run does not wait.” A conversation is a thread of such runs.

What the spec leaves out is as telling. It does not define how a UI renders anything or how an agent framework is built, and it defines no credential: authentication is “a property of the binding and the application, not of the protocol.” AG-UI is the pipe. CopilotKit’s MIT-licensed components and runtime are one front end built on it, with three first-party SDKs behind the protocol itself (TypeScript, Python and .NET) and community ones for Kotlin, Go, Dart, Java, Rust, Ruby and C++.

AG-UI, MCP, A2A and ACP: which connection each one is

The four get named together because one agent can speak all of them at once. They differ in who is on the other end:

ProtocolConnects the agent toHow it travels
AG-UIThe app a person uses: a web page, a phone app, SlackAn HTTP POST per run, events back over SSE
MCPTools and dataJSON-RPC over stdio or Streamable HTTP
A2AOther agents, found through a published Agent CardJSON-RPC, gRPC or plain HTTP+JSON
ACPCode editors and IDEsJSON-RPC over stdio, the agent run as a subprocess

MCP shapes what an agent can do, A2A lets it hand work to other agents, and AG-UI’s docs describe handshakes that let an AG-UI app “front for” agents reached through MCP and A2A, so the layers stack rather than compete. The useful line is between AG-UI and the Agent Client Protocol. An ACP editor launches the agent on the same machine and talks to it over stdin and stdout; its HTTP transport is still a draft proposal. AG-UI assumes the agent is a service somewhere else. If yours lives on another computer and you want it in a browser tab or on your phone, AG-UI is the one shaped for that.

Two near-names cause most of the confusion, and AG-UI’s introduction carries a warning about the first. A2UI, started at Google, is a format for interface an agent wants drawn (cards, forms, buttons as a declarative component tree), and MCP Apps is the MCP extension for interactive UI from tool servers. Both are payloads; AG-UI is a pipe that can carry them, and A2UI can travel inside its activity events. AG2, formerly AutoGen, is an agent framework that happens to support AG-UI.

Who supports it (checked October 2, 2026)

Each of these is confirmed in the framework’s own docs, not only in AG-UI’s list:

On AG-UI’s own list only: Amazon Bedrock AgentCore, marked supported; community bridges for Anthropic’s Claude Agent SDK and Claude Managed Agents, and OpenAI’s Agents SDK and Cloudflare Agents marked in progress. The clients it lists are CopilotKit for the web, a React Native example, a terminal client, Slack and Microsoft Teams through CopilotKit’s Channels SDK, and, in the protocol’s repository, a Kotlin chat app that runs on Android, iOS and desktop. Thoughtworks’ Technology Radar placed AG-UI in Trial in April 2026.

OpenClaw, Hermes and OpenMuse

OpenClaw has no AG-UI support in its docs or code. One of AG-UI’s creators opened a pull request on July 16, 2026 adding a bundled AG-UI channel; OpenClaw’s automated review asked maintainers for “an explicit product and security decision” on bundling “this new browser-facing Gateway protocol,” recommending an external plugin instead, and the PR was closed unmerged for inactivity on September 9. What exists is clawg-ui, a community plugin (openclaw plugins install @contextableai/clawg-ui) that adds a /v1/clawg-ui route to the Gateway and hands each AG-UI message to OpenClaw’s normal reply pipeline, so an AG-UI app becomes one more channel, like Telegram. Its latest release, 0.7.0, dates from April 29, 2026 and is built on AG-UI’s 0.x packages.

Hermes has none either. Its AG-UI adapter, pull request #65845 from the same AG-UI co-creator, has been open since July 16, 2026 and carries a needs-decision label; it would add an agui section to config.yaml and a hermes agui command. A second, community PR has been open since August 8. Hermes does speak ACP, for editors.

OpenMuse is where September’s attention came from: announced on September 22 as “an open source, self-hostable personal assistant that works with any agent harness.” Built by CopilotKit and MIT-licensed, it is an app for iOS, Android and the web with chat, a persistent browser, an optional Linux workspace, tasks and goals. Given a model, it runs CopilotKit’s built-in agent on an OpenAI, Anthropic or Google key. The “any harness” part is one setting: AGENT_BACKEND=agui with an AGENT_URL and AGENT_TOKEN hands its chat to an external AG-UI endpoint (the comment beside it says this “replaces conversational routing only”), which for OpenClaw today means clawg-ui and for Hermes means the unmerged adapter. Two things its README states that the announcement did not: it is an alpha, and every mode needs a CopilotKit Intelligence project key, the service that stores its chat threads, which CopilotKit runs in its cloud or you host yourself on Kubernetes or AWS ECS.

Reaching an agent on its own computer from a phone or web app

This is the arrangement AG-UI is designed around: the agent is a service, and the app is somewhere else. What it asks of you is an HTTP endpoint the app can reach, and every decision about who may use it.

Two properties matter more on a phone than at a desk. A run lives inside one HTTP response, and the SSE binding has no resumption: if the connection drops mid-run, “the producer may have kept going, but this consumer will never see the rest,” and trying again is a new run. And every exchange is opened by the app, since the input is the only message that travels in that direction. Our reading of the 1.0 spec is that an agent cannot start a conversation over plain AG-UI: a morning report reaches you only if the app asks for it or adds a notification channel of its own. CopilotKit’s separate Intelligence service fills part of that gap with runs you can reconnect to, from another device too; OpenMuse keeps its own task worker and in-app notifications, with device push on its roadmap. An agent on a chat app such as Telegram has no such gap: it can write to you first.

One more fit question for personal agents. AG-UI was shaped around apps that hold the conversation: every request carries the thread’s complete history, and the protocol has no other channel through which earlier turns arrive. OpenClaw and Hermes keep their own sessions and memory on their own disk, so as we read it, a bridge for them has to map an AG-UI thread onto the agent’s session rather than replay the app’s copy. clawg-ui does it by routing each message through the Gateway like any channel message; the closed OpenClaw PR kept a session per conversation.

Your own open-source AI agent, set up for you.

Everpod runs OpenClaw on a private, always-on computer of its own: set up, secured and backed up, with model usage included. You name your agent, and say hello about fifteen minutes later.

Create your agent

First month half price, then $29/mo · model usage included · cancel anytime

Wondering what you’d do with one? See what a cloud agent can do