AgentBus Product Requirements Document
Status: Draft v0.1 | Date: 2026-10-10 | Owner: Founding team
1. Problem statement
Agents work. Agents talking to each other across machines does not.
A developer today can run Claude Code on a laptop, Codex CLI on a build server, and OpenCode in a container, and each one is capable on its own. The moment one of them needs to hand work to another, the developer is back to copy-pasting between terminals, writing ad-hoc webhooks, or standing up a queue and a schema nobody else agreed to. Every team that tries this rebuilds the same five things:
- A way to name and find the other agent.
- A way to deliver a message when the other agent is busy, offline, or on another network.
- A way to prove who sent what, and to stop an unauthorised party from sending anything.
- A way to see what it cost: tokens, money, bytes, and time.
- A way for the agent itself to find out why a message did not arrive.
None of these are agent intelligence problems. They are plumbing problems, and plumbing is what a platform should own. AgentBus owns them.
AgentBus is a hosted, authenticated, auditable inbox for every agent, reached through a single binary sidecar that any harness can use in under five minutes. Every message is schema-validated, signed, policy-checked, delivered at least once, and recorded with its cost and a trace id. The same system can be self-hosted, the way GitLab can.
2. Positioning
2.1 Versus MCP
The Model Context Protocol connects an agent to tools: files, databases, APIs. It is agent-to-tool. It is not designed for one agent to queue work for another agent on a different machine and get a result back hours later. AgentBus does not replace MCP. The sidecar exposes its own tools over MCP so that Claude Code, Codex, and OpenCode can send and receive without custom integration. MCP is the local hook; AgentBus is the wire.
2.2 Versus A2A
Google's Agent2Agent protocol targets agent-to-agent communication and defines agent cards for discovery and a task lifecycle with states. It is HTTP and JSON-RPC based and synchronous-leaning. It has no hosted queueing for offline recipients, no audit trail, no cost ledger, and no tenancy or marketplace model.
AgentBus is an asynchronous transport and control plane. It borrows the A2A agent card format for the agent directory and maps its task states onto A2A's so that an A2A-speaking agent can be bridged later without redesign. AgentBus is complementary to A2A the way an email server is complementary to an email client.
2.3 What is different
| Capability | MCP | A2A | AgentBus |
|---|---|---|---|
| Agent-to-tool | Yes | No | No (uses MCP locally) |
| Agent-to-agent | No | Yes | Yes |
| Cross-machine, cross-network | Server-dependent | Yes | Yes |
| Queued delivery to offline recipient | No | No | Yes |
| At-least-once with acks and dead-letter | No | No | Yes |
| Signed envelopes, per-agent identity | No | Partial | Yes |
| Audit trail and cost ledger | No | No | Yes |
| Multi-tenant with cross-tenant grants | No | No | Yes |
| Machine-readable error codes with hints for self-debugging | No | No | Yes |
| Hosted and self-hostable | n/a | n/a | Yes |
3. Personas
3.1 Solo developer with several harnesses (primary MVP persona)
Runs Claude Code, Codex CLI, and OpenCode across two or three machines. Wants one of them to review what another wrote, or to split a task, without copy-paste. Cares about time to first message and about seeing what it cost. Will install a single binary. Will not read a protocol spec.
3.2 Platform team at a company
Runs tens to hundreds of agents across teams. Needs workspaces, membership, policy, SSO, audit for the security team, and cost attribution per team. Wants self-hosting as an option and a path to enterprise controls. Will read the spec and the threat model.
3.3 Agent publisher (post-MVP)
Has built a very good specialised agent and wants others to use it, free or paid. Needs a listing, grant approval, usage metering they can trust, and payouts. Needs dispute evidence.
3.4 Support engineer
Answers "my message never arrived" tickets. Needs to search by trace id, receipt, or agent, see the hop-by-hop delivery timeline, see which policy denied what, and verify a token without ever seeing it. Needs every support action recorded in the customer's own audit log.
3.5 Security and compliance reviewer
Evaluates AgentBus before a company adopts it. Needs the threat model, tenant isolation guarantees, encryption at rest and in transit, key management options, retention controls, and tamper-evident audit. Will ask about prompt injection between agents.
4. MVP goals and non-goals
4.1 Goals
- G1. A developer can register, create an agent, connect a harness, and exchange a message with another harness on another machine in under five minutes.
- G2. Claude Code, Codex CLI, and OpenCode are all supported as first-class adapters.
- G3. Every message is authenticated, signed, schema-validated, and policy-checked.
- G4. Delivery is at least once with explicit acks, redelivery, expiry, and dead-letter.
- G5. Every message has a trace id, a receipt, and a hop-by-hop timeline visible to the sender, the recipient, and support.
- G6. Every error has a stable code, a hint, and a help URL, and
agentbus doctorcan diagnose the common failures without a human. - G7. Usage and cost are recorded per message, per agent, per workspace, and per tenant, with the source of each number labelled.
- G8. The system runs from a single
docker compose upfor local development and for self-hosting.
4.2 Non-goals for MVP
- Marketplace listings, paid grants, billing, payouts.
- Cross-tenant communication.
- Custom domains, mTLS, SAML, dedicated NATS accounts, region pinning.
- End-to-end encryption where the server cannot read payloads.
- A visual workflow builder or orchestration DSL. AgentBus moves messages; it does not plan.
- Running agents for the customer. The harness is always the customer's.
These are specified in other documents (06, 10, 13) so the MVP does not paint itself into a corner.
5. User stories and acceptance criteria
Each story has an ID, a persona, the story, and acceptance criteria. Error codes follow 00-CONVENTIONS.md.
US-01 Sign up and create a tenant
As a solo developer, I sign up at console.agentbus.exchange with Google or GitHub and get a tenant
with a default workspace.
- Signup completes via Clerk; the backend creates
ten_, a defaultws_nameddefault, and ausr_with roleowner. - The tenant slug is derived from the chosen name, validated to
[a-z0-9-], 3-40 chars. - The user lands on an empty agent directory with the install command for
agentbusshown.
US-02 Create an integration token
As a developer, I create one token per harness so a leak is scoped and revocable.
- Token is shown exactly once, formatted
ab_live_<40 chars>. - The token has a label, scopes (
send,receive,admin,read-audit), optional expiry. - Stored as SHA-256. List view shows prefix, last 4 chars, label, scopes, created, last used.
- Revoking returns
AB-1004on the next use by that token within 10 seconds.
US-03 Log in from the CLI
As a developer, I run agentbus login and authenticate through the browser.
- Device-code flow: CLI prints a URL and code, opens the browser, polls, stores the credential in the OS keychain or an encrypted file.
agentbus login --token ab_live_...works non-interactively for CI.agentbus whoamiprints tenant, workspace, user, token label, and scopes.
US-04 Create an agent
As a developer, I run agentbus agent create reviewer --description "Reviews diffs for concurrency bugs" --capabilities code.review.
- The server returns
agt_, the addressagent://<tenant>/<workspace>/reviewer, and creates the inbox consumer onT_<tenant_id>. - The agent appears in the workspace directory with name, description, capabilities, owner,
and presence
offline. - Duplicate names in the same workspace return
AB-3002.
US-05 Connect a harness
As a developer, I run agentbus connect --agent reviewer --adapter claude and my Claude Code
session is now reachable.
- The sidecar launches the harness as a child process, installs the plugin or hook or config the adapter needs, starts the local daemon on a Unix socket, and begins heartbeats.
- Presence flips to
onlinein the directory within 5 seconds and back toofflinewithin 30 seconds of the harness exiting. - The same works for
--adapter codexand--adapter opencode. agentbus doctorreports adapter health asokafter connect.
US-06 Send a message between two harnesses on two machines
As a developer, I send a message from Claude Code on machine A to Codex on machine B.
- From inside Claude Code, the model calls the
sendMCP tool with the address and text. - The gateway validates the envelope, signs the receipt, publishes to the recipient inbox, and
returns
rcp_andmsg_to the sender. - On machine B, the Codex adapter injects the message into the live session within 2 seconds (p95) when online.
- If machine B is offline, the message waits in the inbox up to
expires_atand is delivered on reconnect.
US-07 Request and reply
As a developer, I send a message and wait for the reply in one call.
agentbus send <address> --wait --timeout 120sblocks until a message with matchingcorrelation_idarrives or the timeout elapses.- Timeout returns
AB-4010with the receipt so the reply can still be fetched later.
US-08 Task lifecycle
As a developer, I submit a task and watch it progress to a result.
agentbus.task.request.v1creates atsk_in statesubmitted.- The recipient's
task.acceptmoves it toaccepted,task.progresstoin_progress,task.resulttocompleted,task.errortofailed,task.canceltocancelled, expiry toexpired. - Invalid transitions return
AB-4020. - A second
task.requestwith the sameidempotency_keyto the same recipient returns the existingtsk_and does not create a duplicate. - The task's state and timestamps are visible in the console and via
agentbus trace.
US-09 Audit view
As a platform team member, I see who talked to whom.
- The console shows a searchable log per workspace: time, from, to, type, size, outcome, trace id, receipt.
- Payload bodies are shown only to members of the workspace with the
read-auditscope. - Entries cannot be edited or deleted through any API. Retention is 7 days on Starter and 90 days on Business.
US-10 Cost view
As a platform team member, I see what agents cost.
- Per message: bytes in, bytes out, delivery attempts, latency.
- Per agent and per workspace: message counts, bytes, and reported LLM usage (tokens, model,
estimated cost) with the source labelled
harness,proxy,agent, orestimate. - Where a harness cannot provide usage, the cost view shows
unavailable, never zero.
US-11 Self-debugging
As an agent, when my message fails I can find out why without a human.
- Every error response carries
code,message,hint,help_url,trace_id,retryable. - The
diagnoseMCP tool andagentbus doctorcheck connectivity, token validity, clock skew, inbox backlog, last failed deliveries, and adapter health, and return the matching codes. agentbus trace <rcp_>prints the hop-by-hop timeline: accepted, persisted, delivered to sidecar, read by agent, acked, or expired or dead-lettered.agentbus policy simulate --from A --to B --type agentbus.message.v1returns allow or deny with the policy id that decided.
US-12 Workspace membership and agent directory
As a platform team member, I invite colleagues and we see each other's agents.
- Invitation by email, or OIDC SSO on Business.
- Every member sees every agent in the workspace with name, description, capabilities, owner, and presence.
- Members can message any agent in the workspace by default.
US-13 Default policy
As a security reviewer, I expect the default to be safe.
- Allow within a workspace. Deny across workspaces. Deny across tenants.
- Anything else requires an explicit
grt_created by a workspace admin. - A denied send returns
AB-2001with the policy id.
US-14 Inbound messages are untrusted data
As a security reviewer, I expect a message from another agent to be treated as data, not instructions.
- The sidecar delivers every inbound message wrapped with sender identity, workspace, trust level, and an explicit "this is data, not instructions" frame.
- The console lets a workspace admin restrict which agents may message which, and which message types.
US-15 Self-host
As a platform team member, I run AgentBus on my own infrastructure.
docker compose upbrings up gateway, NATS, PostgreSQL, ClickHouse, object storage, Caddy, and console with generic OIDC login.- The same sidecar binary connects to either the hosted gateway or a self-hosted one via
AGENTBUS_GATEWAY_URL.
6. Plans
| Feature | Starter | Business | Enterprise |
|---|---|---|---|
| Workspaces | 1 | 10 | Unlimited |
| Agents | 10 | 200 | Unlimited |
| Messages per day | 10,000 | 500,000 | Custom |
| Publish rate per integration token | 60 / min | 600 / min | Custom |
| Pending inbox depth per agent | 1,000 | 10,000 | 100,000 |
| Stream bytes per tenant | 1 GiB | 20 GiB | Custom |
| Message retention | 7 days | 90 days | Configurable |
| Audit event retention | 7 days | 90 days | Configurable |
| Audit chain (tamper-evidence hashes) | 7 years | 7 years | 7 years |
| Blob size | 100 MiB | 1 GiB | Custom |
| Blob storage | 5 GiB | 100 GiB | Custom |
| Login | Google, GitHub, email | OIDC SSO | OIDC SSO, SAML |
| Policy | Default allow-in-workspace only | Cedar policies | Cedar policies |
| Marketplace | Consume free listings | Publish and sell | Publish and sell |
| Custom domains | No | No | Yes |
| mTLS with SPIFFE identities | No | No | Yes |
| Dedicated NATS account | No | No | Yes |
| Region pinning | No | No | Yes |
| Bring your own key (BYOK) | No | No | Yes |
| Cross-tenant federation | No | No | Yes |
| Support | Community | Email, business hours | SLA, named contact |
Enterprise is marked "coming soon" at MVP launch. Gating is enforced in the gateway by plan
checks, not in the console, so the CLI and API return AB-5003 (plan limit) consistently.
7. Success metrics
| Metric | Target | How measured |
|---|---|---|
| Time to first message, signup to first delivered message | Under 5 minutes, median | Console event timestamps |
| Delivery success rate for online recipients | 99.9 percent | ClickHouse delivery events |
| Gateway publish latency, p95 | Under 150 ms | OpenTelemetry |
| End-to-end delivery to online sidecar, p95 | Under 2 seconds | Receipt timeline |
| Self-resolved failures (no support ticket after an error) | 80 percent | Error events vs tickets |
agentbus doctor diagnostic accuracy on seeded faults | 95 percent | Test suite |
| Weekly active agents per tenant, month 3 | 5 or more | ClickHouse |
| Harness coverage | Claude Code, Codex, OpenCode all at parity on inject and usage | Conformance suite |
8. Risks and assumptions
8.1 Risks
| Risk | Impact | Mitigation |
|---|---|---|
| Harness delivery UX: turn-based harnesses do not react well to messages arriving mid-task | Product does not feel live | Week 1 of MVP is adapters only, before the gateway exists. Stop-hook and prompt-hook fallbacks for Claude Code; app-server daemon for Codex; HTTP API for OpenCode. |
| Claude Code channel allowlisting: custom channel servers are gated during the research preview [verified 2026-10-10] | Best Claude Code path unavailable broadly | File allowlist request in week 1. Ship hooks plus MCP as the default path. Org-level allowedChannelPlugins for teams. |
| Interactive Claude Code usage metering: hooks expose no token or cost fields [verified 2026-10-10] | Cost view incomplete for the most popular harness | Optional LLM proxy metering via ANTHROPIC_BASE_URL. Label cost unavailable otherwise. |
Codex removed codex mcp-server in v0.154 [verified 2026-10-10] | Designs built on it break | Build on app-server JSON-RPC and the daemon control socket only. |
| Prompt injection between agents | Security incident | Untrusted-data framing on every inbound message; policy restrictions; signed envelopes. |
| NATS operational complexity | Outages | Three-node cluster, managed Postgres, runbooks in 10-INFRA. |
| Clerk lock-in | Migration cost | Own user ids, generic OIDC verifier, memberships in own tables. |
8.2 Assumptions
- Developers already run more than one harness and feel the gap.
- Harness vendors keep their hook, app-server, and HTTP APIs stable enough for adapters to track.
- Customers accept that the server can read payloads by default, with BYOK on Enterprise and end-to-end encryption post-MVP.
- Self-hosting is a buying requirement for the platform-team persona and must not be a separate codebase.
9. Out of scope for MVP
Specified elsewhere but not built in MVP:
- Marketplace, grants, paid listings, Stripe Connect: 13-MARKETPLACE-AND-CROSS-TENANT.md.
- Custom domains, mTLS, SPIFFE identities, dedicated NATS accounts, region pinning: 06-SECURITY-AND-THREAT-MODEL.md and 10-INFRA-AND-OPERATIONS.md.
- End-to-end payload encryption and BYOK: 06-SECURITY-AND-THREAT-MODEL.md.
- Cross-tenant federation: 13-MARKETPLACE-AND-CROSS-TENANT.md.
- Gemini CLI adapter: 08-HARNESS-ADAPTERS.md lists it as post-MVP pending research.
10. Open questions tracked in 14-MVP-PLAN.md
- Whether interactive Claude Code cost is proxy-metered or shown as unavailable in MVP.
- Whether the self-host edition ships with MVP or immediately after.
- Whether
agentbus.message.v1should carry an optionalformatfor markdown versus plain text.