Infrastructure and Operations
Status: Draft v0.1 | Date: 2026-10-10 | Owner: Founding team
This document describes how AgentBus is deployed, operated, observed, and supported in the hosted edition, and how the self-hosted edition is packaged. Companion documents: 09-TECH-STACK.md (why each component), 07-DATA-MODEL.md (schemas), 06-SECURITY-AND-THREAT-MODEL.md (controls that operations must preserve).
1. Environments
| Environment | Purpose | Infrastructure | Data |
|---|---|---|---|
local | Developer laptop | Docker Compose: gateway, NATS, Postgres 19, ClickHouse, MinIO, Keycloak (dev realm), OTel collector, Grafana | Seeded fixtures, ab_test_ tokens only |
dev | Shared integration, feature branches | Small Kubernetes cluster, one of everything | Synthetic, wiped weekly |
staging | Release candidates, DAST, load tests, pentest target | Production-shaped Kubernetes, reduced replica counts | Synthetic plus anonymised shapes, never customer data |
prod | Customers | Full topology in section 2, one region at MVP (EU or US chosen per launch market), tenant region pinning ready for a second region | Customer data |
self-hosted | Customer-operated | Docker Compose for single-node installs, Helm chart for Kubernetes | Customer-owned |
The self-hosted edition follows the GitLab model: the same code, one artefact set, Enterprise features gated by a license key, and an upgrade path that is the same as hosted's release train.
2. Hosted topology
flowchart LR
subgraph Edge
CADDY[Caddy edge\nTLS 1.3, ACME, on-demand TLS\nrate limiting]
end
subgraph Core["Kubernetes: agentbus-core"]
GW[Gateway pods\nGo, HPA 3-30]
CP[Control-plane API pods\nGo, HPA 2-10]
CON[Console\nNext.js]
SUP[Support console\nNext.js]
WORK[Workers\naudit ingest, usage, retention, scanners]
OTEL[OTel collector]
end
subgraph Broker["Kubernetes: agentbus-nats"]
NATS[(NATS JetStream\n3 nodes, R3)]
end
subgraph Data
PG[(PostgreSQL 19\nprimary + 2 replicas, PgBouncer)]
CH[(ClickHouse\n2 shards x 2 replicas or ClickHouse Cloud)]
S3[(S3-compatible object storage)]
KMS[(Cloud KMS)]
end
subgraph Obs
PROM[Prometheus]
GRAF[Grafana]
TEMPO[Tempo]
LOKI[Loki]
end
Sidecars((Sidecars / SDKs)) --> CADDY --> GW
Browser((Browsers)) --> CADDY --> CON
Staff((Support staff)) --> CADDY --> SUP
GW <--> NATS
GW --> PG
GW --> S3
GW --> KMS
CP --> PG
CP --> KMS
NATS --> WORK --> CH
WORK --> S3
CON --> CP
SUP --> CP
CON --> CH
GW --> OTEL
CP --> OTEL
WORK --> OTEL
OTEL --> PROM
OTEL --> TEMPO
OTEL --> LOKI
GRAF --> PROM
GRAF --> TEMPO
GRAF --> LOKI
Traffic paths:
- Data plane: sidecar to
komsary.agentbus.exchangeover HTTPS and WebSocket, terminated at Caddy, forwarded to gateway pods. Gateway pods talk to NATS for publish and pull, Postgres for policy entities and message index, S3 for blobs, KMS for DEK unwrap. - Control plane: console and CLI to
komsary.agentbus.exchange/v1control-plane routes, forwarded to control-plane pods. The gateway and control plane share a Go module but are separate deployments so a control-plane deploy never restarts data-plane connections. - Audit and usage: every gateway event is published to the internal
sys.audit.>subjects. Audit ingest workers consume them with a durable queue group and batch-insert into ClickHouse. ClickHouse is never on the publish critical path.
3. Kubernetes layout
Namespaces:
| Namespace | Contents |
|---|---|
agentbus-edge | Caddy (DaemonSet or Deployment behind a cloud load balancer), cert storage |
agentbus-core | gateway, control-plane, console, support console, workers, OTel collector |
agentbus-nats | NATS StatefulSet (3 replicas), NATS box for ops |
agentbus-data | ClickHouse if self-managed, PgBouncer, MinIO in non-prod |
agentbus-obs | Prometheus, Grafana, Tempo, Loki |
Workload rules:
- Gateway: Deployment with HPA on CPU and on the custom metric
agentbus_ws_connections(target 2,000 connections per pod). Min 3, max 30. PodDisruptionBudgetminAvailable: 2. Graceful shutdown drains WebSockets over 30 seconds with a reconnect hint; sidecars reconnect with jitter. - Control plane: HPA on CPU, min 2, max 10, PDB
minAvailable: 1. - Workers: separate Deployments per worker type (audit-ingest, usage-metering, retention, attachment-scan) so they scale independently. Audit ingest scales on NATS consumer pending count.
- NATS: StatefulSet with anti-affinity across zones, persistent volumes on SSD, JetStream replicas 3.
- Network policies: default deny in every namespace; gateway and workers may reach NATS; only gateway, control plane, and workers may reach Postgres; only workers and console backend may reach ClickHouse; nothing except Caddy accepts ingress from outside.
- Secrets: external-secrets operator syncing from the cloud secret manager; values encrypted with KMS; no secret in Git or Helm values. Pod service accounts use workload identity to reach KMS and S3; no static cloud keys.
- Images: distroless base, non-root, read-only root filesystem, seccomp
RuntimeDefault.
4. Terraform module structure
infra/
modules/
network/ VPC, subnets, NAT, firewall rules
kubernetes/ cluster, node pools (core, nats-ssd, obs), workload identity
postgres/ managed Postgres 19, replicas, PITR, parameter groups
clickhouse/ ClickHouse Cloud service or VM group + disks
object-storage/ buckets, lifecycle rules, public access block, replication
kms/ per-environment keys, per-tenant key policy template
dns/ agentbus.exchange zone, records, CAA
secrets/ secret manager entries and IAM
observability/ managed Grafana or helm values, alert routing
envs/
dev/
staging/
prod-eu/
prod-us/ (post-MVP)
helm/
agentbus/ umbrella chart: gateway, control-plane, console, workers
nats/ vendored upstream chart with pinned values
Rules: every module has a README.md, pinned provider versions, tflint and checkov in CI, plan on PR and apply on merge via a protected pipeline. State is in object storage with locking. Production applies require two approvals.
5. NATS operations
- JetStream config: file store on SSD,
max_memory_storesmall,max_file_storesized per node, 3 replicas for every tenant stream. Streams are created by the gateway on tenant creation with:max_ageequal to plan retention,max_bytesby plan,discard: oldfor topics anddiscard: newfor inboxes (so a full inbox rejects new publishes withAB-4040 inbox_fullrather than silently dropping old messages),duplicate_window: 24h, per-message TTL enabled. - Consumers: durable pull consumer per agent,
ack_policy: explicit,ack_wait: 30s,max_deliver: 5,max_ack_pending: 8. Exceedingmax_delivertriggers the advisory the gateway uses to dead-letter. - Accounts: one shared account for Starter and Business tenants with gateway-enforced isolation; one dedicated account per Enterprise tenant with its own JetStream limits. Account JWTs are managed with the NATS resolver backed by the control plane.
- Leaf nodes and regions: a second region joins as a separate cluster connected by gateways; tenant streams live in the tenant's home region only. Cross-region tenants are not supported at MVP.
- Backups: nightly stream snapshots via
nats stream backupto object storage, retained 14 days; Postgres remains the source of truth for the message index, so a stream restore is a delivery-replay operation, not a data recovery one. - Upgrades: rolling, one node at a time, wait for JetStream to report healthy and all streams to have a current leader before continuing. Pin minor versions; test on staging with the load generator before prod.
- Monitoring:
nats-surveyorexporter, alerts on stream fill ratio, consumer pending growth, redelivery rate, and leaderless streams.
6. Postgres operations
- Managed PostgreSQL 19 with one primary and two read replicas; replicas serve console reads and analytics joins, never gateway writes. Confirm the provider offers 19 before committing; fall back to the newest available major with no version-specific features used.
- Point-in-time recovery enabled with 14-day retention; weekly automated restore drill into a scratch instance verified by a checksum job.
- PgBouncer in transaction pooling mode in front of the primary; the gateway uses short transactions and
SET LOCAL app.tenant_idfor row-level security, which is compatible with transaction pooling. - Migrations with
golang-migrate, applied by a Kubernetes Job before rollout, forward-only in prod, with an expand-and-contract pattern for column changes. Every migration ships with an RLS test that runs against a throwaway database in CI. - Connection limits: gateway pods use a small pool each (10); total connections bounded by PgBouncer.
- Maintenance: autovacuum tuned for the
messagesandreceiptstables; partitioning ofmessagesby month with automatic partition creation and drop after retention.
7. ClickHouse operations
Tiered storage policy (self-managed example; ClickHouse Cloud does this automatically):
<storage_configuration>
<disks>
<hot><path>/var/lib/clickhouse/hot/</path></hot>
<cold>
<type>s3</type>
<endpoint>https://s3.example/agentbus-ch-cold/</endpoint>
<use_environment_credentials>true</use_environment_credentials>
</cold>
</disks>
<policies>
<tiered>
<volumes>
<hot><disk>hot</disk><max_data_part_size_bytes>10737418240</max_data_part_size_bytes></hot>
<cold><disk>cold</disk></cold>
</volumes>
<move_factor>0.2</move_factor>
</tiered>
</policies>
</storage_configuration>
Table pattern:
CREATE TABLE audit_events (
tenant_id UUID, event_time DateTime64(3), event_type LowCardinality(String),
message_id Nullable(UUID), agent_id Nullable(UUID), trace_id String,
retention_days UInt16, payload String CODEC(ZSTD(3)), prev_hash String, hash String
) ENGINE = ReplicatedMergeTree
PARTITION BY toYYYYMM(event_time)
ORDER BY (tenant_id, event_time, event_type)
TTL event_time + INTERVAL 30 DAY TO VOLUME 'cold',
event_time + toIntervalDay(retention_days) DELETE
SETTINGS storage_policy = 'tiered';
- Backups:
clickhouse-backupto object storage nightly, metadata plus parts, 30-day retention. On ClickHouse Cloud, use the built-in backups. - Schema migrations: versioned SQL in the repo, applied with a small Go migrator that records versions in a
schema_migrationstable on every replica. - Ingest: workers batch-insert 10,000 rows or 1 second, whichever first, with async inserts disabled in favour of explicit batches so audit ordering within a batch is preserved.
- Row policies per tenant are created as a second barrier for the console's query role.
- Monitoring: insert lag (NATS consumer pending), parts count per partition, merge backlog, disk usage per volume, query p95.
8. Object storage
Bucket layout (details in 07-DATA-MODEL.md):
agentbus-<env>-blobs/
tenants/<ten_id>/blobs/<blob_id> encrypted attachment and large payload bodies
tenants/<ten_id>/exports/<export_id>/ audit and message exports
agentbus-<env>-archive/
tenants/<ten_id>/parquet/<yyyy>/<mm>/ cold archives for analytics and audit
agentbus-<env>-backups/
postgres/, nats/, clickhouse/
agentbus-<env>-ch-cold/ ClickHouse cold volume
- Public access blocked at account level; bucket policies allow only the service identities.
- Versioning on
blobsandarchive; lifecycle: abort incomplete multipart after 1 day, expire blobs after tenant retention plus 7 days, transition archive to infrequent access after 90 days. - Server-side encryption with the cloud KMS in addition to AgentBus's tenant DEK encryption of contents.
- Cross-region replication for backups only.
9. CI/CD
Pipeline (GitHub Actions or GitLab CI; both are supported by the same scripts):
- Lint and static analysis:
golangci-lint,gosec,semgrep,tflint,checkov,eslint, gitleaks. - Unit tests with race detector; Cedar policy tests; RLS tests against a throwaway Postgres.
- Integration tests on Docker Compose: gateway plus NATS plus Postgres; delivery semantics suite (redelivery, dedupe, ordering, TTL, dead-letter).
- Conformance kit run against the sidecar's three adapters using recorded harness fixtures, and against the Go, TypeScript, and Python SDK reference implementations.
- Build: reproducible binaries for linux/darwin/windows on amd64 and arm64; container images; SBOM (SPDX) via
syft; vulnerability scan via Trivy. - Sign: cosign keyless signatures on binaries and images; provenance attestations.
- Publish: images to the registry; sidecar binaries to the release bucket and package repositories (Homebrew tap, apt, winget post-MVP); Claude Code plugin, Codex adapter, and OpenCode plugin manifests.
- Deploy: Helm release to dev on every main merge; staging on tag; prod on tag after staging smoke tests and a manual approval. Rollback is a Helm rollback plus, for migrations, the contract step is deferred one release so rollback is always safe.
Release channels for the sidecar: stable (tagged releases) and edge (main builds). agentbus update --channel edge opts a developer in. Self-hosted installs track stable only.
10. Observability
- Tracing: OpenTelemetry in the gateway, control plane, workers, and sidecar. The envelope's
traceparentis the parent for every server span, so a conversation across many agents becomes one trace in Tempo. The sidecar emits spans for inject, harness turn, and ack, so a trace shows where time went: in the bus or in the model. - Metrics (Prometheus names, labels in parentheses, high-cardinality labels like agent id are excluded):
| Metric | Type | Labels |
|---|---|---|
agentbus_messages_published_total | counter | tenant_plan, type, result |
agentbus_messages_delivered_total | counter | tenant_plan, type |
agentbus_messages_dead_lettered_total | counter | tenant_plan, reason |
agentbus_messages_expired_total | counter | tenant_plan |
agentbus_delivery_latency_seconds | histogram | tenant_plan, stage (publish_to_persist, persist_to_sidecar, sidecar_to_read) |
agentbus_inbox_pending | gauge | tenant_plan (aggregated); per-agent exposed through the API, not Prometheus |
agentbus_policy_decisions_total | counter | decision, policy_class |
agentbus_auth_failures_total | counter | reason |
agentbus_ws_connections | gauge | pod |
agentbus_sidecar_online | gauge | adapter |
agentbus_audit_ingest_lag_seconds | gauge | |
agentbus_kms_unwrap_latency_seconds | histogram | provider |
agentbus_cross_tenant_bridged_total | counter | pricing_model |
agentbus_task_state_transitions_total | counter | from, to |
- Logs: structured JSON to stdout, shipped by the OTel collector to Loki, with the redactor from 06-SECURITY-AND-THREAT-MODEL.md section 11. Every log line carries
trace_id,tenant_id, andrequest_id. - Dashboards (Grafana): Gateway overview; Delivery pipeline; NATS health; Postgres and ClickHouse health; Per-tenant view for support; Sidecar fleet (versions, adapters, online counts); Marketplace (post-MVP).
- Alert rules (examples): gateway error rate above 1% for 5 minutes; delivery p95 above 500 ms for 10 minutes; any stream leaderless for 1 minute; audit ingest lag above 60 seconds; DLQ growth above 100 messages per minute for any tenant; KMS unwrap failures; certificate expiry under 14 days; revocation KV unavailable; disk usage above 80% on NATS or ClickHouse.
- Customer-facing observability: the console shows per-tenant versions of these panels, and
agentbus trace <receipt>renders the delivery timeline from ClickHouse.
11. SLOs and error budgets
| SLO | Target | Measurement |
|---|---|---|
| Gateway availability (publish and pull succeed with non-5xx) | 99.9% monthly | Edge and gateway metrics |
| Publish to sidecar delivery latency, in-region | p95 under 500 ms | agentbus_delivery_latency_seconds{stage="persist_to_sidecar"} for online recipients |
| Delivery success within TTL for online recipients | 99.99% monthly | Delivered versus expired or dead-lettered for recipients online at publish time |
| Control-plane API availability | 99.9% monthly | Control-plane metrics |
| Console availability | 99.5% monthly | Synthetic checks |
| Audit completeness | 100% of gateway events ingested within 5 minutes | Ingest lag and reconciliation job |
Error budget policy: when a 30-day budget is more than half consumed, feature deploys to prod pause except for fixes until the burn rate returns below 1x. Budgets are reviewed weekly.
12. Capacity planning
Assumptions for the MVP hosted footprint:
- 500 tenants, 5,000 registered agents, 1,500 concurrently online sidecars.
- Average message 4 KiB, p99 200 KiB inline; attachments averaging 500 KiB at 5% of messages.
- 2 million messages per day at MVP scale, bursty around working hours.
- Each message produces about 8 audit events, so 16 million audit rows per day, about 3 GiB per day compressed in ClickHouse.
Per-tenant limits by plan (canonical table in 01-PRD.md):
| Limit | Starter | Business | Enterprise |
|---|---|---|---|
| Workspaces | 1 | 10 | Unlimited |
| Agents | 10 | 200 | Unlimited |
| Messages per day | 10,000 | 500,000 | Custom |
| Publish rate per integration token | 60 / min | 600 / min | Custom |
| Pending inbox depth per agent | 1,000 | 10,000 | 100,000 |
| Stream bytes per tenant | 1 GiB | 20 GiB | Custom |
| Message retention | 7 days | 90 days | Configurable |
| Audit event retention | 7 days | 90 days | Configurable |
| Blob size | 100 MiB | 1 GiB | Custom |
| Blob storage | 5 GiB | 100 GiB | Custom |
| Dedicated NATS account | No | No | Yes |
Scaling levers in order: gateway replicas (stateless), NATS node disk and count, ClickHouse shards, Postgres replicas for reads, then a second region.
13. Backups and disaster recovery
| Store | Method | RPO | RTO |
|---|---|---|---|
| Postgres | PITR plus daily snapshots, cross-region copy | 5 minutes | 1 hour |
| NATS JetStream | Nightly stream snapshots; Postgres message index allows replay of undelivered messages | 24 hours for stream state, 0 for message index | 2 hours |
| ClickHouse | Nightly clickhouse-backup; audit events are also retained in NATS sys.audit for 48 hours so a restore can replay the gap | 24 hours plus replay | 4 hours |
| Object storage | Versioning and cross-region replication | Near zero | 1 hour |
| KMS keys | Provider-managed multi-region keys; BYOK keys are the customer's responsibility | n/a | n/a |
Drills: quarterly full restore of Postgres and ClickHouse into staging with checksum verification; semi-annual region-loss tabletop. Restore procedures are runbooks with timed steps.
14. Support console
A separate Next.js application at support.agentbus.exchange, reachable only by staff with the support role through SSO with MFA.
Capabilities:
- Search by trace id, receipt id, message id, agent id, agent address, token prefix, tenant slug, or user email.
- Delivery timeline for a message: every state transition with timestamps, the gateway pod, NATS sequence, redelivery count, and the final outcome, rendered from ClickHouse.
- Policy decision view: for any denied publish, the Cedar decision, determining policy ids, and a one-click simulate with modified parameters.
- Token and credential state: validity, scopes, last used, revocation status. The token value is never retrievable.
- Agent and sidecar state: online status, adapter type and version, last heartbeat, inbox depth, recent errors with codes.
- Tenant health: plan, limits versus usage, stream fill ratio, DLQ depth, KMS status, open incidents.
- Payload access: disabled by default. A tenant admin enables a consent window (1 to 72 hours) from the customer console; within it, support can view payloads for a named ticket. Every view is recorded in the customer's audit chain with staff id, ticket id, and message id.
- Actions: trigger redelivery of a dead-lettered message, pause and resume a tenant stream, extend a TTL, revoke a token on request, all audited.
The support console reads through the control-plane API with a staff identity; it has no direct database access.
15. Status page
Public at status.agentbus.exchange, hosted by a third-party status provider so it stays up when AgentBus is down. Components: gateway (per region), control plane, console, marketplace (post-MVP), sidecar update service. Incidents are posted within 15 minutes of Sev1 detection. Tenants can subscribe by email and webhook. Historical uptime is published against the SLOs in section 11.
16. Runbooks
Each runbook lives in ops/runbooks/ with symptoms, diagnosis queries, actions, and rollback.
- NATS consumer stuck: symptoms are rising
agentbus_inbox_pendingfor one agent with an online sidecar; check consumer info forack_pendingat the limit; inspect sidecar logs for ack failures; action is to reset the consumer or instruct the sidecar to reconnect; verify withagentbus trace. - DLQ growth: identify the tenant and reason label; common causes are offline recipients past TTL, schema rejections after a sidecar upgrade, or a policy change; notify the tenant; redeliver after the fix.
- ClickHouse ingest lag: check worker replica count and NATS pending; scale workers; if ClickHouse is the bottleneck, check merge backlog and parts; the 48-hour
sys.auditretention bounds data loss risk. - Token leak response: revoke the token, confirm revocation propagation in the KV, list agents affected, notify the owner, review audit events from the token since the suspected leak time, rotate agent keys if the sidecar host is suspect.
- KMS unavailable: confirm provider status; gateway serves from the DEK cache for 10 minutes then fails closed; communicate on the status page; for BYOK, contact the tenant.
- Tenant offboarding with crypto-shredding: deliver the audit and message export, confirm receipt, revoke all tokens, delete streams and consumers, schedule KEK deletion with a 7-day pending window, record completion in the retained anchor log.
- Gateway rollout gone wrong: Helm rollback, verify WebSocket reconnect storm is within limits, check delivery latency returns to baseline.
- Certificate issuance failure for a custom domain: check the
askendpoint logs, the tenant's DNS TXT verification, and ACME rate limits.
17. Cost estimate for the MVP hosted footprint
Rough monthly ranges for a single production region at the capacity assumptions in section 12, excluding staff. Prices vary by provider; the ranges assume a mid-tier cloud with managed Postgres and ClickHouse Cloud.
| Item | Monthly range (USD) | Assumptions |
|---|---|---|
| Kubernetes nodes (core, 6 to 12 vCPU-heavy nodes) | 900 to 1,800 | Gateway, control plane, workers, console, OTel |
| NATS nodes (3 x 4 vCPU, SSD 500 GiB) | 400 to 700 | Dedicated node pool |
| Managed Postgres 19 (primary 8 vCPU, 2 replicas) | 600 to 1,200 | PITR included |
| ClickHouse Cloud (development to small production tier) | 300 to 900 | About 100 GiB hot, S3 cold |
| Object storage (2 to 5 TiB with egress) | 100 to 300 | Blobs, archives, backups |
| Load balancer, egress, DNS | 150 to 400 | |
| KMS, secrets, logging retention | 100 to 250 | |
| Observability (managed Grafana stack or self-hosted nodes) | 200 to 500 | |
| Clerk, Stripe, status page, email | 100 to 400 | Usage-based |
| Total | 2,850 to 6,450 |
Staging adds roughly 40% of production; dev adds about 15%. The self-hosted edition's minimum footprint is one 8 vCPU, 32 GiB node.
18. Self-hosting guide outline
- Requirements: Docker Compose v2 on one node with 8 vCPU, 32 GiB RAM, 500 GiB SSD for evaluation; Kubernetes 1.30 or newer with the Helm chart for production; a public DNS name and either ACME reachability or operator-supplied certificates; an OIDC provider (Keycloak or Zitadel; Clerk also works if the operator wants hosted identity); an S3-compatible store; a KMS or Vault.
- Compose services:
caddy,gateway,control-plane,console,support-console,workers,nats(single node with JetStream for evaluation, 3 nodes recommended),postgres,clickhouse,minio,otel-collector,grafana, optionalkeycloak. - OIDC configuration: set
AGENTBUS_OIDC_ISSUER,AGENTBUS_OIDC_CLIENT_ID,AGENTBUS_OIDC_CLIENT_SECRET,AGENTBUS_OIDC_JWKS_URL; map an IdP group claim to the tenantownerrole for bootstrap; example realm exports for Keycloak and Zitadel are shipped indeploy/oidc/. - Keys: set
AGENTBUS_KMS_PROVIDERtoaws,gcp,vault, orfile;fileis evaluation only and the console shows a warning. - First run:
agentbus-admin bootstrapcreates the first tenant and owner from the OIDC login, prints the console URL, and runs the samedoctorchecks the sidecar uses. - Upgrades: follow the hosted release train;
docker compose pull && docker compose up -dorhelm upgrade; migrations run automatically on start with the expand-and-contract guarantee; apre-upgrade-checkcommand validates compatibility. - Backups: the compose file includes a backup sidecar for Postgres and ClickHouse to the configured S3 bucket; NATS snapshots are scheduled with a cron container.
- License: Starter-equivalent features are free to self-host; Business and Enterprise features (SSO group mapping, Cedar custom policies, BYOK, custom domains, dedicated NATS accounts, marketplace participation) require a license key issued from the hosted console; the key is verified offline by signature and does not phone home unless telemetry is enabled.
- Telemetry: off by default; the documented aggregate counters can be enabled with
AGENTBUS_TELEMETRY=on. - Support: self-hosted customers on a paid license get the same support console in their installation and ticket-based support from AgentBus with log bundle export via
agentbus-admin support-bundle.
Implemented operations (as of 2026-10-11)
What exists in the repo and on the hosted server, versus the design above:
| Area | State |
|---|---|
| Backups | deploy/server/backup.sh nightly 03:00 UTC (systemd timer agentbus-backup.timer): pg_dump + ClickHouse native BACKUP + JetStream snapshot/copy + /etc/agentbus, one age-encrypted .tar.zst.age under /var/backups/agentbus (14-day rotation), optional S3 upload via /etc/agentbus/backup.env (deploy/server/env/backup.env.example). Recipient key /etc/agentbus/backup.age.pub; the private key must be kept off-host. |
| Restore | deploy/server/restore.sh <archive> --drill restores Postgres into agentbus_restore_drill, verifies, drops. First drill passed on 2026-10-10 (tenants=2, 24 tasks, 165 receipts). --postgres, --clickhouse, --secrets modes documented in the script. |
| Status page | deploy/status/status-probe.sh every minute → /var/www/status/{index.html,status.json,history.jsonl} served by status.agentbus.exchange (nginx block installed; DNS A record pending; 24 h uptime figure). Checks: gateway, console, ingest, postgres, clickhouse, nats, disk <90 %, TLS >7 d. |
| Observability | deploy/observability/docker-compose.yml: Prometheus + Alertmanager + blackbox + node + NATS exporters, Loki + promtail, Grafana with three provisioned dashboards (Gateway, Ingest & platform, Business via the ClickHouse datasource) and ten alert rules. The gateway exports no /metrics yet; required metric names listed in deploy/observability/README.md. |
| Self-host | deploy/selfhost/: Compose edition (Caddy auto-TLS, gateway, console, ingest, NATS, Postgres, ClickHouse, RustFS, Keycloak with an imported realm whose mappers emit org_id/org_slug/org_role/email/azp like Clerk), distroless images under images/, README with first-run, identity, upgrade and backup. Verified locally: readyz, well-known, Keycloak discovery through Caddy, conformance 26/26 PASS. Console web UI still needs Clerk keys (documented limitation). |