AI AutomationCase study
Self-Hosted Buzz: A Private Team Workspace Where AI Agents Are Teammates
I took Block's open-source Buzz platform from a GitHub repo to two live production servers for a cybersecurity company: private team chat where AI agents sit in channels, answer @mentions and do real work, all on infrastructure the company owns.
- Role
- Solo engineer: infra, agents, debugging
- Client
- Cybersecurity company (confidential)
- Year
- 2026
- Status
- Live in production
- Google Cloud
- AWS EC2
- Docker Compose
- Kubernetes (k3s)
- nginx
- Let's Encrypt
- PostgreSQL
- Redis
- MinIO
- Rust
- Goose
- Ollama Cloud
- Nostr
- systemd
The goal
The company wanted a Slack-style workspace where AI agents are real members of the team. You @mention an agent in a channel and it answers, checks a server, opens a GitHub issue or reviews a document. The non-negotiable requirement was ownership: every message, file and agent had to run on servers the company controls, not on someone else's cloud.
Buzz by Block is an open-source platform built for exactly this. But an open-source repo is not a working product. Over about six weeks I turned it into one: two production servers, more than fifteen agents, a way for non-developers to create their own agents, and documentation so the team can run it without me.
- production servers (GCP + AWS)
- 2
- AI agents deployed
- 15+
- silent failures traced and documented
- 20+
- agent idle limit, patched in Rust
- 2h → 30d
How Buzz works
Understanding the moving parts was the first job, because most of the later bugs lived in the gaps between them:
- The relay is the server. It speaks Nostr over WebSockets and stores everything in PostgreSQL, Redis and MinIO.
- Buzz Desktop (Mac and Windows) is the actual app. Channels, agents and member management only exist there. The relay's web page is just a Git repository browser, which surprised everyone at first.
- An agent is a small bridge process (
buzz-acp) that connects outbound to the relay, listens for @mentions and hands each one to an AI harness like Goose, which talks to the LLM and posts the reply. Because the connection is outbound, an agent can run on any server without opening a single port.
Buzz Desktop app
│ @mention in a channel
▼
nginx (HTTPS, WebSocket) ──▶ Buzz relay ──▶ PostgreSQL · Redis · MinIO
▲
│ outbound connection from each agent
│
buzz-acp ──▶ Goose harness ──▶ LLM (Ollama Cloud)
(systemd service, Docker container or Kubernetes pod)
Phase 1: Production server
The first server ran on Google Cloud. I deployed the relay stack with Docker Compose and put nginx in front as the only public entry point, with WebSocket upgrades configured (the relay protocol depends on them). After DNS went live I issued a Let's Encrypt certificate with automatic renewal.
Two decisions here mattered later:
- Data lives in named Docker volumes, not inside containers, so restarting or upgrading the relay never touches the database or uploaded files.
- Secrets stay in a server-side env file, never in an image. The relay image is public, so baking keys into it would have leaked them permanently.
Then I logged in as the owner from the desktop app, created the first invite links for the team and handed the owner identity over to the client.
Phase 2: The first agents
The first five agents ran as systemd services on the same server, each in its own channel: an infra agent, a dev agent, a general helper, a support agent and a welcome agent. A sixth ran on a completely different server and joined over its outbound connection, with no inbound ports opened.
- The infra agent reports CPU, memory and disk for four servers through a small Python metrics service I wrote, bound to localhost only, plus a Docker health check on one production box.
- The dev agent is connected to GitHub: it can create and close issues and review pull requests.
- Every agent got a prompt written for the company's real context, plus a shared folder of project guides to answer from.
Midway through, I migrated every agent from OpenRouter to the Goose harness with Ollama Cloud, removing that dependency completely.
Getting the very first agent to say anything took fixing five silent bugs in a row. The agent looked healthy every time; it just never replied.
The system prompt was being ignored
- Symptom
- The agent ran but behaved like it had no instructions at all.
- Root cause
- The prompt file was set with an environment variable name the bridge does not read.
- Fix
- Found the variable the bridge actually reads in its source and switched every agent to it.
The agent had no tools
- Symptom
- It could think about a reply but had no way to post one.
- Root cause
- The tool server (MCP) that gives the agent its shell and messaging tools was never configured.
- Fix
- Configured the MCP server for every agent.
Replies were generated but never posted
- Symptom
- Logs showed full answers being written, but nothing appeared in the channel.
- Root cause
- The bridge does not publish the agent's text automatically. The agent has to post it, and the setting that forces it to was off by default.
- Fix
- Turned on the require-reply setting for every agent. This was the most important fix of the whole first week.
Agents showed up as a long key instead of a name
- Symptom
- In channels, the bot appeared as a random-looking public key.
- Root cause
- No profile had ever been published for the agent's identity.
- Fix
- Published a profile (name and bio) signed with the agent's own key. It updates live, no restart needed.
The agent could not see its own channel
- Symptom
- The agent was a member in the database but never subscribed to the channel.
- Root cause
- The relay serves membership from cached events, and those were stale. Deleting only one of them reconciled nothing.
- Fix
- Deleted both cached events for the channel, then ran the relay's reconcile command so it rebuilt them from the database.
Phase 3: Agents from the desktop app
Server-side agents worked, but every new one needed SSH, database queries and a systemd file. The goal was for anyone on the team to create an agent from the desktop app by filling in a form. That meant running a Kubernetes cluster (k3s) for agents and getting the app's Run on Kubernetes option to work. It took a full day of failures, each documented so it never happens again.
Installing Kubernetes broke HTTPS for every user
- Symptom
- Right after installing k3s, every desktop client got stuck on Reconnecting. The relay itself was healthy.
- Root cause
- k3s ships with Traefik, which quietly took over ports 80 and 443 from nginx and served its own self-signed certificate.
- Fix
- Removed Traefik and disabled it permanently in the k3s config, so nginx and the real certificate own HTTPS again.
Remote Kubernetes access timed out
- Symptom
- kubectl from a laptop timed out on the Kubernetes API port, even with a firewall rule in place.
- Root cause
- Three stacked problems: the firewall rule targeted a network tag the VM did not have, its destination filter used the external IP while the cloud matches the internal one, and the API certificate did not include the public IP.
- Fix
- Fixed all three in order: tagged the VM, removed the wrong destination filter, and regenerated the API certificate with the public IP included.
The app said Deployed, but no pod existed
- Symptom
- The agent showed as deployed in the app, yet nothing was running in the cluster.
- Root cause
- The deploy binary silently does nothing when any part of the agent config is invalid.
- Fix
- Called the deploy binary directly with hand-written JSON to learn exactly what payload it expects, then fixed the config.
Agents replied only with emoji, never text
- Symptom
- The pod showed Running and reacted to @mentions with an emoji, but never sent text. No error anywhere in the UI.
- Root cause
- The deploy binary always starts the Goose harness, whatever you pick in the UI, and the official agent image does not contain Goose. Overriding it was impossible: the app rejects that variable as reserved.
- Fix
- Reported it to the Buzz maintainers (issue #6473) and built my own image with Goose included.
Switching harness brought new errors
- Symptom
- Trying the other harness instead crashed on start with unsupported provider, then unsupported API mode.
- Root cause
- The two harnesses use completely different provider names and settings. One of the variables I set was not the base URL at all, it selected the API mode.
- Fix
- Mapped the correct provider and variables for each harness, which also proved the harness choice was being ignored for pods.
My own image still could not find Goose
- Symptom
- The custom image worked locally, but in the cluster it failed with No such file or directory.
- Root cause
- Goose was linked from the agent's home folder, and the deploy binary mounts an empty volume over that folder when the pod starts, so the link pointed at nothing.
- Fix
- Copied the binary into a system path instead of linking it, rebuilt and republished the image. Agents finally replied with text from the desktop app.
Phase 4: Keeping agents alive
Agents went offline every 2 hours and never came back
- Symptom
- Agents created from the desktop app disappeared after a couple of hours of silence and stayed offline until someone clicked Start again. Server-side agents never did this.
- Root cause
- Reading the provider source: every pod gets a default 2-hour inactivity limit, the restart policy is hard-coded to Never, and the provider is a one-shot command with no background process watching the cluster. Setting the limit to 0 is rejected by this version.
- Fix
- Changed the default to 30 days in the provider's Rust source, rebuilt it, ran its full test suite (157 + 4 tests, all passing) and installed it. The app now pre-fills the safe value for every new agent.
I was explicit about what this fix does not cover: existing agents keep the value they were created with, and a pod that actually crashes still stays down, because there is still no watchdog. Both are written into the team's runbook.
A desktop update made the Run on Kubernetes option vanish
- Symptom
- After updating the app, the whole Run on section disappeared from the Create Agent dialog.
- Root cause
- Buzz moved to a plugin system: the app now scans the machine for a separate provider binary and hides the section if it finds none. This was only clear from the source code and design docs.
- Fix
- Built the provider binary from source, then checked a prebuilt copy into the team repo so other admins get the option back without installing a Rust toolchain.
Phase 5: A second server, done right
The client then needed a separate workspace for another team, so I built a second production server on AWS, following the playbook from the first one. It was much faster, and I used the chance to fix what I would have done differently.
The desktop app could not create invite links
- Symptom
- Couldn't create invite link, every time. Server logs showed zero invite requests had ever arrived.
- Root cause
- The relay's CORS settings allowed the website's origin but not the desktop app's own origins, so the app blocked its own request before sending it. The fix from the first server lived in that server's config and did not carry over.
- Fix
- Proved it with a CORS preflight request, added the Mac and Windows app origins, and made it a day-one checklist item for every new server.
- Least-privilege cluster access. On the first server, admins got a full cluster-admin credential just to deploy agents. Here I created a dedicated service account that can only manage pods and secrets in its own namespaces, and verified everything else (nodes, other namespaces, RBAC) is denied.
- No more local builds. Agents use the published image from a public registry, so any machine can pull it.
- Onboarding written for non-developers. A Mac and Windows install guide that matches what the current app actually shows, including traps like the button that looks right for owners but signs you into Block's hosted service instead, and the fact that display names are case-sensitive, so typing your name differently creates a duplicate user.
Phase 6: Agents at scale
The second server needed eight board-level agents (finance, HR, strategy, operations, market research and more), each with its own source documents. Creating them by clicking through the app eight times was not an option, so I built a headless pipeline that does it from the server in seven steps: generate an identity, register it on the relay, add it to a channel, publish its agent records, set its profile, ship its prompt and documents, and start its container.
Agents created on the server could not be @mentioned
- Symptom
- Could not authorize a mentioned agent, even though the agent was online and in the channel.
- Root cause
- An agent created outside the desktop app is missing two owner-signed records the app checks before it allows a mention. Both are required; either one alone is not enough.
- Fix
- Wrote small Rust tools to generate agent identities and publish both records, and made them part of the pipeline.
The runbook also covers the smaller traps I hit along the way:
- A missing document mount makes an agent politely refuse to cite figures instead of making them up. Correct behaviour, but it means the mount was forgotten.
- An agent occasionally freezes mid-conversation during a tool call. A container restart fixes it every time so far.
- Server-created agents show their owner as unavailable in the app. Cosmetic, and fixable with one database update.
- Removing an agent can leave it visible because the relay serves a cached member list, and the database's soft-delete columns are not read anywhere in the code. A real removal needs the proper remove command.
The agent that builds agents
Seven precise steps, where a single typo silently breaks an agent, is exactly the kind of work an agent should do. So I turned the pipeline into one.
The team now writes something like "create an agent called support-agent and add it to #help", in English, Hindi or Hinglish, and the meta-agent generates the identity, registers it, publishes the records, writes its instructions and starts it. It can also delete agents, move them between channels and update their files.
Because it holds admin-level power over every other agent, it lives in its own dedicated channel away from everyday conversations, and its rollout was gated on an end-to-end test. Its prompt also refuses to create channels: doing that from the command line once produced two channels with the same name, the agent joined the one nobody could see, and @mentions silently went nowhere.
Every agent silently stopped answering at once
- Symptom
- Found while deploying the meta-agent: agents accepted @mentions, started a turn and never replied. No error in the app or the server.
- Root cause
- The LLM provider had retired the model every agent was configured to use, so every request was failing with 410 Gone behind the scenes.
- Fix
- Pulled the provider's live model list, moved the new agent to a current model and documented the same one-line fix for every other agent.
My debugging toolkit
With failures this quiet, I worked layer by layer: is the server up, did the request arrive, did something block it, what is the agent really doing. These are the checks I wrote into the team's runbook (real hosts and names replaced with placeholders):
# 1. Is the relay reachable at all? Rules out DNS, TLS and a dead server.
curl -s https://<your-relay-domain>/ -H "Accept: application/nostr+json"
# 2. Did the request even reach the server?
docker logs <relay-container> --since 1h | grep -i invite
# 3. Is CORS blocking the desktop app before it sends anything?
curl -sI -X OPTIONS https://<your-relay-domain>/api/invites \
-H "Origin: tauri://localhost" \
-H "Access-Control-Request-Method: POST"
# 4. Which certificate is really being served on port 443?
openssl s_client -connect <your-relay-domain>:443 -servername <your-relay-domain> </dev/null \
| openssl x509 -noout -issuer
# 5. Can a laptop reach the Kubernetes API?
nc -zv <server-ip> 6443
# 6. What is the agent actually doing?
kubectl logs <agent-pod> -n <agent-namespace> --tail 50
Making it safe to hand over
-
Scoped cluster access. The provider's role allows exactly what its source code calls, nothing more:
# Only what the agent provider actually calls (names are placeholders) kind: ClusterRole apiVersion: rbac.authorization.k8s.io/v1 metadata: name: <agent-provider-role> rules: - apiGroups: [""] resources: ["namespaces"] verbs: ["get", "list", "create"] - apiGroups: [""] resources: ["pods", "secrets"] verbs: ["get", "list", "watch", "create", "delete"] -
Health checks. One command checks nginx, the relay, the database, Redis, storage, the cluster and certificate expiry, and exits non-zero on failure so it can feed an alert later.
-
Backups. A daily job dumps the database and archives uploaded files, keeping the last seven.
-
Documentation. Every step, bug and fix is written down twice: a full log for engineers and a click-by-click guide for everyone else.
What is still open
Being honest about this is part of handing a system over:
- Backups are stored on the same server. They protect against corruption and mistakes, not against losing the server. Copying them off-site is the next step.
- There is still no watchdog that restarts a crashed agent pod automatically.
- The welcome agent answers @mentions but does not yet greet new members on its own.
What I took away
- Read the source, not just the docs. The plugin system, the idle limit, the CORS origins and the mention authorization were only clear from the code.
- Silent failures need active checks. Health checks, log reviews and end-to-end tests caught what the UI never showed.
- Fix it for the next person. A patched default, a prebuilt binary and a written runbook beat a fix only I remember.
- Know when not to fork. Patching the desktop app was possible. Not doing it was the better engineering decision.