All work

AI AutomationCase study

Self-Hosted Buzz: A Private Team Workspace Where AI Agents Are Teammates

I took Block's open-source Buzz platform from a GitHub repo to two live production servers for a cybersecurity company: private team chat where AI agents sit in channels, answer @mentions and do real work, all on infrastructure the company owns.

Role
Solo engineer: infra, agents, debugging
Client
Cybersecurity company (confidential)
Year
2026
Status
Live in production
  • Google Cloud
  • AWS EC2
  • Docker Compose
  • Kubernetes (k3s)
  • nginx
  • Let's Encrypt
  • PostgreSQL
  • Redis
  • MinIO
  • Rust
  • Goose
  • Ollama Cloud
  • Nostr
  • systemd
Video walkthrough

The goal

The company wanted a Slack-style workspace where AI agents are real members of the team. You @mention an agent in a channel and it answers, checks a server, opens a GitHub issue or reviews a document. The non-negotiable requirement was ownership: every message, file and agent had to run on servers the company controls, not on someone else's cloud.

Buzz by Block is an open-source platform built for exactly this. But an open-source repo is not a working product. Over about six weeks I turned it into one: two production servers, more than fifteen agents, a way for non-developers to create their own agents, and documentation so the team can run it without me.

production servers (GCP + AWS)
2
AI agents deployed
15+
silent failures traced and documented
20+
agent idle limit, patched in Rust
2h → 30d

How Buzz works

Understanding the moving parts was the first job, because most of the later bugs lived in the gaps between them:

Buzz Desktop app
      │  @mention in a channel
      ▼
nginx (HTTPS, WebSocket) ──▶ Buzz relay ──▶ PostgreSQL · Redis · MinIO
      ▲
      │  outbound connection from each agent
      │
buzz-acp ──▶ Goose harness ──▶ LLM (Ollama Cloud)
(systemd service, Docker container or Kubernetes pod)

Phase 1: Production server

The first server ran on Google Cloud. I deployed the relay stack with Docker Compose and put nginx in front as the only public entry point, with WebSocket upgrades configured (the relay protocol depends on them). After DNS went live I issued a Let's Encrypt certificate with automatic renewal.

Two decisions here mattered later:

Then I logged in as the owner from the desktop app, created the first invite links for the team and handed the owner identity over to the client.

Phase 2: The first agents

The first five agents ran as systemd services on the same server, each in its own channel: an infra agent, a dev agent, a general helper, a support agent and a welcome agent. A sixth ran on a completely different server and joined over its outbound connection, with no inbound ports opened.

Midway through, I migrated every agent from OpenRouter to the Goose harness with Ollama Cloud, removing that dependency completely.

Getting the very first agent to say anything took fixing five silent bugs in a row. The agent looked healthy every time; it just never replied.

The system prompt was being ignored

Symptom
The agent ran but behaved like it had no instructions at all.
Root cause
The prompt file was set with an environment variable name the bridge does not read.
Found the variable the bridge actually reads in its source and switched every agent to it.

The agent had no tools

Symptom
It could think about a reply but had no way to post one.
Root cause
The tool server (MCP) that gives the agent its shell and messaging tools was never configured.
Configured the MCP server for every agent.

Replies were generated but never posted

Symptom
Logs showed full answers being written, but nothing appeared in the channel.
Root cause
The bridge does not publish the agent's text automatically. The agent has to post it, and the setting that forces it to was off by default.
Turned on the require-reply setting for every agent. This was the most important fix of the whole first week.

Agents showed up as a long key instead of a name

Symptom
In channels, the bot appeared as a random-looking public key.
Root cause
No profile had ever been published for the agent's identity.
Published a profile (name and bio) signed with the agent's own key. It updates live, no restart needed.

The agent could not see its own channel

Symptom
The agent was a member in the database but never subscribed to the channel.
Root cause
The relay serves membership from cached events, and those were stale. Deleting only one of them reconciled nothing.
Deleted both cached events for the channel, then ran the relay's reconcile command so it rebuilt them from the database.

Phase 3: Agents from the desktop app

Server-side agents worked, but every new one needed SSH, database queries and a systemd file. The goal was for anyone on the team to create an agent from the desktop app by filling in a form. That meant running a Kubernetes cluster (k3s) for agents and getting the app's Run on Kubernetes option to work. It took a full day of failures, each documented so it never happens again.

Installing Kubernetes broke HTTPS for every user

Symptom
Right after installing k3s, every desktop client got stuck on Reconnecting. The relay itself was healthy.
Root cause
k3s ships with Traefik, which quietly took over ports 80 and 443 from nginx and served its own self-signed certificate.
Removed Traefik and disabled it permanently in the k3s config, so nginx and the real certificate own HTTPS again.

Remote Kubernetes access timed out

Symptom
kubectl from a laptop timed out on the Kubernetes API port, even with a firewall rule in place.
Root cause
Three stacked problems: the firewall rule targeted a network tag the VM did not have, its destination filter used the external IP while the cloud matches the internal one, and the API certificate did not include the public IP.
Fixed all three in order: tagged the VM, removed the wrong destination filter, and regenerated the API certificate with the public IP included.

The app said Deployed, but no pod existed

Symptom
The agent showed as deployed in the app, yet nothing was running in the cluster.
Root cause
The deploy binary silently does nothing when any part of the agent config is invalid.
Called the deploy binary directly with hand-written JSON to learn exactly what payload it expects, then fixed the config.

Agents replied only with emoji, never text

Symptom
The pod showed Running and reacted to @mentions with an emoji, but never sent text. No error anywhere in the UI.
Root cause
The deploy binary always starts the Goose harness, whatever you pick in the UI, and the official agent image does not contain Goose. Overriding it was impossible: the app rejects that variable as reserved.
Reported it to the Buzz maintainers (issue #6473) and built my own image with Goose included.

Switching harness brought new errors

Symptom
Trying the other harness instead crashed on start with unsupported provider, then unsupported API mode.
Root cause
The two harnesses use completely different provider names and settings. One of the variables I set was not the base URL at all, it selected the API mode.
Mapped the correct provider and variables for each harness, which also proved the harness choice was being ignored for pods.

My own image still could not find Goose

Symptom
The custom image worked locally, but in the cluster it failed with No such file or directory.
Root cause
Goose was linked from the agent's home folder, and the deploy binary mounts an empty volume over that folder when the pod starts, so the link pointed at nothing.
Copied the binary into a system path instead of linking it, rebuilt and republished the image. Agents finally replied with text from the desktop app.

Phase 4: Keeping agents alive

Agents went offline every 2 hours and never came back

Symptom
Agents created from the desktop app disappeared after a couple of hours of silence and stayed offline until someone clicked Start again. Server-side agents never did this.
Root cause
Reading the provider source: every pod gets a default 2-hour inactivity limit, the restart policy is hard-coded to Never, and the provider is a one-shot command with no background process watching the cluster. Setting the limit to 0 is rejected by this version.
Changed the default to 30 days in the provider's Rust source, rebuilt it, ran its full test suite (157 + 4 tests, all passing) and installed it. The app now pre-fills the safe value for every new agent.

I was explicit about what this fix does not cover: existing agents keep the value they were created with, and a pod that actually crashes still stays down, because there is still no watchdog. Both are written into the team's runbook.

A desktop update made the Run on Kubernetes option vanish

Symptom
After updating the app, the whole Run on section disappeared from the Create Agent dialog.
Root cause
Buzz moved to a plugin system: the app now scans the machine for a separate provider binary and hides the section if it finds none. This was only clear from the source code and design docs.
Built the provider binary from source, then checked a prebuilt copy into the team repo so other admins get the option back without installing a Rust toolchain.

Phase 5: A second server, done right

The client then needed a separate workspace for another team, so I built a second production server on AWS, following the playbook from the first one. It was much faster, and I used the chance to fix what I would have done differently.

The desktop app could not create invite links

Symptom
Couldn't create invite link, every time. Server logs showed zero invite requests had ever arrived.
Root cause
The relay's CORS settings allowed the website's origin but not the desktop app's own origins, so the app blocked its own request before sending it. The fix from the first server lived in that server's config and did not carry over.
Proved it with a CORS preflight request, added the Mac and Windows app origins, and made it a day-one checklist item for every new server.

Phase 6: Agents at scale

The second server needed eight board-level agents (finance, HR, strategy, operations, market research and more), each with its own source documents. Creating them by clicking through the app eight times was not an option, so I built a headless pipeline that does it from the server in seven steps: generate an identity, register it on the relay, add it to a channel, publish its agent records, set its profile, ship its prompt and documents, and start its container.

Agents created on the server could not be @mentioned

Symptom
Could not authorize a mentioned agent, even though the agent was online and in the channel.
Root cause
An agent created outside the desktop app is missing two owner-signed records the app checks before it allows a mention. Both are required; either one alone is not enough.
Wrote small Rust tools to generate agent identities and publish both records, and made them part of the pipeline.

The runbook also covers the smaller traps I hit along the way:

The agent that builds agents

Seven precise steps, where a single typo silently breaks an agent, is exactly the kind of work an agent should do. So I turned the pipeline into one.

The team now writes something like "create an agent called support-agent and add it to #help", in English, Hindi or Hinglish, and the meta-agent generates the identity, registers it, publishes the records, writes its instructions and starts it. It can also delete agents, move them between channels and update their files.

Because it holds admin-level power over every other agent, it lives in its own dedicated channel away from everyday conversations, and its rollout was gated on an end-to-end test. Its prompt also refuses to create channels: doing that from the command line once produced two channels with the same name, the agent joined the one nobody could see, and @mentions silently went nowhere.

Every agent silently stopped answering at once

Symptom
Found while deploying the meta-agent: agents accepted @mentions, started a turn and never replied. No error in the app or the server.
Root cause
The LLM provider had retired the model every agent was configured to use, so every request was failing with 410 Gone behind the scenes.
Pulled the provider's live model list, moved the new agent to a current model and documented the same one-line fix for every other agent.

My debugging toolkit

With failures this quiet, I worked layer by layer: is the server up, did the request arrive, did something block it, what is the agent really doing. These are the checks I wrote into the team's runbook (real hosts and names replaced with placeholders):

# 1. Is the relay reachable at all? Rules out DNS, TLS and a dead server.
curl -s https://<your-relay-domain>/ -H "Accept: application/nostr+json"

# 2. Did the request even reach the server?
docker logs <relay-container> --since 1h | grep -i invite

# 3. Is CORS blocking the desktop app before it sends anything?
curl -sI -X OPTIONS https://<your-relay-domain>/api/invites \
  -H "Origin: tauri://localhost" \
  -H "Access-Control-Request-Method: POST"

# 4. Which certificate is really being served on port 443?
openssl s_client -connect <your-relay-domain>:443 -servername <your-relay-domain> </dev/null \
  | openssl x509 -noout -issuer

# 5. Can a laptop reach the Kubernetes API?
nc -zv <server-ip> 6443

# 6. What is the agent actually doing?
kubectl logs <agent-pod> -n <agent-namespace> --tail 50

Making it safe to hand over

What is still open

Being honest about this is part of handing a system over:

What I took away