run a company of ai agents on a raspberry pi: cloud brain or fully self-hosted

run a company of ai agents on a raspberry pi: cloud brain or fully self-hosted


A Raspberry Pi is enough to run a whole company of AI agents. A persistent team that takes work off a queue, hands tasks between agents, and pings your phone only when a human has to decide.

It fits on a $100 board because 5dive is Linux wired together, not a platform on top of it: bash for the CLI every agent calls, systemd for supervision, SQLite for the shared task queue, journald for logs, separate Linux users for isolation. No broker, no vector database, no Kubernetes. It adds nothing to the board that isn’t already ARM-native, so it runs wherever the coding CLI runs, and coordinating a team barely touches the Pi.

That leaves one decision: where the brain, the model itself, actually runs. You get two ways, and the second is one no hosted agent platform will give you.

first, put 5dive on the Pi

Either option starts the same way. Flash Ubuntu Server 24.04 LTS ARM64 (official Pi 5 support), then:

curl -fsSL https://install.5dive.ai | sudo bash
sudo 5dive init
uname -m        # aarch64

option A: a hosted brain (works on any Pi)

The simplest setup, and the right one for serious coding. The Pi runs the company, a hosted model does the thinking. Point the agents at Anthropic, or at a BYO provider (DeepSeek, GLM, Qwen’s hosted tier, and others) through the same Claude Code path, on your own key, with no account with us.

sudo 5dive agent create coder --type=claude --auth-profile=work

No GPU, no extra hardware. The frontier-class model lives in the cloud. Everything else, the agents, the repos, the state, the coordination, lives on the board in your hand.

option B: go fully self-hosted (your own GPU)

This is the path nothing else gives you. 5dive drives the real Claude Code harness, and that harness will point at any Anthropic-compatible endpoint, so the model can be one you run yourself. Put it on a GPU box on your LAN, one you may already own, and nothing in the loop is rented: your agents, your hardware, your model, and a key that never leaves your network.

The model doing the work is Qwen3.8-27B, a 27B dense open-weight model, roughly 18GB, with a 256K context window. It fits a single 24GB-class GPU. On the GPU box:

ollama pull qwen3.8:27b
# then put TLS in front of it (caddy or nginx) so the Pi can reach it over https

Then, on the Pi, point the agents at it:

sudo 5dive agent create coder \
  --type=claude \
  --base-url=https://gpu.lan.example:8443 \
  --api-key=<your-token> \
  --auth-profile=selfhost \
  --model=qwen3.8:27b

Two things worth knowing:

--type=claude selects the Claude Code harness. --base-url and --model select the actual inference backend.

One security rule shapes the setup. 5dive accepts http:// only for the same box (localhost, 127.0.0.1, [::1]). Anything on the open network, including your own LAN, must be https://, because the API key rides the URL on every request and the LAN is not a place to trust it in the clear. That is why the GPU box gets a TLS front.

Want to see how the open models stack up against the frontier before you commit a GPU to one? We keep a live, independently-scored board at 5dive.ai/models, refreshed hourly, with benchmarks plus real token usage, including the open-weight sizes people actually self-host. Meta’s open-weight Muse Glimmer is a reasonable alternative to Qwen here, in the same 24GB-GPU class.

give the team a real job

Whichever brain you chose, the company runs the same way:

5dive task add "find one small bug, fix it, run the tests, write a summary" --assignee=lead

Close the SSH session. The heartbeat wakes an agent only when there’s work, the team hands tasks off through the queue, and Telegram pings your phone only if a human has to decide. Come back to a changed repo, passing tests, and a log of who did what.

how much fits on one board?

The question worth answering is how big a team a single Pi holds as the control plane, and how the 27B holds up as the brain, where tool-call reliability matters more than raw speed. We are not going to quote numbers we haven’t measured on the bench yet. When we run this exact setup on hardware, the results go in a follow-up, including where the open model falls short.

the shape of it

An always-on agent team on a board the size of a deck of cards, with a brain you can put wherever you want it: in the cloud when you want frontier muscle, or on your own GPU when you want to own the whole stack. Same company, same box, your choice of where the thinking happens.

The Raspberry Pi is the company. The brain is yours to place.

Hero photo: Raspberry Pi 5 by SimonWaldherr, CC BY 4.0, resized.