last deploy · 2026-09-02

working note · confidence: medium · 2026-08-26 · [02]

Running a local model next to the rest of the rack

The useful question is not whether you can serve a model at home. It is whether the box still feels like a tool you own after the second week.

The rack, printed as ink on paper
Fig 01 The rack, printed as ink on paper

Most of the writeups start at the GPU and end at a screenshot of tokens per second. That is a bench, not a system. A homelab already has a job: DNS, backups, a git forge, the site you are reading. The model has to sit in that list without becoming the only thing the rack is for.

This is a living note. Serving aliases will move. I will update the note, not republish the essay. Updated when the work changes, not when the feed wants a post.

What I actually kept

One role per machine. The Halos serve with vLLM and the official weights. They do not also do CI. They do not also publish this site. Speech stays on the 4070. The small boxes plan. The image box can be off and the rest of the house still has to work.

The public site is this one. It stays off the gate on purpose. The rack stays gated. I do not need strangers hitting a model to prove the lab is real.

The interface I trust is still text. A file. A feed. A list that writes when something ships. I do not need the site to look like the model. I need it to look like I still know which machine is which.