What a private homelab AI build costs and runs
How much memory you need, measured speed under load and where the $2,500 setup fee goes.
The Homelab AI setup from TheFireDev costs $2,500 for the work. Hardware is separate. If you already own a supported machine, you can start with the free Harness and pay no new software license fee for its core.
Separate the costs before comparing quotes
- Hardware: the computer, memory, storage and networking in the agreed design. TheFireDev does not resell hardware, so get a vendor quote for the exact configuration.
- Setup work: $2,500 for the Homelab AI package, including sizing, installation, task tests and handoff.
- Operation: electricity, backups, replacement parts and maintenance time. Cloud API or client subscriptions are separate.
For equipment you already own, the hardware cost can be $0, so the total is the $2,500 package fee plus tax. A new build costs the vendor quote for the exact configuration plus that fee. The fee includes no hardware.
Compare delivered prices for the same configuration. RAM, storage, country, tax and stock change the bill, and a used GPU also needs a compatible computer, cooling and a power supply. Estimate electricity from a wall meter: average kilowatts × hours on × your rate per kWh, idle time included. The benchmarks below record speed, not wall power.
The Harness core is free, the Pro skills pack is $99 and the one-machine install is $499. None is required to serve a local model.
Hardware classes and what they can hold
Memory capacity decides whether the weights and working state fit. Bandwidth and the runtime decide how fast tokens arrive. A mixture-of-experts model keeps every expert in memory even when only some are active.
An existing laptop or small desktop
The Harness catalog starts at 8 GB RAM with Qwen3-1.7B Q4_K_M, and the 4B model needs 12 GB, with 16 GB recommended. These small models suit short drafts and reference tasks.
The lab laptop has 30 GiB of RAM and an Intel integrated GPU. It ran Qwen3.5-2B Q4_K_M using Vulkan at 24.6 output tokens per second on October 5.
A discrete-GPU workstation
A 16 GB card can provide quick small-model replies, but competing applications take some of its memory. The RTX 4080 in the lab also served speech and other model workloads. A 4B Q5 model allocation failed under that shared load; the tested fast tier used a 2B Q4 model.
An October 6 sizing calculation estimated Qwen3-8B at Q4_K_M with 32,768 tokens of context at 10.05 GB, and Qwen3-30B-A3B at the same settings at 21.65 GB, which does not fit on one 16 GB RTX 4080. These are estimates, not benchmarks.
CPU offload can run a model that does not fit in GPU memory, at a speed penalty. Check quantization, context length, simultaneous users and headroom before buying a card.
Apple unified memory
Apple Silicon shares memory between CPU and GPU. The 24 GiB M4 Pro Mac mini in the lab served Qwen3.5-9B Q4_K_M with Metal at 8.9 tokens per second on October 5 while also running a heavy build. A fresh 1,600-token prompt took 14.45 seconds to the first token.
A higher-memory Mac can fit larger weights. Match the purchase to the runtime and model you will use; the M4 Pro result does not transfer to other Macs.
One or two GB10-class machines
The lab uses an NVIDIA DGX Spark and an ASUS Ascent GX10, two GB10-class machines with 128 GB unified memory each. Their 200 Gb/s interconnect supports a tensor-parallel model spanning both machines. Two memory pools are not automatically one larger GPU: the model runtime and communication backend must support the arrangement.
The large-model tests used a locally converted NVFP4 GLM-5.3-Flash uncensored model on vLLM with tensor parallelism across both nodes. An October 6 metadata check reported the served ID glm-5.3-flash-uncensored and a maximum context of 262,144 tokens. That maximum is not a promise that many users can each hold a full window at once.
A GB10 pair buys capacity for that configuration and says nothing about a single GB10 or a different quantization. Check model licensing, conversion support and task quality before reproducing the build.
Measured speed, with dates and workload
Output tokens per second measure generation after prompt processing. First-token latency also includes cold prompt processing, network time and queueing. Compare the same model, quantization, context and load when evaluating hardware.
| Configuration | Output tok/s | Date and conditions |
|---|---|---|
| Two GB10s, GLM-5.3-Flash NVFP4, prose | 38.4 | Oct 2, idle-gated perf-loop |
| Same pair and model, code | 68.3 | Oct 2, idle-gated perf-loop |
| Same pair, six simultaneous streams | 91 aggregate | Oct 2, combined rate, not per user |
| Same pair, one benchmark stream alongside live traffic | 6.2 | Oct 5, loaded; decode first token 1.689 s |
| RTX 4080, Qwen3.5-2B Q4_K_M, CUDA | 274.4 | Oct 5, concurrency 1; decode first token 15 ms |
| Intel laptop iGPU, Qwen3.5-2B Q4_K_M, Vulkan | 24.6 | Oct 5, concurrency 1; decode first token 138 ms |
| M4 Pro 24 GiB, Qwen3.5-9B Q4_K_M, Metal | 8.9 | Oct 5, shared build load; decode first token 413 ms |
| 16-thread CPU, Qwen3-4B-Instruct-2507 Q4_K_M | 11.97 and 12.29 | Oct 6, two direct llama.cpp runs; 65,536-token configured context |
The October 2 cluster rates are historical figures from TheFireDev cluster records. The October 6 engine had live traffic, so no new idle benchmark was run.
The October 5 runs used temperature zero, thinking off and a roughly 100-word prompt capped at 128 output tokens. Rates are medians of two to five runs after warmup, taken while the cluster served other requests. The October 6 CPU runs used a 23-token prompt capped at 160 output tokens.
Read the dated measurement record and method. These are observations from hardware operated by TheFireDev.
Choose with a task test, not a token-rate chart
- Collect representative tasks with expected answers. Keep private data on your own machine, and include the failures you most want to avoid.
- Run the candidate model at the context length you actually need. Measure first-token time, output rate and answer quality separately.
- Repeat with your expected number of simultaneous users. Watch the queue and memory headroom.
- Test restart after power loss and restoration of the configuration before making it part of daily work.
For a small local agent, install the free Harness first, then test your own workload before you buy hardware.
What the $2,500 package includes
- Sizing advice for hardware you own or plan to buy (I do not resell hardware)
- A runtime installed to fit the machine: llama.cpp, vLLM or Ollama
- One or two open-weight models tested against your own tasks, with the numbers
- An OpenAI-compatible endpoint on your private network, your agent harness pointed at it
- Monitoring, restart after a power loss and a backup of the configuration
- A written runbook and 30 days of email support
Required: A computer with a supported GPU or unified memory, and admin access to it. The published timeline is 1 to 2 weeks. Hardware, cloud API subscriptions, ongoing hosting and a custom business workflow are separate.
After payment, the confirmation email asks for two or three suitable times. A time is confirmed within two business days. If your setup needs a multi-user rollout or a specific compliance requirement, send that scope before buying.
Sources and dates
- Dated TheFireDev lab measurements: October 2, 5 and 6, 2026. Private network addresses and account details are excluded.
- Harness 1.0.0 model catalog and memory and install requirements, checked October 6.
- Memory estimates from a sizing calculator, run October 6, 2026 and labeled as estimates, not measurements.
- Current Homelab AI package scope, checked October 6, 2026.
Questions
How much does a private homelab AI build cost?
The setup work is $2,500 and covers sizing, installation, task tests and handoff. Hardware is separate: $0 if you already own a supported machine, otherwise the vendor quote for the exact configuration. Electricity and backups are the ongoing costs.
How much memory does a local AI model need?
Memory decides whether the weights and the context fit. In an October 6, 2026 sizing calculation, Qwen3-8B at Q4_K_M with 32,768 tokens of context needed about 10 GB and Qwen3-30B-A3B needed about 22 GB. Leave headroom for other programs.
How fast do local models run?
In dated lab runs on October 5, 2026, an RTX 4080 ran Qwen3.5-2B Q4_K_M at 274.4 output tokens per second and a 24 GiB M4 Pro ran Qwen3.5-9B Q4_K_M at 8.9 under build load. Speed depends on the model, quantization, context length and load.
Do I need a cloud subscription to run models at home?
No. Open-weight models served from your own hardware carry no per-token bill. You pay for electricity, backups and the time it takes to maintain the machine.