What a private homelab AI build costs and runs
How much memory you need, measured speed under load and where the $2,500 setup fee goes.
The Homelab AI setup from TheFireDev costs $2,500 for the work. Hardware is separate. If you already own a supported machine, you can start with the free Harness and pay no new software license fee for its core.
Choose a model for the task, then size the memory and measure a reply on the proposed hardware. Buying more capacity can let a larger model fit while leaving short replies slower than they were on a smaller GPU.
Separate the costs before comparing quotes
- Hardware: the computer, memory, storage and any networking required by the agreed design. TheFireDev does not resell hardware. Obtain a current vendor quote for the exact configuration.
- Setup work: $2,500 for the Homelab AI package, including sizing, installation, task tests and handoff.
- Operation: electricity, backups, replacement parts and time spent maintaining the machine. Cloud API or client subscriptions remain separate if you choose to use them.
For equipment you already own, the new hardware purchase can be $0. That makes the package fee $2,500 before taxes or any separately quoted work. A new build costs the vendor hardware quote plus that fee. There is no hardware included in the advertised $2,500.
Do not compare a launch MSRP with the current delivered price of an available machine. RAM capacity, storage, country, tax and stock can change the bill. A used GPU quote also needs the rest of the computer, cooling and a compatible power supply.
Estimate recurring electricity from a wall meter: average measured kilowatts × hours powered on × your utility rate per kWh. Include idle time. We have not measured wall power for these benchmarks, so they do not establish a monthly electricity cost or cost per token.
Software price and setup price also differ. The Harness core is free; its Pro skills pack is $99 once and its one-machine install is $499. Connectors Pro is $19 a month for public-data tools. Neither is required to serve a local model.
Hardware classes and what they can hold
Memory capacity controls whether the weights and working state fit. Memory bandwidth and the runtime affect how quickly tokens arrive. For a mixture-of-experts model, all experts still occupy memory even when only a subset is active for each token.
An existing laptop or small desktop
The Harness catalog starts at 8 GB RAM with Qwen3-1.7B Q4_K_M. Its 4B model threshold is 12 GB, with 16 GB recommended for the machine. Those are small models for short drafts and reference tasks. Test tool use and factual accuracy before trusting agent work.
The lab laptop has 30 GiB of RAM and an Intel integrated GPU. It ran Qwen3.5-2B Q4_K_M using Vulkan at 24.6 output tokens per second on October 5. That measures a 2B model on that configuration, not an 8B model or every laptop.
A discrete-GPU workstation
A 16 GB card can provide quick small-model replies, but competing applications take some of its memory. The RTX 4080 in the lab also served speech and other model workloads. A 4B Q5 model allocation failed under that shared load; the tested fast tier used a 2B Q4 model.
An October 6 sizing-tool call estimated Qwen3-8B at Q4_K_M and 32,768-token context needed 10.05 GB, including weights, KV cache and runtime overhead. Qwen3-30B-A3B at the same settings needed 21.65 GB and did not fit entirely on one 16 GB RTX 4080. These are estimates from models_fit_check, not an executed benchmark or a guarantee of allocation.
CPU offload may let a model run that does not fit in GPU memory, with a performance penalty. Advertised GPU memory alone is not a reason to buy a card. Check the exact quantization, context length, number of simultaneous users and remaining headroom.
Apple unified memory
Apple Silicon shares memory between CPU and GPU. The 24 GiB M4 Pro Mac mini in the lab served Qwen3.5-9B Q4_K_M with Metal. It produced 8.9 tokens per second while also handling a heavy build workload on October 5. A fresh 1,600-token prompt took 14.45 seconds before the first token. This was a busy build machine, not an idle Mac performance test.
A higher-memory Mac can fit larger weights, but the purchase should match the runtime and model you will use. We have no measured result here for the current Mac mini generation or a 128 GB Mac. Do not extrapolate the M4 Pro result to those machines.
One or two GB10-class machines
The lab uses an NVIDIA DGX Spark and an ASUS Ascent GX10, two GB10-class machines with 128 GB unified memory each. Their 200 Gb/s interconnect supports a tensor-parallel model spanning both machines. Two memory pools are not automatically one larger GPU: the model runtime and communication backend must support the arrangement.
The historical large-model tests used a locally converted NVFP4 GLM-5.3-Flash uncensored model on vLLM with tensor parallelism across both nodes. An October 6 read-only model metadata check reported the served ID glm-5.3-flash-uncensored and a maximum context of 262,144 tokens. That maximum is not a promise that many users can each hold a full window simultaneously.
A GB10 pair buys capacity for that configuration. It does not establish the speed of a single GB10 or a different quantization. Model licensing, conversion support and task quality should be checked before reproducing the build.
Measured speed, with dates and workload
Output tokens per second measure generation after prompt processing. First-token latency also includes cold prompt processing, network time and queueing. Compare the same model, quantization, context and load when evaluating hardware.
| Configuration | Output tok/s | Date and conditions |
|---|---|---|
| Two GB10s, GLM-5.3-Flash NVFP4, prose | 38.4 | Oct 2, idle-gated perf-loop |
| Same pair and model, code | 68.3 | Oct 2, idle-gated perf-loop |
| Same pair, six simultaneous streams | 91 aggregate | Oct 2, combined rate, not per user |
| Same pair, one benchmark stream alongside live traffic | 6.2 | Oct 5, loaded; decode first token 1.689 s |
| RTX 4080, Qwen3.5-2B Q4_K_M, CUDA | 274.4 | Oct 5, concurrency 1; decode first token 15 ms |
| Intel laptop iGPU, Qwen3.5-2B Q4_K_M, Vulkan | 24.6 | Oct 5, concurrency 1; decode first token 138 ms |
| M4 Pro 24 GiB, Qwen3.5-9B Q4_K_M, Metal | 8.9 | Oct 5, shared build load; decode first token 413 ms |
| 16-thread CPU, Qwen3-4B-Instruct-2507 Q4_K_M | 11.97 and 12.29 | Oct 6, two direct llama.cpp runs; 65,536-token configured context |
The October 2 cluster rates are historical reference figures from TheFireDev cluster-verification and inference-operations records. The October 6 engine had active traffic, so no new idle benchmark was run. The small 2B GPU result and the large GLM result answer different questions about model capacity and output quality.
The October 5 measurements used temperature zero, thinking off and a roughly 100-word decode prompt capped at 128 output tokens. Rates are medians of two to five runs after warmup. The cluster was already serving other requests. Its queue changed throughout the test window; low per-stream rates under load cannot be attributed to hardware alone.
The October 6 CPU runs used a 23-token prompt asking for about 100 words on local-model privacy, capped at 160 output tokens. The server returned 125 and 114 completion tokens. Two short runs do not predict full agent-session latency.
Read the dated measurement record and method. These are observations from hardware operated by TheFireDev, not an independent product comparison or a service-level commitment.
Choose with a task test, not a token-rate chart
- Collect representative tasks with expected answers. Keep private data on your own machine, and include the failures you most want to avoid.
- Run the candidate model at the context length you actually need. Measure first-token time, output rate and answer quality separately.
- Repeat with your expected number of simultaneous users. Watch the queue and memory headroom.
- Test restart after power loss and restoration of the configuration before making it part of daily work.
For a small local agent, install the free Harness first. Use the model-sizing MCP tools to reject memory configurations that do not fit. Their speed ceilings are estimates; only a real run measures your task.
What the $2,500 package includes
- Sizing advice for hardware you own or plan to buy (I do not resell hardware)
- A runtime installed to fit the machine: llama.cpp, vLLM or Ollama
- One or two open-weight models tested against your own tasks, with the numbers
- An OpenAI-compatible endpoint on your private network, your agent harness pointed at it
- Monitoring, restart after a power loss and a backup of the configuration
- A written runbook and 30 days of email support
Required: A computer with a supported GPU or unified memory, and admin access to it. The published timeline is 1 to 2 weeks. Hardware, cloud API subscriptions, ongoing hosting and a custom business workflow are separate. The package does not promise that an open model equals a frontier model on hard tasks.
After payment, the confirmation email asks for two or three suitable times. A time is confirmed within two business days. If your setup needs a multi-user rollout or a specific compliance requirement, send that scope before buying.
Sources and dates
- Dated TheFireDev lab measurements: October 2, 5 and 6, 2026. Private network addresses and account details are excluded.
- Harness 1.0.0 model catalog and memory and install requirements, checked October 6.
- Connector memory estimator and GPU catalog, with sizing calls made October 6. Estimates are labeled separately from measurements.
- Current Homelab AI package scope, checked October 6, 2026. New hardware needs a current vendor quote; no unverified shopping price is quoted here.