Install Claude Code, Codex and a local model on one machine
The installer, the separate coding clients, working commands and the limits of a small local model.
The free MeshVault Harness installs llama.cpp, a Qwen3 model, Hermes Agent and OMP. Claude Code and Codex are separate installs. You can keep cloud coding tools beside a local agent without sending every task to a paid API.
On October 6, 2026, the published installer finished in 6 minutes 31 seconds on a Linux workstation, with the 4B model selected and an isolated home directory.
Before you install
Use Linux or macOS as a normal user. Windows needs WSL2. Have curl, git and tar available. The install docs recommend 30 GB of free disk and 16 GB of RAM. Linux may need the system libraries libatomic and libgomp; the installer can request admin access for those.
Minimums are installer thresholds from the version 1.0.0 catalog, which allows for a 65,536-token context. They are not free memory on a busy machine.
| Model ID | Minimum RAM | Download |
|---|---|---|
qwen3-1.7b | 8 GB | 1.11 GB |
qwen3-4b-2507 | 12 GB | 2.50 GB |
qwen3-8b | 20 GB | 5.03 GB |
qwen3-30b-a3b-2507 | 40 GB | 18.56 GB |
Sizes are decimal GB from the catalog. More RAM lets a model fit but does not speed up a CPU reading a long prompt.
1. Read the installer, then run it
The entry script fetches the newest release tag and runs its setup script. Model and llama.cpp downloads are pinned and checked with SHA-256.
curl -fsSL https://raw.githubusercontent.com/thefiredev-cloud/meshvault-harness/main/install.sh -o meshvault-install.sh
less meshvault-install.sh
bash meshvault-install.sh --model qwen3-4b-2507The one-command version:
curl -fsSL https://raw.githubusercontent.com/thefiredev-cloud/meshvault-harness/main/install.sh | bash -s -- --model qwen3-4b-2507Leave out --model to let the installer choose by RAM. Linux defaults to CPU; add --gpu vulkan only with Vulkan drivers installed. macOS uses Metal. Pass a CUDA build with --llama-server PATH.
Output from the October 6 run, unrelated lines omitted:
MeshVault Harness 1.0.0
Downloading llama.cpp b11430 (linux-x64)
ok llama.cpp b11430 installed
ok model verified (sha256)
ok model server up on http://127.0.0.1:8484 (model: qwen3-4b-2507)
ok synced 11 skill(s) from free pack
ok everything checks out2. Check the model before asking an agent to work
export PATH="$HOME/.local/bin:$PATH"
meshvault status
meshvault doctor
meshvault ask "What is 17 times 3? Answer with the number only."
meshvault skills listA healthy doctor run checks a model answer, both agents and their configuration. Output after the install:
MeshVault Harness 1.0.0 on linux/x64, 124 GB RAM
ok local model answers
ok hermes installed (Hermes Agent v0.21.5+8160.g4ca6f9e (2026.9.24) · upstream 4ca6f9e9)
ok hermes points at the local model
ok omp installed (omp/18.6.3)
skills installed: 11
ok everything checks outFor a first read-only workflow, launch hermes and ask for a daily standup from a notes folder, or launch omp inside a project. Start with sample files and inspect changes before granting real access.
3. Install Claude Code and Codex separately
The Harness installs neither client and signs you into no cloud account. Install them with npm, using a Node.js version the client supports, and read the Claude Code README and Codex README first.
npm install -g @anthropic-ai/claude-code
npm install -g @openai/codex
claude --version
codex --versionExamples use Claude Code 2.1.284 and Codex CLI 0.159.0, checked October 6. Sign in to your own provider account; that access is separate from the free Harness.
4. Understand the local API before changing a client
The bundled llama.cpp b11430 server answers OpenAI-style chat requests on loopback. This tests the model directly, without the large agent prompt:
curl -s http://127.0.0.1:8484/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3-4b-2507","messages":[{"role":"user","content":"Say OK"}],"max_tokens":20}'The same server returned HTTP 200 and OK through /v1/messages and /v1/responses on October 6. A full coding session also depends on prompt size and on how well the model uses tools, so test that separately.
Claude Code: a small local connection test
Set environment variables for this one invocation so your cloud configuration stays unchanged. The token is a dummy value for the loopback-only server.
ANTHROPIC_BASE_URL=http://127.0.0.1:8484 \
ANTHROPIC_AUTH_TOKEN=local \
ANTHROPIC_MODEL=qwen3-4b-2507 \
ANTHROPIC_SMALL_FAST_MODEL=qwen3-4b-2507 \
CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536 \
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 \
claude -p --tools '' --disable-slash-commands \
--system-prompt 'You are testing a local endpoint. Respond to text without using tools.' \
'Reply with exactly these words: CLAUDE LOCAL OK'Claude Code returned CLAUDE LOCAL OK in 43 seconds on October 6 and warned that the model was missing from its catalog. This test turns tools off and uses a short system prompt. The default prompt gave no answer in 15 minutes on the same CPU server, so use this as a connection check and run full sessions on a GPU or Apple Silicon machine.
Codex: a separate local provider configuration
Use a separate Codex configuration directory so your cloud profile stays intact. The provider uses the Responses API.
mkdir -p "$HOME/.codex-meshvault"
cat > "$HOME/.codex-meshvault/config.toml" <<'EOF'
model = "qwen3-4b-2507"
model_provider = "meshvault"
[model_providers.meshvault]
name = "MeshVault local model"
base_url = "http://127.0.0.1:8484/v1"
wire_api = "responses"
EOF
CODEX_HOME="$HOME/.codex-meshvault" codex exec \
--sandbox read-only --skip-git-repo-check \
'Reply with exactly these words: CODEX LOCAL OK'With this provider, Codex returned CODEX LOCAL OK in 316 seconds on October 6 and reported 6,097 tokens used. It warned that model metadata was missing. That was one text reply. A cloud coding model next to a small local assistant often works better than forcing every client onto the small model.
You can point an existing Ollama, LM Studio or vLLM server at the Harness instead. This supported installer option skips its bundled model download:
bash meshvault-install.sh \
--endpoint http://127.0.0.1:11434/v1 \
--endpoint-model qwen3:8bStart that server with a context window of at least 64,000 tokens, because Hermes sends a long first prompt.
Troubleshooting without reinstalling everything
meshvault: command not found: open a new terminal or add$HOME/.local/binto PATH with the export above.- The model will not start: inspect
meshvault logs. If memory is exhausted, usemeshvault model use qwen3-1.7band close competing workloads. - The first Hermes reply takes minutes: CPU prefill must process its long prompt. Test
meshvault askfirst. Keep the toolset small withhermes tools list; consider GPU acceleration. - OMP has no local provider: an existing
models.ymlis preserved. Use the provider block printed by the installer and compare the shipped template.
Update with meshvault update, hermes update and omp update. Back up your configuration before changing models or providers.
What stays private, and what still needs approval
The bundled model listens on loopback only and the Harness adds no telemetry. Cloud models, public data tools and messaging services send data to their providers when you enable them. The installer does not sandbox agents, so they can read files and run commands as your user.
Keep write approvals enabled. Skills that ask before sending or deleting are instructions, not an operating-system boundary. Keep payments, account security and patient information away from untested local models.
The $499 Harness install covers one machine and a first workflow. Agent Setup at $450 covers your choice of Claude Code, Codex, Hermes or OMP with project rules and up to three tool connections.
Sources and test dates
- Harness 1.0.0 install instructions, setup implementation and troubleshooting, checked October 6, 2026.
- Current delivery and tier scope, checked October 6, 2026.
- TheFireDev clean-container proof, October 5, 2026, and workstation checks, October 6, 2026, using sample prompts.
Questions
Can Claude Code or Codex run on a local model?
Both accept a custom endpoint, and the Harness serves an OpenAI-style API on 127.0.0.1:8484. On October 6, 2026 a short no-tools prompt worked with each. A full default Claude Code prompt on a CPU-only server gave no answer within 15 minutes, so coding sessions need a GPU, Apple Silicon or a larger local model.
Is the MeshVault Harness free?
The installer, the meshvault command and 11 skills are MIT licensed and free. The Pro skills pack is $99 and a done-for-you install is $499. Neither is required to run a local model.
How much RAM do I need for a local model?
The installer asks for 8 GB for the 1.7B model, 12 GB for 4B, 20 GB for 8B and 40 GB for the 30B-A3B model. 16 GB is the practical minimum once your editor and browser are open.
Does the Harness send my data anywhere?
The Harness adds no telemetry and the bundled model listens on loopback only. Data goes to a cloud provider only when you enable that provider, a public data tool or a messaging service.