Install Claude Code, Codex and a local model on one machine
The installer, the separate coding clients, working commands and the limits of a small local model.
The free MeshVault Harness installs llama.cpp, a Qwen3 model, Hermes Agent and OMP. Claude Code and Codex are separate installs. You can keep cloud coding tools beside a local agent without sending every task to a paid API.
On October 6, 2026, the published installer completed in 6 minutes 31 seconds on a Linux workstation. It ran in an isolated home directory, with the 4B model selected explicitly. That is one installation, not a promise about download speed on your connection.
Before you install
Use Linux or macOS as a normal user. Windows needs WSL2. Have curl, git and tar available. The install docs recommend 30 GB of free disk and 16 GB of RAM. Linux may need the system libraries libatomic and libgomp; the installer can request admin access for those.
The version 1.0.0 catalog includes memory allowance for a 65,536-token context and quantized KV cache. The minimums below are installer thresholds, not free memory on an already busy machine.
| Model ID | Minimum RAM | Download |
|---|---|---|
qwen3-1.7b | 8 GB | 1.11 GB |
qwen3-4b-2507 | 12 GB | 2.50 GB |
qwen3-8b | 20 GB | 5.03 GB |
qwen3-30b-a3b-2507 | 40 GB | 18.56 GB |
Download sizes are rounded decimal GB from the catalog. More RAM lets a model fit; it does not make a CPU process a long prompt quickly. Reserve memory for your editor, browser and operating system.
1. Read the installer, then run it
The entry script fetches the newest release tag and runs its setup script. The model and llama.cpp downloads are pinned and checked with SHA-256. The official Hermes and OMP installers have their own update paths.
curl -fsSL https://raw.githubusercontent.com/thefiredev-cloud/meshvault-harness/main/install.sh -o meshvault-install.sh
less meshvault-install.sh
bash meshvault-install.sh --model qwen3-4b-2507The published one-command version is:
curl -fsSL https://raw.githubusercontent.com/thefiredev-cloud/meshvault-harness/main/install.sh | bash -s -- --model qwen3-4b-2507The 4B choice makes this example repeatable across larger machines. Leave out --model to let the installer choose from the catalog. Linux defaults to CPU. Use --gpu vulkan only with the Vulkan driver and libraries installed; macOS uses Metal. A CUDA build can be supplied with --llama-server PATH.
These are actual lines from the October 6 workstation run, with unrelated lines omitted:
MeshVault Harness 1.0.0
Downloading llama.cpp b11430 (linux-x64)
ok llama.cpp b11430 installed
ok model verified (sha256)
ok model server up on http://127.0.0.1:8484 (model: qwen3-4b-2507)
ok synced 11 skill(s) from free pack
ok everything checks out2. Check the model before asking an agent to work
export PATH="$HOME/.local/bin:$PATH"
meshvault status
meshvault doctor
meshvault ask "What is 17 times 3? Answer with the number only."
meshvault skills listA healthy doctor run checks an answer from the model, the agents and their configuration. This is the October 6 output after the install:
MeshVault Harness 1.0.0 on linux/x64, 124 GB RAM
ok local model answers
ok hermes installed (Hermes Agent v0.21.5+8160.g4ca6f9e (2026.9.24) · upstream 4ca6f9e9)
ok hermes points at the local model
ok omp installed (omp/18.6.3)
skills installed: 11
ok everything checks outFor a first read-only workflow, launch hermes and ask for a daily standup from a local notes folder. Launch omp inside a project for coding work. Start with sample files, then inspect the changes before giving an agent access to real work.
A clean Ubuntu 24.04 container with a 16 GB memory limit produced HERMES LOCAL OK in 8 minutes 10 seconds and OMP LOCAL OK in 44 seconds during the October 5 install proof. The shared host was heavily loaded. Those are full reply times, including prompt processing, not tokens per second. A simple meshvault ask reply took 3.44 seconds in that run.
3. Install Claude Code and Codex separately
The Harness does not install either client or sign you into a cloud account. These are their npm install commands. Use a supported Node.js version for the current client release. Read the Claude Code README and Codex README before installing.
npm install -g @anthropic-ai/claude-code
npm install -g @openai/codex
claude --version
codex --versionThe connection examples use Claude Code 2.1.284 and Codex CLI 0.159.0, checked October 6. Run claude or codex and complete provider sign-in yourself when using cloud models. Your subscription or API access is separate from the free Harness.
4. Understand the local API before changing a client
The bundled llama.cpp b11430 server answers OpenAI-style chat requests at the local loopback address below. This command tests the model directly, without the much larger agent instructions:
curl -s http://127.0.0.1:8484/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3-4b-2507","messages":[{"role":"user","content":"Say OK"}],"max_tokens":20}'The same bundled server returned HTTP 200 and OK through /v1/messages and /v1/responses in direct protocol tests on October 6. A complete coding session also depends on request format, prompt size and model tool use. One successful HTTP request does not prove reliable agent behavior.
Claude Code: a small local connection test
Use environment settings for this one invocation, so your normal cloud configuration stays unchanged. The token below is a dummy value for the loopback-only model server, not a provider credential.
ANTHROPIC_BASE_URL=http://127.0.0.1:8484 \
ANTHROPIC_AUTH_TOKEN=local \
ANTHROPIC_MODEL=qwen3-4b-2507 \
ANTHROPIC_SMALL_FAST_MODEL=qwen3-4b-2507 \
CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536 \
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 \
claude -p --tools '' --disable-slash-commands \
--system-prompt 'You are testing a local endpoint. Respond to text without using tools.' \
'Reply with exactly these words: CLAUDE LOCAL OK'Claude Code returned CLAUDE LOCAL OK in 43 seconds on October 6. It also warned that the local model was unknown to its catalog. This test disables tools and uses a short system prompt. The earlier default-prompt invocation produced no answer within 15 minutes on the same CPU server. Treat the successful short test as a connection check, not proof of a usable full coding session.
Codex: a separate local provider configuration
Use a dedicated Codex configuration directory for this example. The provider uses the Responses API. Keep an existing cloud profile intact rather than overwriting its configuration.
mkdir -p "$HOME/.codex-meshvault"
cat > "$HOME/.codex-meshvault/config.toml" <<'EOF'
model = "qwen3-4b-2507"
model_provider = "meshvault"
[model_providers.meshvault]
name = "MeshVault local model"
base_url = "http://127.0.0.1:8484/v1"
wire_api = "responses"
EOF
CODEX_HOME="$HOME/.codex-meshvault" codex exec \
--sandbox read-only --skip-git-repo-check \
'Reply with exactly these words: CODEX LOCAL OK'With this provider configuration, Codex returned CODEX LOCAL OK in 316 seconds on October 6 and reported 6,097 tokens used. It warned that model metadata was missing. This was one text reply, not a tool-use or code-editing benchmark. A cloud coding model beside a smaller local assistant can be a more useful setup than forcing every client onto that small model.
You can point an existing Ollama, LM Studio or vLLM server at the Harness instead. This supported installer option skips its bundled model download:
bash meshvault-install.sh \
--endpoint http://127.0.0.1:11434/v1 \
--endpoint-model qwen3:8bStart that server with a context window large enough for Hermes: at least 64,000 tokens. A short chatbot prompt fitting in memory says little about processing the long first agent prompt.
Troubleshooting without reinstalling everything
meshvault: command not found: open a new terminal or add$HOME/.local/binto PATH with the export above.- The model will not start: inspect
meshvault logs. If memory is exhausted, usemeshvault model use qwen3-1.7band close competing workloads. - The first Hermes reply takes minutes: CPU prefill must process its long prompt. Test
meshvault askfirst. Keep the toolset small withhermes tools list; consider GPU acceleration. - Hermes reports a small context window: retain
--ctx 65536, or raise the context on the external server. - Port 8484 is occupied: choose an unused port with
--port 8585. Do not stop a process you do not own. - OMP has no local provider: an existing
models.ymlis preserved. Use the provider block printed by the installer and compare the shipped template. - A skill is missing: run
meshvault skills sync. A same-named folder you created is intentionally skipped.
Update the free catalog and skills with meshvault update. Hermes and OMP update separately with hermes update and omp update. Back up your configuration before changing models or providers.
What stays private, and what still needs approval
The bundled model listens on loopback only. The Harness adds no telemetry. Cloud-model calls, public data tools and messaging services still send data to their respective providers when you enable them. The installer does not sandbox agents: they can read files and run commands as your user.
Keep write approvals enabled. Markdown skills that ask before sending or deleting are useful instructions, not an operating-system security boundary. Never give an untested local model unattended control of payments, account security or patient information.
The $499 Harness install includes one machine, a first workflow, the Pro pack and 30 days of email support. The separate $450 Agent Setup covers your choice of Claude Code, Codex, Hermes or OMP, project rules and up to three tool connections. Neither buys hardware or a cloud subscription.
Sources and test dates
- Harness 1.0.0 install instructions, setup implementation and troubleshooting, checked October 6, 2026.
- Current delivery and tier scope, checked October 6, 2026.
- TheFireDev clean-container proof, October 5, 2026; isolated Linux workstation install and direct local API checks, October 6, 2026. These tests used sample prompts, not client data.