Install Claude Code, Codex and a local model on one machine

The installer, the separate coding clients, working commands and the limits of a small local model.

The free MeshVault Harness installs llama.cpp, a Qwen3 model, Hermes Agent and OMP. Claude Code and Codex are separate installs. You can keep cloud coding tools beside a local agent without sending every task to a paid API.

On October 6, 2026, the published installer completed in 6 minutes 31 seconds on a Linux workstation. It ran in an isolated home directory, with the 4B model selected explicitly. That is one installation, not a promise about download speed on your connection.

Before you install

Use Linux or macOS as a normal user. Windows needs WSL2. Have curl, git and tar available. The install docs recommend 30 GB of free disk and 16 GB of RAM. Linux may need the system libraries libatomic and libgomp; the installer can request admin access for those.

The version 1.0.0 catalog includes memory allowance for a 65,536-token context and quantized KV cache. The minimums below are installer thresholds, not free memory on an already busy machine.

Qwen3 model choices in Harness 1.0.0
Model IDMinimum RAMDownload
qwen3-1.7b8 GB1.11 GB
qwen3-4b-250712 GB2.50 GB
qwen3-8b20 GB5.03 GB
qwen3-30b-a3b-250740 GB18.56 GB

Download sizes are rounded decimal GB from the catalog. More RAM lets a model fit; it does not make a CPU process a long prompt quickly. Reserve memory for your editor, browser and operating system.

1. Read the installer, then run it

The entry script fetches the newest release tag and runs its setup script. The model and llama.cpp downloads are pinned and checked with SHA-256. The official Hermes and OMP installers have their own update paths.

curl -fsSL https://raw.githubusercontent.com/thefiredev-cloud/meshvault-harness/main/install.sh -o meshvault-install.sh
less meshvault-install.sh
bash meshvault-install.sh --model qwen3-4b-2507

The published one-command version is:

curl -fsSL https://raw.githubusercontent.com/thefiredev-cloud/meshvault-harness/main/install.sh | bash -s -- --model qwen3-4b-2507

The 4B choice makes this example repeatable across larger machines. Leave out --model to let the installer choose from the catalog. Linux defaults to CPU. Use --gpu vulkan only with the Vulkan driver and libraries installed; macOS uses Metal. A CUDA build can be supplied with --llama-server PATH.

These are actual lines from the October 6 workstation run, with unrelated lines omitted:

MeshVault Harness  1.0.0
Downloading llama.cpp b11430 (linux-x64)
 ok  llama.cpp b11430 installed
 ok  model verified (sha256)
 ok  model server up on http://127.0.0.1:8484 (model: qwen3-4b-2507)
 ok  synced 11 skill(s) from free pack
 ok  everything checks out

2. Check the model before asking an agent to work

export PATH="$HOME/.local/bin:$PATH"
meshvault status
meshvault doctor
meshvault ask "What is 17 times 3? Answer with the number only."
meshvault skills list

A healthy doctor run checks an answer from the model, the agents and their configuration. This is the October 6 output after the install:

MeshVault Harness 1.0.0 on linux/x64, 124 GB RAM
 ok  local model answers
 ok  hermes installed (Hermes Agent v0.21.5+8160.g4ca6f9e (2026.9.24) · upstream 4ca6f9e9)
 ok  hermes points at the local model
 ok  omp installed (omp/18.6.3)
skills installed: 11
 ok  everything checks out

For a first read-only workflow, launch hermes and ask for a daily standup from a local notes folder. Launch omp inside a project for coding work. Start with sample files, then inspect the changes before giving an agent access to real work.

A clean Ubuntu 24.04 container with a 16 GB memory limit produced HERMES LOCAL OK in 8 minutes 10 seconds and OMP LOCAL OK in 44 seconds during the October 5 install proof. The shared host was heavily loaded. Those are full reply times, including prompt processing, not tokens per second. A simple meshvault ask reply took 3.44 seconds in that run.

3. Install Claude Code and Codex separately

The Harness does not install either client or sign you into a cloud account. These are their npm install commands. Use a supported Node.js version for the current client release. Read the Claude Code README and Codex README before installing.

npm install -g @anthropic-ai/claude-code
npm install -g @openai/codex
claude --version
codex --version

The connection examples use Claude Code 2.1.284 and Codex CLI 0.159.0, checked October 6. Run claude or codex and complete provider sign-in yourself when using cloud models. Your subscription or API access is separate from the free Harness.

4. Understand the local API before changing a client

The bundled llama.cpp b11430 server answers OpenAI-style chat requests at the local loopback address below. This command tests the model directly, without the much larger agent instructions:

curl -s http://127.0.0.1:8484/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3-4b-2507","messages":[{"role":"user","content":"Say OK"}],"max_tokens":20}'

The same bundled server returned HTTP 200 and OK through /v1/messages and /v1/responses in direct protocol tests on October 6. A complete coding session also depends on request format, prompt size and model tool use. One successful HTTP request does not prove reliable agent behavior.

Claude Code: a small local connection test

Use environment settings for this one invocation, so your normal cloud configuration stays unchanged. The token below is a dummy value for the loopback-only model server, not a provider credential.

ANTHROPIC_BASE_URL=http://127.0.0.1:8484 \
ANTHROPIC_AUTH_TOKEN=local \
ANTHROPIC_MODEL=qwen3-4b-2507 \
ANTHROPIC_SMALL_FAST_MODEL=qwen3-4b-2507 \
CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536 \
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 \
claude -p --tools '' --disable-slash-commands \
  --system-prompt 'You are testing a local endpoint. Respond to text without using tools.' \
  'Reply with exactly these words: CLAUDE LOCAL OK'

Claude Code returned CLAUDE LOCAL OK in 43 seconds on October 6. It also warned that the local model was unknown to its catalog. This test disables tools and uses a short system prompt. The earlier default-prompt invocation produced no answer within 15 minutes on the same CPU server. Treat the successful short test as a connection check, not proof of a usable full coding session.

Codex: a separate local provider configuration

Use a dedicated Codex configuration directory for this example. The provider uses the Responses API. Keep an existing cloud profile intact rather than overwriting its configuration.

mkdir -p "$HOME/.codex-meshvault"
cat > "$HOME/.codex-meshvault/config.toml" <<'EOF'
model = "qwen3-4b-2507"
model_provider = "meshvault"

[model_providers.meshvault]
name = "MeshVault local model"
base_url = "http://127.0.0.1:8484/v1"
wire_api = "responses"
EOF
CODEX_HOME="$HOME/.codex-meshvault" codex exec \
  --sandbox read-only --skip-git-repo-check \
  'Reply with exactly these words: CODEX LOCAL OK'

With this provider configuration, Codex returned CODEX LOCAL OK in 316 seconds on October 6 and reported 6,097 tokens used. It warned that model metadata was missing. This was one text reply, not a tool-use or code-editing benchmark. A cloud coding model beside a smaller local assistant can be a more useful setup than forcing every client onto that small model.

You can point an existing Ollama, LM Studio or vLLM server at the Harness instead. This supported installer option skips its bundled model download:

bash meshvault-install.sh \
  --endpoint http://127.0.0.1:11434/v1 \
  --endpoint-model qwen3:8b

Start that server with a context window large enough for Hermes: at least 64,000 tokens. A short chatbot prompt fitting in memory says little about processing the long first agent prompt.

Troubleshooting without reinstalling everything

  • meshvault: command not found: open a new terminal or add $HOME/.local/bin to PATH with the export above.
  • The model will not start: inspect meshvault logs. If memory is exhausted, use meshvault model use qwen3-1.7b and close competing workloads.
  • The first Hermes reply takes minutes: CPU prefill must process its long prompt. Test meshvault ask first. Keep the toolset small with hermes tools list; consider GPU acceleration.
  • Hermes reports a small context window: retain --ctx 65536, or raise the context on the external server.
  • Port 8484 is occupied: choose an unused port with --port 8585. Do not stop a process you do not own.
  • OMP has no local provider: an existing models.yml is preserved. Use the provider block printed by the installer and compare the shipped template.
  • A skill is missing: run meshvault skills sync. A same-named folder you created is intentionally skipped.

Update the free catalog and skills with meshvault update. Hermes and OMP update separately with hermes update and omp update. Back up your configuration before changing models or providers.

What stays private, and what still needs approval

The bundled model listens on loopback only. The Harness adds no telemetry. Cloud-model calls, public data tools and messaging services still send data to their respective providers when you enable them. The installer does not sandbox agents: they can read files and run commands as your user.

Keep write approvals enabled. Markdown skills that ask before sending or deleting are useful instructions, not an operating-system security boundary. Never give an untested local model unattended control of payments, account security or patient information.

The $499 Harness install includes one machine, a first workflow, the Pro pack and 30 days of email support. The separate $450 Agent Setup covers your choice of Claude Code, Codex, Hermes or OMP, project rules and up to three tool connections. Neither buys hardware or a cloud subscription.

Sources and test dates