Free Shipping on all orders · Priority Mail Shipping with fee of $8.00
🧠
Divine Tribe Software · Open Source

claude-code-local

Run Claude Code on your own Mac — no cloud, no subscriptions, no internet required.

Python⭐ 3,139 stars🍴 595 forksMIT licensedOpen source
⭐ 3,139
GitHub Stars
🍴 595
Forks
💻 Python
Primary Language
📅 August 2026
Last Updated
See it in action

Watch the demo

What it is

All the power of Claude Code, without sending a single byte to the cloud.

Claude Code Local turns your Mac into a complete AI coding workstation. The same chat-with-your-code experience you get from the cloud — except the model actually lives on your laptop. It's fast enough to feel instant, smart enough to handle real work, and private enough that your code never leaves the room.

We built it after getting frustrated paying monthly fees for an AI that couldn't be used on sensitive projects. A year of tinkering later, you can read an NDA on an airplane with the Wi-Fi off, and the model will still summarize it in twenty seconds.

Why it's different

What makes claude-code-local special

Blazing fast

Up to 65 tokens per second — faster than cloud Opus, running on the laptop you already own.

🔒

100% private

Your code, questions, and documents never touch the internet. Verified with live network audits.

💸

$0 per month

Forever. No subscription, no API fees, no surprise bills.

✈️

Works offline

Airplane, cabin, hotel Wi-Fi from hell — doesn't matter. Everything runs locally.

🎛️

Three models, one click

Gemma for daily driving, Llama for hard reasoning, Qwen for raw speed. Swap between them instantly.

🗣️

Talk to it hands-free

Pair with NarrateClaude to speak to your computer and hear it answer in your own cloned voice.

Who it's for

Is this for you?

  • Developers who want faster, free Claude Code without the meter running
  • Lawyers and accountants who can't upload client documents to the cloud
  • Anyone who wants to use AI on a plane, in a cabin, or off-grid
  • Privacy-minded folks who'd rather own their tools than rent them
How to get it

Getting started in minutes

1

Clone the repo

One git clone and you're halfway there. No Docker, no complicated setup.

2

Run the setup script

It picks the right model for your Mac's RAM, downloads it, and installs everything. One command.

3

Double-click the launcher

A Claude Local.command file lands on your desktop. Open it and you're coding locally.

Ready to try claude-code-local?

It's free, open source, and runs on the hardware you already own. Head to GitHub to get started, or drop a star to help us keep building in public.

For firms · Confidential workflows

Need this for a law firm, healthcare org, or anywhere documents can't leave the machine?

claude-code-local runs the same on your firm's MacBook as it does on mine. AirGap AI is the commercial pilot — a 14-day engagement that ships claude-code-local (and the rest of the local-AI stack) into a real legal/medical workflow with verified network audits. Privileged docs in, answers out, never a byte to a cloud.

Explore the AirGap pilot → Get in touch
Stay in the tribe

More from Divine Tribe

Full technical docs

The complete README

Open the GitHub README — every detail, every benchmark, every code block

🧠⚡ Claude Code Local

Run Claude Code 100% on-device with local AI on Apple Silicon.
No cloud, no API key, no proxy — an MLX-native server that speaks the Anthropic API.
🥊 Pick your fighter: Gemma 4 31B · Llama 3.3 70B · Qwen 3.5 122B · DeepSeek V4 Flash (1M context via ds4).

GitHub stars GitHub forks 5 Models Qwen 3.5 speed Claude Code task time 100% Local Hands-Free Voice MIT Join the NiceDreamzApps Discord

🛑 Usage Limit · 🤔 What Is This · 🚀 Quick Start · 🥊 Lineup · 🎮 Modes · 🔒 Privacy · 📊 Benchmarks · 🎤 Voice · 🌐 Browser · 📱 Phone · 🔌 MCP · 🛣️ Roadmap


🛑 Hit your Claude usage limit?

If Claude Code just told you "you've reached your usage limit" and gave you a reset time hours away, that's what this is for. You keep working — same Claude Code, same terminal, same project — except the model answering is running on your own Mac.

curl -fsSL https://raw.githubusercontent.com/nicedreamzapp/claude-code-local/main/install.sh | bash

No API key. No second subscription. No waiting until 3pm. It works on a 16 GB MacBook and gets better the more RAM you have — see what runs on your Mac.

Claude Code editing a file with Gemma 4 31B running locally on a Mac, no cloud

Real session, unedited. Claude Code reads and edits the file — the model answering is Gemma 4 31B on the laptop.


🤔 What Is This?

Your Mac has a powerful GPU built right into the chip. This project uses that GPU to run massive AI models — the same kind that power ChatGPT and Claude — entirely on your computer, and plugs them into Claude Code so the whole coding experience works offline.

🚫 No internet needed 💰 No monthly subscription 🔒 No one sees your code or data ✅ Full Claude Code experience — write code, edit files, manage projects, control your browser, or run a full hands-free voice session

         📱 You (Mac or Phone)
          │
     🤖 Claude Code           ← the AI coding tool you know
          │  HTTP localhost:4000
     ⚡ MLX Native Server      ← this repo (~1000 lines of Python)
          │
     🥊 Pick your fighter     ← Gemma 4 31B · Llama 3.3 70B · Qwen 3.5 122B
          │
     🖥️  Apple Silicon GPU    ← your M-series chip does all the work

The trick: Claude Code speaks the Anthropic API. Local model servers speak the OpenAI API. So everyone bolts a translation proxy in between — and the proxy is slow and fragile. This server speaks Anthropic natively. One process, zero translations:

🐌 What everyone else does 🚀 What we did
Claude Code → Proxy → Ollama → Model Claude Code → Our Server → Model
3 processes, 2 API translations 1 process, 0 translations
133 seconds per task 17.6 seconds per task

🎯 That one change — eliminating the proxy — made it 7.5× faster.


🎬 Watch It Run — AirGap AI

A real NDA. Llama 3.3 70B. Wi-Fi physically OFF. lsof running live. Watch a 70-billion-parameter model audit a confidential legal document, on-device, with the receipts on screen.

AirGap AI — Wi-Fi OFF NDA Demo

AirGap is this whole build running as one private workstation — a capability, not a product. Everything you need is in this repo. If your firm needs one built, here's what it looks like.

More local-AI demos on the channel:

Video What happens
🌌 The Rematch 4 AI engines build northern lights, 3 fully local — the local challenger painted the best aurora
🏁 Hexagon Shootout Gemma 31B vs Llama 70B vs cloud Claude, same physics prompt, live counters — 2 of 3 with zero cloud calls
🐳 DeepSeek Three-Way DeepSeek V4 Flash local beats cloud Claude on wall-clock, same MacBook
🎤 NarrateClaude Speak to Claude Code, hear replies in a cloned voice — 100% on-device
🏠 Mac mini as home AI Chat with the Mac mini at home from any browser on any phone

🥊 The Lineup — Pick Your Fighter

We started with one model. Now we ship a roster. Same MLX server, same Anthropic API — swap one env var and you swap the brain. Plus the ds4 engine for DeepSeek V4 Flash via its own native Metal runtime.

🟡 Hermes 4 14B 🟢 Gemma 4 31B 🟠 Llama 3.3 70B 🔵 Qwen 3.5 122B 🐳 DeepSeek V4 Flash
Nickname The One That Runs On Your Laptop The Quick One The Wise One The Beast The 1M-Context Whale
Build 4-bit abliterated 4-bit IT abliterated 8-bit abliterated 4-bit MoE (A10B) 2-bit asymmetric (ds4 GGUF)
Speed not benchmarked yet ~15 tok/s ~7 tok/s 65 tok/s 🚀 ~32 tok/s
Params 14 B dense (Qwen3 base) 31 B dense 71 B dense 122 B / 10 B active 284 B / 37 B active
Context 40 K 128 K 128 K 256 K 1 M tokens
RAM ~8 GB ~18 GB ~70 GB ~75 GB ~81 GB
Min RAM to run 16 GB 32 GB 96 GB 96 GB 128 GB
Best at Everyday edits on a stock MacBook Daily coding Hardest reasoning, full precision Max throughput, active sparsity Long context, agentic loops
Engine MLX Native MLX Native MLX Native MLX Native antirez/ds4
Launcher Claude Local.command Gemma 4 Code.command Llama 70B.command Claude Local.command DeepSeek V4 Flash.app

💻 Got a 16 GB MacBook Air? Start with Hermes. setup.sh picks it for you automatically — you don't need 96 GB of RAM to use this.

💡 Fun fact: Qwen wins raw speed because it's an MoE — only 10B of 122B params activate per token. DeepSeek V4 Flash is even bigger (284B) but only ~37B active per token, and it ships with on-disk KV cache so a 25k-token Claude Code system prompt prefills exactly once, ever.

🐳 DeepSeek V4 Flash via ds4

We tested it the day Antirez (the Redis guy) shipped ds4. Local DeepSeek beat cloud Claude on wall-clock time on the same MacBook, same prompt — watch the three-way.

🧠 Engine antirez/ds4 — pure C + Metal kernels, ~few thousand lines
🤗 Weights antirez/deepseek-v4-gguf (q2: 81 GB, q4: 153 GB)
📦 Server wrapper ~/.local/bin/ds4-server-up (boots on demand)
🚀 Claude Code wrapper ~/.local/bin/claude-ds4 (drop-in replacement for claude)
📏 Context 1 M tokens; 200 K is sane for most agent runs
💾 Disk KV cache Persists across restarts — first prefill is the only one that ever happens

⭐ Our Own MLX Abliterated Uploads

The models in this lineup aren't from generic mirrors — we package and upload our own abliterated MLX builds to HuggingFace so anyone running this repo can pull them with one command. Browse the full set at huggingface.co/divinetribe.

# Llama 3.3 70B — full-precision feel
MLX_MODEL=divinetribe/Llama-3.3-70B-Instruct-abliterated-8bit-mlx \
  bash scripts/start-mlx-server.sh

# Gemma 4 31B — fast daily driver
MLX_MODEL=divinetribe/gemma-4-31b-it-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh

# Hermes 4 14B — sweet spot for 16/32 GB Macs
MLX_MODEL=divinetribe/Hermes-4-14B-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh
Model Quant Disk Params Context Best for
Llama-3.3-70B-Instruct-abliterated-8bit-mlx 8-bit, g64 ~75 GB 71 B dense 128 K Hardest reasoning on 96 GB+ Macs
gemma-4-31b-it-abliterated-4bit-mlx 4-bit, g64 ~17 GB 31 B dense 128 K Daily coding on a 32 GB+ Mac
Hermes-4-14B-abliterated-4bit-mlx 4-bit, g64 ~8 GB 14 B dense (Qwen3 base) 40 K 16 GB Macs, instruction-following, tool use

Abliteration sources: huihui-ai (Llama, Gemma) and Babsie (Hermes). MLX conversion + quantization by us. See what abliteration means.

⚠️ Use it responsibly. "Abliterated" suppresses the model's built-in refusal direction so it doesn't refuse benign-but-edgy requests. It is not a general capability upgrade, and you remain bound by each upstream license (Llama 3.3, Gemma, Hermes/Qwen3).


🎮 The Modes

Four ways to run the lineup. Each one is a double-clickable launcher in launchers/.

Mode What it does Launcher
🤖 Code Run Claude Code with a local model — same UX, no API key Claude Local.command, Gemma 4 Code.command, Llama 70B.command
🌐 Browser Local AI controls real Brave browser via Chrome DevTools Browser Agent.command
🎤 Hands-Free Voice Speak in, hear replies in your cloned voice — full loop, 100% on-device Narrative Gemma.command + NarrateClaude
📱 Phone iMessage in → text/image/video out, via claude-screen-to-phone ~/.claude/imessage-*.sh

💻 What You Need

Your Mac RAM What setup.sh installs for you
MacBook Air / base M1-M4 16 GB 🟡 Hermes 4 14B — yes, this works
M1/M2/M3/M4 Pro 32-48 GB 🟢 Gemma 4 12B
M2/M3/M4/M5 Max 64-95 GB 🟢 Gemma 4 31B
M3/M4/M5 Max · Ultra 96 GB+ 🔵 Qwen 3.5 122B, 🟠 Llama 70B, 🐳 DeepSeek

Also need:

  • 🐍 Python 3.12+ (for MLX)
  • 🤖 Claude Code (npm install -g @anthropic-ai/claude-code)

🚀 Quick Start (One Command)

curl -fsSL https://raw.githubusercontent.com/nicedreamzapp/claude-code-local/main/install.sh | bash

Or clone it yourself if you'd rather read the script first:

git clone https://github.com/nicedreamzapp/claude-code-local
cd claude-code-local
bash setup.sh

setup.sh auto-detects your RAM, picks a model from the lineup, downloads it, installs the MLX server, and creates a Claude Local.command launcher on your Desktop.

Then double-click Claude Local.command. You're coding locally.

🐛 If the launcher asks you to sign in to a Claude account: your claude CLI is too old. The launchers pass --bare to force local-only API-key auth; older CLIs don't support it. Fix: npm install -g @anthropic-ai/claude-code

🛠️ Note for contributors: setup.sh installs the server as a symlink at ~/.local/mlx-native-server/server.py pointing back at this repo's proxy/server.py. Edit the file in the repo, restart the server, done — one source of truth, no silent drift.

Or do it manually

# 1. Set up the MLX virtualenv
python3.12 -m venv ~/.local/mlx-server
~/.local/mlx-server/bin/pip install mlx-lm

# 2. Pick a fighter and download (one time, ~18-75 GB)
bash scripts/download-and-import.sh gemma   # or 'llama' or 'qwen'

# 3. Start the server
MLX_MODEL=divinetribe/gemma-4-31b-it-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh

# 4. Launch Claude Code
ANTHROPIC_BASE_URL=http://localhost:4000 \
ANTHROPIC_API_KEY=sk-local \
claude --model claude-sonnet-4-6

🔧 How It Works

┌──────────────────────────────────────────────────┐
│              Your MacBook (M-series)             │
│                                                  │
│  📝 You type ──> 🤖 Claude Code                  │
│                      │                           │
│                      ▼                           │
│                 ⚡ MLX Server (port 4000)        │
│                      │                           │
│                      ▼                           │
│                 🥊 Local model ──> 🖥️  GPU        │
│                 (Gemma·Llama·Qwen)               │
│                      │                           │
│                      ▼                           │
│  📝 Answer <─── ✨ Clean response                │
│                                                  │
│         🔒 Nothing leaves this box. Ever.        │
└──────────────────────────────────────────────────┘

The server (proxy/server.py) is one file, ~1000 lines. It does six things:

  1. 📦 Loads the model — Apple's MLX framework, native Metal GPU, unified memory. Handles Gemma's RotatingKVCache quirk automatically.
  2. 🔌 Speaks Anthropic API — Claude Code thinks it's talking to Anthropic's cloud. It's not.
  3. 🔧 Translates tool use — Three tool-call formats in and out: Gemma 4 native, Llama 3.3 raw JSON, and HuggingFace <tool_call> JSON (Qwen and others). All converted ↔ Anthropic tool_use blocks, with garbled-output recovery for small models.
  4. 🧹 Cleans the output — A real-time ThinkingFilter strips <think> blocks token-by-token during generation, then clean_response handles stop markers and reasoning preamble.
  5. Reuses prompt caches across requests — Claude Code's system prompt doesn't get re-prefilled every turn. Huge speedup for short questions.
  6. 🎯 Code mode — auto-detects Claude Code coding sessions, swaps the ~10K-token harness prompt for a slim ~150-token one, and strips verbose tool descriptions to name + parameter types. A 28× prompt reduction that cuts prefill from ~60 s to ~2 s on Gemma 4 31B.

🛤️ The Journey

We didn't start here. Three generations in one night:

Gen What We Tried Speed 💡 What We Learned
1️⃣ Ollama + custom proxy 30 tok/s Ollama works but Claude Code can't talk to it directly
2️⃣ llama.cpp TurboQuant + proxy 41 tok/s TurboQuant compresses KV cache 4.9x, but the proxy is the bottleneck
3️⃣ MLX native server 65 tok/s Kill the proxy. Speak Anthropic API directly. 7.5x faster.
4️⃣ The lineup 65 / 15 / 7 tok/s Three brains, one server — swap one env var to change the fighter

🔒 Privacy + How the Data Flows

This is the part we're proudest of. Your code never leaves your Mac. Not for a model call. Not for telemetry. Not for "anonymous analytics". Not ever.

   ┌─────────────────────────────────────────────────────────────┐
   │                    🖥️  YOUR MACBOOK                          │
   │                                                             │
   │   📝 Your code ──> 🤖 Claude Code ──> ⚡ MLX Server          │
   │                     (localhost:4000)      │                 │
   │                                           ▼                 │
   │                    🧠 Local model ──> 🖥️  Apple GPU          │
   │                                                             │
   │             🚫 ZERO outbound network calls                  │
   │             🚫 ZERO telemetry                               │
   │             🚫 ZERO phone-home                              │
   └─────────────────────────────────────────────────────────────┘
                   │
                   ✗  ←  Nothing from our code crosses this line.
                   │
   ┌─────────────────────────────────────────────────────────────┐
   │                    ☁️  THE INTERNET                          │
   │                  (your code never goes here)                 │
   └─────────────────────────────────────────────────────────────┘

🔍 What We Audited (Every Component)

Component Source Outbound calls Verdict
server.py (ours) We wrote it line by line 0 ✅ Safe
browser agent nicedreamzapp/browser-agent — we wrote it 0 (localhost CDP only) ✅ Safe
mlx-lm Apple ML team 0 ✅ Safe
MLX framework Apple 0 ✅ Safe
Model weights HuggingFace verified repos 0 at runtime ✅ Safe
Claude Code CLI Anthropic (closed-source binary) 0 with our launchers — lsof-verified ✅ Safe

Verified offline. Claude Code's own binary previously reached out to api.anthropic.com on startup for telemetry, statsig feature flags, marketplace auto-install, and the autoupdater. The launchers plug all four channels via documented Anthropic env vars (thanks @tadrianonet, PR #32):

CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
DISABLE_AUTOUPDATER=1
CLAUDE_CODE_DISABLE_OFFICIAL_MARKETPLACE_AUTOINSTALL=1
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1

Verify it yourself: run lsof -p $(pgrep -f claude) during a session — you'll see only localhost:4000. Run lsof -i -P while the server is up — nothing leaves your Mac.

⚠️ We removed LiteLLM after supply-chain attack concerns. Every dependency was re-audited from scratch. If a package had unexplained network calls, it didn't ship.

✈️ When To Use This

Situation Use This? Why
On a plane (no wifi) Full AI coding, no internet needed
NDA / sensitive client code Nothing leaves your machine — air-gapped, lsof-verified
Healthcare / legal / finance review 100% on-device, audit-friendly
Don't want API fees $0/month forever
Want fastest possible ☁️ Cloud Sonnet is still slightly faster
Need Claude-level reasoning ☁️ Local models are good, not Claude-level

📊 Benchmarks

Three generations of optimization. Each one got faster.

⚡ Speed Comparison

Generation Approach Speed Real Claude Code task
🐌 Gen 1 Ollama + Proxy 30 tok/s 133 s
🏃 Gen 2 llama.cpp + Proxy 41 tok/s 133 s
🚀 Gen 3 MLX Native (ours) 65 tok/s 17.6 s

🥊 Lineup Comparison

Model tok/s RAM Best For
🟢 Gemma 4 31B Abliterated ~15 ~18 GB Daily coding on a 64 GB Mac
🟠 Llama 3.3 70B Abliterated ~7 ~70 GB Hardest reasoning, full precision
🔵 Qwen 3.5 122B-A10B 65 ~75 GB Maximum throughput, MoE sparsity

☁️ vs Cloud APIs

🖥️ Our Local Setup ☁️ Claude Sonnet ☁️ Claude Opus
Speed 65 tok/s ~80 tok/s ~40 tok/s
Monthly cost $0 🎉 $20-100+ $20-100+
Privacy 100% local 🔒 Cloud Cloud
Works offline Yes ✈️ No No

💡 Our local setup beats cloud Opus on raw speed (65 vs 40 tok/s) at $0/month. Qwen numbers measured on M5 Max 128 GB — full details in BENCHMARKS.md.


🔧 Tool-Call Reliability

Local models don't format tool calls perfectly. They want to call a tool but mix XML and JSON syntax — Claude Code sees no valid tool call, re-prompts, and the model garbles it the same way again. The result: infinite loops where the AI says "let me do that" but never does anything.

We fixed this with 4 changes to server.py:

Change What Why
KV Cache 4-bit → 8-bit, quantization starts at token 1024 Model retains conversation context
Temperature 0.7 → 0.2 Less randomness = more consistent tool formatting
Garbled Recovery recover_garbled_tool_json() Catches XML-in-JSON hybrids, infers tool names from parameter keys
Retry Logic Up to 2 retries when tool intent is detected but parsing fails Re-prompts with explicit formatting instructions

🧪 Results: 98/98 tests passed across 7 consecutive runs. Zero failures. The multi-step scenario that used to trigger infinite loops — create 12 month folders, delete all but September, verify — now passes every time. Run it yourself:

python3 scripts/test_mlx_server.py

⚙️ Tuning

Variable Default What It Does
MLX_MODEL divinetribe/gemma-4-31b-it-abliterated-4bit-mlx Pick which fighter to load
MLX_KV_BITS 8 KV cache quantization bits (4 saves memory, 8 improves coherence)
MLX_KV_QUANT_START 1024 Token position where KV quantization begins
MLX_TOOL_RETRIES 2 Max retries when a garbled tool call is detected
MLX_MAX_TOKENS 8192 Max output tokens per response
MLX_SUPPRESS_THINKING 1 Skip the model's reasoning chain (~1 min/request saved). Set 0 to let it think.
MLX_BROWSER_MODE 0 Optimize for chrome-devtools MCP sessions — keeps only the 9 essential browser tools (~99% fewer tokens)

📚 More

Everything above gets you running. These live in docs/ so this page stays short:

🎤 Hands-Free Voice Mode Talk to Claude Code, hear it answer in a cloned voice
🌐 Browser Agent Let the local model drive your real browser
📱 Control From Your Phone Run a session on your Mac from anywhere
🔌 MCP Servers Claude Code's whole plugin ecosystem, 100% local
📁 What's In This Repo File-by-file tour
📊 Benchmarks · 🔧 Tool-Call Reliability The numbers and how they were measured
📱 Apps From The Same Workshop · 🙏 Credits Everything else

🧩 The Local-First Stack

claude-code-local is the brain. It pairs with sibling repos — each stands alone, together they take Claude Code off the keyboard and off the screen:

Repo Role What it does
🤖 claude-code-local Brain (you are here) MLX Anthropic server · launcher lineup · tool-call translation
🎤 NarrateClaude Ears + Mouth Talk to Claude, hear replies in your cloned voice — both directions on-device
🌐 browser-agent Hands Drives real Brave via CDP — iframes, Shadow DOM, ProseMirror
📱 claude-screen-to-phone Remote iPhone → Claude Code over iMessage; text/screenshots/videos back
🛟 claude-failover Backstop Keep cloud Claude primary, flip one command to local when limits pinch or Anthropic is down

🛣️ What's Next

We ship fast and in public. If any of these excite you, hit Watch to get the release ping.

  • 🟡 Full Qwen 3.5 122B benchmark suite — reliability, tool-call pass rate, long-context behavior vs Gemma
  • 🟡 Fully-local Whisper fallback — alternative to the Apple SFSpeechRecognizer path for older Macs and non-English voices
  • 🟡 One-click DMG installer — no terminal needed
  • 🟡 MLX_MODEL=<hf-url> — point at any HuggingFace repo and auto-register a new fighter
  • 🟡 More fighters — open to PRs adding launchers for DeepSeek, Mistral, Phi, anything MLX-compatible

💡 Want something that's not on this list? Open an issue → Every serious request gets read and usually replied to within 24h.

🤝 Contributing

Ideas, bug reports, a new launcher for a model I don't run, a better code-mode prompt — open an issue or a PR, I read them all. Especially interested in: folks on older Apple Silicon (M1/M2, 16–36 GB) who know which models actually fit; anyone stress-testing the voice loop on different hardware or accents; TTS recipes beyond Pocket TTS (Piper, MLX-TTS, Kyutai Moshi); and edge cases I'll never hit on an M5 Max with 128 GB.


🙏 Credits

🧑‍🔧 Contributors

Every one of these landed on hardware I don't own, on a bug I hadn't hit. Thank you.

Who What they fixed
@0xshugo Client disconnects handled, retries skipped when there are no tools (#4)
@asdmoment Gemma inference crash — auto-disable KV quantization (#7)
@kulveersingh ArraysCache has no attribute offset (#10)
@tripathiprateek uninstall.sh — reverses setup.sh cleanly (#23)
@tadrianonet Mac base/Pro 16 GB support: Qwen 2.5 14B, ChatML stop markers, <tools> parser, offline leak fix (#32)
@kevbarns Gemma 4 thinking suppression + slimmer tool descriptions — ~4× latency cut (#33)
@KaoCSC Stop on the tokenizer's real EOS, and tolerate empty env ints (#41) · bare JSON tool calls, which took Qwen 2.5 Coder from 0/12 to 14/14 (#43)

Tested on Apple M5 Max with 128 GB unified memory.

Built by Matt Macosko in Arcata, CA — part of Nice Dreamz LLC. More open-source at nicedreamzwholesale.com/software · demos at youtube.com/@nicedreamzapps.

X YouTube GitHub


📜 MIT License — Use it however you want.

💬 Builders hang out on Discord — share what you're building, swap MLX tips.

Star this repo if it helped you!

Upstream projects this is built on are listed in docs/CREDITS.md.