The interesting thing about local AI in 2026 isn’t that it’s possible — it’s that the boring parts finally work. A Mac with 16 GB of unified memory, or a PC with an 8 GB GPU, runs a 7B–13B model well enough for transcription, OCR, routine chat and the small repetitive tasks that make up most of the day. No API key, no per-token meter, no outage on someone else’s status page.

What follows is the toolbox: thirteen tools, grouped by what they actually do. I verified each one’s repository, language and license while writing this, because half the “top tools” lists circulating right now cite projects that have been renamed or abandoned. Where something surprised me, I say so.


1. The runtime: Ollama or LM Studio

Everything else assumes you can run a model. There are two sane entry points and they are not competitors so much as different moments in the same workflow.

Ollama (Go, MIT) is the one that ends up in production. It’s a daemon with a CLI and an OpenAI-compatible HTTP API, which is the property that matters: anything that speaks the OpenAI protocol — and by now that’s nearly everything — can be pointed at it by changing a base URL.

ollama pull qwen2.5:14b
ollama run qwen2.5:14b
ollama serve            # OpenAI-compatible API on :11434

LM Studio is a GUI. You browse models, click download, and chat in the same window. It is genuinely the faster way to answer “is this model any good on my hardware” before you commit to scripting anything.

One honest caveat that most round-ups skip: LM Studio itself is not open source. The app is free for personal use and its CLI and SDKs are open, but the core is proprietary. If open source is a requirement rather than a preference, Ollama is the answer and LM Studio is a convenience you should know the shape of.

My rule: LM Studio to explore, Ollama to build.


2. Knowing what your hardware can actually run

This is where most people waste an afternoon. You download a 70B model, it swaps, you conclude local AI is a toy. The failure was in model selection, not in the idea.

Three tools attack this from different angles.

CanIRun.ai — the thirty-second answer

canirun.ai detects your GPU or Mac in the browser and tells you which open LLMs, coding, image and video models fit, with VRAM requirements, speed estimates and letter grades. No install. It’s the right first stop, and the right thing to send to someone who asks you what they can run.

llmfit — the scriptable version

llmfit (Rust, MIT) is the same question as a CLI: one command to find what runs on your hardware, across hundreds of models and providers. It scans the system, offers a TUI to filter, and — the part I find most useful — works in reverse: give it a model, get the hardware you’d need. That’s the mode to use before buying a machine rather than after.

whichllm — ranked by benchmarks, not parameter count

whichllm (Python, MIT) ranks the models that fit your hardware by measured performance rather than size. This distinction is the whole point: a well-tuned 7B routinely beats a badly-quantized 13B on the same box, and parameter count alone will never tell you that.

Rough sizing, if you want a number before running anything:

7B  (Q4_K_M)   8–16 GB RAM      comfortable on any M-series Mac
13B             16–32 GB RAM     the sweet spot for daily use
30B+            64 GB, or dedicated GPU VRAM

A GPU is not required. CPU inference on a 7B Q4 runs roughly 5–15 tokens/sec against 30–60 with a GPU — slow enough to notice in chat, entirely fine for batch transcription or OCR you’re not watching.


3. Apple Silicon: the model you already paid for

apfel — Apple Intelligence from the terminal

If you’re on Apple Silicon, there’s already a foundation model on the disk, reachable only through Siri and system UI. apfel (Swift, MIT) exposes it three ways: a pipe-friendly CLI, an interactive chat, and an OpenAI-compatible server.

apfel "summarize this" < notes.txt
apfel serve                        # OpenAI-compatible, localhost:11434

Zero cost, zero API keys, zero download — the weights ship with the OS. The constraints are real and worth stating: it needs macOS 26 Tahoe or later, Apple Silicon, Apple Intelligence enabled, and the context window is 4096 tokens. That last number rules out document work, and makes it excellent for exactly the short, high-frequency tasks you’d otherwise burn cloud calls on — classification, rewriting a sentence, extracting a field.

Note the port: it defaults to the same 11434 Ollama uses, so pick one or move it.


4. Speech to text

This is the category where local models are not a compromise. On-device transcription in 2026 is accurate, fast, and removes the awkward question of which vendor is holding recordings of your meetings.

Murmure — cross-platform dictation

Murmure (Rust, AGPL-3.0) is fully local speech-to-text with LLM post-processing, running on Mac, Windows and Linux. It’s fast enough on CPU alone.

Scriberr — the archive, not the dictation

Scriberr (Go, MIT) is self-hosted batch transcription with speaker diarization and an API. Different job entirely: Murmure is for the sentence you’re saying now, Scriberr is for the four hundred hours of podcasts, lectures and recorded meetings you want indexed and searchable. Self-hosted, so it slots next to whatever else is in your homelab.

Petal — multi-engine, in the menu bar

Petal (Swift, MIT) is a native macOS menu bar app that puts several engines behind one interface — Apple Speech, Qwen3 ASR, Parakeet via MLX Audio, Voxtral, and Whisper through WhisperKit. If you want to compare engines on your own audio rather than trusting a benchmark, this is the least tedious way to do it.

Ghost Pepper — hold-to-talk, and a nice trick

Ghost Pepper (Swift, macOS 14+, Apple Silicon) is hold-Control-to-record, release-to-paste. WhisperKit or Parakeet v3 for recognition, then — this is the part worth stealing — a small local LLM (Qwen 3.5) cleans up filler words and self-corrections before the text lands at your cursor. The two-stage pipeline is the difference between a raw transcript and something you’d send.

One caution: the repository carries no license file at the time of writing. Free and open to read is not the same as licensed for reuse; treat it as an app to run, not code to vendor.

Also: ghostpepper.app is not this tool. That domain now redirects to an unrelated web-based image generator. The project lives at matthartman/ghost-pepper on GitHub.


5. Text to speech

Voicebox — the other direction

Voicebox (TypeScript, MIT) is a local voice studio built on Qwen3-TTS: clone a voice from a few seconds of reference audio, compose multi-voice narration on a timeline, and drive it all from a REST API. macOS and Windows now, Linux in progress. Unlimited generation with no per-second billing is the practical argument — it changes what you’re willing to re-record.

The obvious caveat, which I’d rather state than skip: voice cloning from three seconds of audio is trivially misusable. Clone your own voice, or one you have explicit permission to use.


6. Images and documents

SnapOtter — 30+ file tools in one container

The article that prompted this post lists this one as Stirling-Image. It has since been renamed to SnapOtter (TypeScript, AGPL-3.0) — the old GitHub path still redirects, which is how I noticed. It’s a self-hosted file-processing tool: convert, compress, OCR, transcribe, background removal, upscaling, and more, all local, no telemetry.

Docker Compose deployment, which makes it a natural fit next to anything else you’re already running behind a reverse proxy — the same pattern as serving a static site securely.

LiteParse — documents into text, properly

LiteParse (Rust, Apache-2.0) comes from the LlamaIndex team and is described as a fast, model-free document parser. That phrase is the selling point: it extracts structure without an inference pass, which makes it fast, deterministic, and cheap enough to run over a whole corpus. If you’re assembling a local RAG pipeline, this is the ingestion stage, and it means the documents never leave the machine to become text.


Which model, for what

Once the tooling is in place, the remaining question is which weights to pull. Reasonable defaults:

general chat        Qwen 2.5 14B, or Llama 3.3 70B if you have the memory
code                Qwen 2.5 Coder 14B, or DeepSeek Coder V3
multilingual (EU)   Mistral Small 3, or Llama 3.3
tight on resources  Phi 4, or Llama 3.2 3B

Verify against your own hardware with one of the tools in section 2 rather than trusting the table — that’s the entire reason those tools exist.


Why bother

Four arguments, in the order they actually convince people:

  1. The data stays put. No third party holds your recordings, your documents, or your codebase. For anything under NDA this isn’t a preference, it’s a requirement.
  2. No meter. Inference is free once the hardware is bought. It changes behaviour — you stop rationing calls, and you start using the model for things that were never worth a per-token cost.
  3. Latency. No round trip, no queue, no rate limit at the wrong moment.
  4. It keeps working. Offline, on a plane, and after the vendor pivots, raises prices, or deprecates the model you built on.

Where local still loses is hard reasoning. For genuinely difficult problems, frontier models are better and it isn’t close — which is why my Claude Code setup is a separate post, and why wiring a terminal agent to a local model is a different exercise with different expectations.

The sensible position isn’t local instead of cloud. It’s local for the ninety percent that’s transcription, OCR, embedding, extraction and routine chat — and the frontier reserved for the problems that actually need it.


References