Guide

Running a local LLM on a Mac mini

Yes, a Mac mini runs local LLMs well, but memory decides which ones. A 16 GB Mac mini fits models up to about 8 billion parameters with room to spare. Models around 30 billion need 48 GB. 70 billion needs 64 GB. Below are the real file sizes, the install commands for Ollama and MLX, and how to keep a model server up.

Why memory is the limit

Apple silicon has unified memory. The CPU and GPU share one pool. That is good for LLMs, because the GPU can use most of the RAM. But macOS, your apps, the model weights and the model's working memory all come from that same pool.

The working memory grows with context length. Ollama's docs say a larger context needs more memory. They also pick the default context from available GPU memory (docs.ollama.com/context-length, October 2026).

  • Under 24 GiB: 4k tokens.
  • From 24 to 48 GiB: 32k tokens.
  • 48 GiB or more: 256k tokens.

So a 16 GB Mac starts with the smallest context. LM Studio recommends 16 GB or more on a Mac. It says 8 GB machines should stick to smaller models (lmstudio.ai system requirements, October 2026).

Our rule of thumb is simpler, and it is ours, not a vendor's. Keep the model file at least 6 to 8 GB smaller than the Mac's memory. That leaves room for macOS, the context and an agent or a browser. Leave more if you also run builds on the same Mac.

What fits, with real file sizes

These are download sizes for the default builds, which are 4-bit for these models. They come from the Ollama library pages in October 2026.

  • qwen3:8b is 5.2 GB. llama3.1:8b is 4.9 GB. Comfortable on 16 GB.
  • gemma3:12b is 8.1 GB. qwen3:14b is 9.3 GB. On 16 GB that is a tight fit with little room for anything else.
  • gpt-oss:20b is 14 GB. On a 16 GB Mac that leaves almost nothing for macOS. Treat it as out of reach.
  • gemma3:27b is 17 GB, qwen3:30b is 19 GB and qwen3.5:35b is 24 GB. These want 48 GB.
  • llama3.1:70b is 43 GB. That needs 64 GB.
  • gpt-oss:120b is 65 GB. It does not fit any Mac mini, since 64 GB is the most a Mac mini holds.

Small models also make more mistakes. In our test below, a 0.6 billion parameter model called the Mac mini "a portable computer." Fine for a smoke test. Not for real work.

Install Ollama

Ollama needs macOS 14 Sonoma or newer (docs.ollama.com/macos). Get the app from the download page, or use the install script. On macOS the script puts Ollama.app in Applications and links the ollama command into the local bin folder.

curl -fsSL https://ollama.com/install.sh | sh

We ran it on GitHub hosted macOS 26 machines on October 2, 2026. It installed in about 7 seconds. In two of our runs, the last step failed with the message Unable to find application named 'Ollama'. The app was installed anyway. Opening it by its path worked.

open /Applications/Ollama.app

We then pulled a tiny model and ran it.

$ ollama --version
ollama version is 0.35.0
$ ollama pull qwen3:0.6b
$ ollama list
NAME          ID              SIZE      MODIFIED
qwen3:0.6b    7df6b6e09427    522 MB    Less than a second ago

For real use, pull a model that fits your memory. Then check that it runs on the GPU. The Processor column should say 100% GPU.

ollama pull qwen3:8b
ollama run qwen3:8b
ollama ps

Or use MLX or LM Studio

MLX is Apple's own machine learning framework. Its LLM package installs with pip. We installed it in a virtual environment and ran a 1 billion parameter 4-bit model.

python3 -m venv ~/mlx && source ~/mlx/bin/activate
pip install mlx-lm
mlx_lm.generate --model mlx-community/Llama-3.2-1B-Instruct-4bit --prompt "What is a Mac mini?" --max-tokens 40

It printed mlx-lm version 0.32.0 and a peak memory of 0.773 GB. For large models, the MLX README shows a setting that raises the memory the GPU may hold. It needs macOS 15 or newer. Keep N above the model size in MB and below your total RAM.

sudo sysctl iogpu.wired_limit_mb=N

LM Studio is a desktop app with a model browser. Its lms command ships with the app. The docs show lms server start for an API server and lms daemon up for running without the window. We did not test LM Studio for this page.

Keep the model server running

By default Ollama unloads a model after 5 minutes idle, per its FAQ. The next request then waits for a reload. For an always-on server, keep models loaded. The FAQ says to set variables for the Mac app with launchctl, then restart the app.

launchctl setenv OLLAMA_KEEP_ALIVE "-1"

Turn off sleep too, or the server freezes with the Mac.

sudo pmset -a sleep 0
sudo pmset -a disksleep 0
sudo pmset -a displaysleep 10

Security notes

Ollama listens on 127.0.0.1 port 11434 by default. Its API did not ask for any password in our test. The FAQ shows how to listen on all networks with OLLAMA_HOST. Do not do that on a Mac that other people can reach. To use the model from your laptop, forward the port over SSH instead.

ssh -N -L 11434:127.0.0.1:11434 user@your-mac

Then point your tools at localhost port 11434 on your laptop. Hermes Agent and OpenClaw can both use a local endpoint like this.

Which MacRun Mac fits

Our stocked M6 has 16 GB of memory and a 256 GB SSD. It runs 8B models well next to an agent. It is the wrong machine for 30B models, and we would rather say so now.

  • M5 Pro with 48 GB and a 1 TB SSD: $299 a month on the agent plan, plus $199 build fee. Fits the 27B to 35B models above.
  • M5 Pro with 64 GB and a 1 TB SSD: $349 a month on the agent plan, plus $249 build fee. Fits 70B models at 4-bit.

Both are pre-orders, ready in about a week. Claude Code and Codex come installed on the agent plan. You install Ollama, MLX or LM Studio yourself with admin rights.

Who should look elsewhere. If you need more than 64 GB, a Mac mini cannot hold it. We do not offer Mac Studio or GPU servers. Compare all tiers on the pricing page.

Frequently asked questions

Can a 16 GB Mac mini run a local LLM?

+

Yes. Models up to about 8 billion parameters fit with room to spare, like qwen3:8b at 5.2 GB. Models of 12 to 14 billion are a tight fit with little memory left for anything else.

How much memory do I need for a 70B model?

+

A 4-bit 70B model such as llama3.1:70b is a 43 GB download on Ollama. You need a 64 GB Mac mini for it, plus room for context.

Is Ollama or MLX faster on a Mac mini?

+

Both use the Apple GPU. MLX is Apple's own framework and Ollama is simpler to run as a server. Test both with the same model on your own machine.

How do I keep an Ollama model loaded all the time?

+

Set OLLAMA_KEEP_ALIVE to -1 with launchctl setenv and restart the Ollama app. By default it unloads a model after 5 idle minutes.

Related guides