Haplo

How to Run an LLM Locally on a Mac

Five ways to run a language model on your own Mac, with the commands we ran on an M2 MacBook Pro: how much memory each model size needs, which open models to start with in October 2026, how fast they run, and what stays on the machine.

17 min read

Close-up of a MacBook Pro's backlit keyboard and speaker grille, lit orange by its open screen and reflected in a glossy desk
A 16-inch MacBook Pro with the M2 Max chip. On Apple silicon, the GPU works from the same memory as everything else · Photo: SimonWaldherr, CC BY-SA 4.0 (Resized)

To run an LLM locally on a Mac, install a local runner such as Ollama, LM Studio, llama.cpp or Apple’s MLX, download an open model that fits in your Mac’s memory, and chat with it on the machine itself, with Wi-Fi off if you like. Memory decides what fits: plan on about 0.6 GB per billion parameters for a 4-bit model, plus room for the conversation, so a 16 GB Mac handles models of up to about 14 billion parameters. We make Haplo AI, a Mac and iPhone app that downloads open models such as Qwen and Phi and runs them on the device, so it’s the fifth way below.

We ran the commands marked as tested on an M2 MacBook Pro with 24 GB of memory in October 2026, and quote the rest from each project’s documentation. For phones, see how to run AI offline on your iPhone.

Five ways to run an LLM on a Mac, compared

All five run open-weight models on your Mac’s own chip. They differ in how you use them and what else they do.

Way What it is Install with Intel Macs Serves other apps
Ollama Command-line tool and desktop app App or script CPU only Yes
LM Studio Desktop app with a model browser and a command-line tool App No Yes
llama.cpp The open-source engine that Ollama, LM Studio and Haplo AI use Homebrew CPU only Yes
MLX LM Apple’s machine learning framework, used from Python pip No Yes
Haplo AI Our app for Mac, iPhone and iPad Mac App Store No No

In short: Ollama if you like Terminal or want a local API for other tools, LM Studio to browse and try models in a window, llama.cpp for the most control, MLX if you work in Python, and Haplo AI for an app that picks models for your Mac with no setup.

How much memory do you need to run an LLM on a Mac?

About 0.6 GB per billion parameters for a 4-bit model, 1.1 GB at 8-bit and 2.2 GB at 16-bit, plus memory for the conversation and for macOS.

On Apple silicon, the CPU and GPU share one pool of memory. Apple describes the M1’s “unified memory architecture” as “high-bandwidth, low-latency memory” in “a single pool” that every part of the chip can use “without copying it” (Apple). So your Mac’s memory is also the GPU’s memory, and its size decides which models fit.

Apple's M1 chip on a Mac mini logic board: a silver lid with the Apple logo and part numbers, and two dark memory chips beside it in the same package
An M1 on a 2020 Mac mini's logic board. The two dark chips beside the silver lid are its memory, inside the same package · Photo: Sonic8400, CC BY-SA 4.0 (cropped)

The rule of thumb. A model’s size is its parameter count times the bits stored per parameter. Google’s Gemma 4 documentation puts its 31-billion-parameter model at 17.5 GB in 4-bit, 34.9 GB in 8-bit and 69.9 GB in 16-bit, including a 20 percent overhead: about 0.6, 1.1 and 2.2 GB per billion parameters. llama.cpp’s usual 4-bit format, Q4_K_M, averages 4.89 bits per weight against 16.0 uncompressed (llama.cpp), and Ollama’s Llama 3.1 8B is 4.9 GB at Q4_K_M, 8.5 GB at 8-bit and 16 GB at 16-bit (Ollama).

The conversation. A model keeps a cache of the conversation so far, the KV cache, which grows with every token. In llama.cpp, Qwen2.5-Coder 14B (8.4 GB of weights) needed 768 MB more at a 4,096-token context and 6 GB more at 32,768. So Ollama picks its default context by GPU memory: 4k tokens below 24 GB, 32k from 24 to 48 GB and 256k from 48 GB up (Ollama). Ours got 4,096.

What macOS keeps. On our 24 GB M2, Ollama, llama.cpp and MLX all reported the same GPU limit: 17.8 GB, about three quarters of the total. The MLX LM documentation describes raising it on macOS 15 or later with sudo sysctl iogpu.wired_limit_mb=N, with N between the model’s size and your memory in megabytes; we left ours alone.

Mixture of experts. A name like Gemma 4 26B A4B means 26 billion parameters in all, about 4 billion used per token. These run about as fast as a small model, but Google notes “all 26 billion parameters must be loaded into memory,” so size them by the first number.

How big a model fits on your Mac

Our working limit is a download of about half your memory on an 8 GB Mac and about 60 percent from 16 GB up, which leaves room for some context and your other apps. Sizes are Ollama’s 4-bit downloads.

Mac memory Models that fit at 4-bit Download size to aim for Examples
8 GB Up to about 4 billion parameters Up to about 4 GB Qwen3.5 4B (3.3 GB), Ministral 3 3B (3.0 GB), Phi-4-mini (2.5 GB), Llama 3.2 3B (2.0 GB)
16 GB Up to about 14 billion Up to about 10 GB Qwen3.5 9B (6.6 GB), Gemma 4 12B (7.7 GB), Phi-4 and Ministral 3 14B (9.1 GB each)
24 to 32 GB About 20 to 30 billion About 15 GB on 24 GB, 20 GB on 32 GB gpt-oss-20b (14 GB), Mistral Small 3.2 24B (15 GB), Gemma 4 26B A4B (16 GB), Qwen3.8 27B (18 GB), Gemma 4 31B (19 GB)
64 GB and up 70 billion, or large mixture-of-experts models About 60 percent of memory Llama 3.3 70B (43 GB, tight on 64 GB); gpt-oss-120b (65 GB) and Qwen3.5 122B-A10B (81 GB) on 128 GB

Apple doubled the MacBook Air’s starting memory from 8 GB to 16 GB in October 2024 (Apple), so 8 GB Macs are mostly older ones.

How to run an LLM on a Mac with Ollama

Install Ollama, then type ollama run and a model name in Terminal: it downloads the model the first time and opens a chat. We tested Ollama 0.32.12.

  1. Install it. Ollama needs macOS Sonoma (14) or newer (docs). Download the app from ollama.com/download, or use the install script from Ollama’s README:

    curl -fsSL https://ollama.com/install.sh | sh
    
  2. Run a model that fits. The docs’ example, ollama run gemma4, fetches Gemma 4 E4B, a 6.6 GB download. Add a size tag from Ollama’s library to choose; on an 8 GB Mac, for example:

    ollama run qwen3.5:4b
    
  3. Check the speed. Add --verbose and Ollama prints timings after each reply. With Qwen2.5 7B, our M2 showed:

    eval count:           108 token(s)
    eval rate:            17.99 tokens/s
    
  4. Check it’s on the GPU. In ollama ps, 100% GPU under PROCESSOR means the whole model is on the GPU (FAQ):

    NAME          ID              SIZE      PROCESSOR    CONTEXT    UNTIL
    qwen2.5:7b    845dbda0ea48    4.7 GB    100% GPU     4096       4 minutes from now
    
  5. Use it from other apps. Ollama serves a REST API at localhost:11434. This is the README’s request, with a model we had:

    curl http://localhost:11434/api/chat -d '{
      "model": "qwen2.5:3b",
      "messages": [{
        "role": "user",
        "content": "Why is the sky blue?"
      }],
      "stream": false
    }'
    
  6. Tidy up. ollama ls lists your models, ollama stop unloads one and ollama rm deletes it. They live in ~/.ollama/models.

Ollama also has cloud models, which run on Ollama’s servers and carry a :cloud tag, such as gemma4:cloud (docs). To keep it local only, set OLLAMA_NO_CLOUD=1 and restart it; our server log then read Ollama cloud disabled: true.

How to run an LLM on a Mac with LM Studio

Download LM Studio, get a model in its Discover tab, load it in the Chat tab and chat. LM Studio isn’t installed on our test Mac, so these steps come from its documentation.

  1. Check your Mac. LM Studio needs Apple silicon and macOS 14.0 or newer, recommends 16 GB or more, and doesn’t support Intel Macs. On 8 GB Macs, its docs say to “stick to smaller models and modest context sizes” (LM Studio).

  2. Install it from lmstudio.ai/download. The Free plan runs local models; a $20-a-month Bionic+ plan adds models hosted in the cloud (pricing).

  3. Download a model in the Discover tab (⌘2): search by name, then pick a version. The docs’ advice is 4-bit or higher if your machine can run it (LM Studio). On Apple silicon it runs both GGUF models, with llama.cpp, and MLX models.

  4. Load it and chat. In the Chat tab, open the model loader (⌘L), choose the model and start typing.

  5. Or use Terminal. The lms command comes with the app; open LM Studio once first (CLI docs):

    lms get llama-3.1-8b
    lms load <model_key>
    lms chat
    lms server start
    

lms load --estimate-only prints a memory estimate without loading, and the server speaks the OpenAI API (LM Studio’s examples use http://localhost:1234/v1).

How to run an LLM on a Mac with llama.cpp

Install llama.cpp with Homebrew, then run llama cli -hf with a model’s Hugging Face name: it downloads the file and opens a chat in Terminal. We tested Homebrew’s build b9730.

  1. Install it (docs):

    brew install llama.cpp
    

    The README also offers curl -LsSf https://llama.app/install.sh | sh. Older guides use llama-cli and llama-server; our build has those alongside the newer llama command.

  2. Run a model from Hugging Face. The README’s example is llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF. That repository has no Q4_K_M file, in which case -hf falls back to the first file in the repo, so we named the 563 MB 4-bit version:

    llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF:Q4_0
    

    After each reply it prints the speed; on our M2, [ Prompt: 111.5 t/s | Generation: 51.2 t/s ]. Qwen3.5 reasoned before answering in our test, and --reasoning off skips that.

  3. Serve it. llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF:Q4_0 started a server at http://127.0.0.1:8080 with a built-in web chat and an OpenAI-compatible /v1/chat/completions endpoint, which we called with curl.

  4. Go offline. Our build kept downloads in ~/.cache/huggingface/hub (llama cli --cache-list lists them), and --offline makes it use that cache without touching the network.

How to run an LLM on a Mac with Apple’s MLX

Install the mlx-lm Python package and run mlx_lm.generate or mlx_lm.chat with an MLX model from Hugging Face. MLX is Apple’s open-source machine learning framework, and it needs an Apple silicon Mac, a native Python 3.10 or newer, and macOS 14.0 or newer (MLX docs). We tested mlx-lm 0.32.0.

  1. Install it, here in a throwaway virtual environment:

    python3 -m venv mlx-venv
    source mlx-venv/bin/activate
    pip install mlx-lm
    
  2. Generate. The README’s quick start, mlx_lm.generate --prompt "How tall is Mt Everest?", uses mlx-community/Llama-3.2-3B-Instruct-4bit, a 1.8 GB download. We named the 713 MB 1B version:

    mlx_lm.generate --model mlx-community/Llama-3.2-1B-Instruct-4bit --prompt "How tall is Mt Everest?"
    

    It ended with:

    Generation: 32 tokens, 69.662 tokens-per-sec
    Peak memory: 0.772 GB
    
  3. Chat. mlx_lm.chat --model mlx-community/Llama-3.2-1B-Instruct-4bit opens a chat in Terminal; ours replied at 73.8 tokens per second.

  4. Serve it. mlx_lm.server --model <path_to_model_or_hf_repo> runs an HTTP API on port 8080 of localhost, though its docs say it isn’t recommended for production as it “only implements basic security checks” (MLX LM).

  5. Run offline. MLX fetches models with Hugging Face’s library, which checks for a newer version each time a model loads. Set HF_HUB_OFFLINE=1 and “no HTTP calls will be made” (Hugging Face); we re-ran the generate command that way, from the cache.

Hugging Face’s MLX Community has ready-made MLX versions of thousands of models.

How to run an LLM on a Mac with Haplo AI

Haplo AI is our app for running open models without Terminal: download one from a list that only offers what your Mac’s memory can hold, and chat. It’s a native macOS app, not the iPhone app running on a Mac, and needs macOS 15.0 or later and an M1 or later (App Store). It’s free to download, but new users subscribe before they can chat: the US App Store lists Haplo AI Premium at $1.99 a month or $12.99 a year. One App Store app covers iPhone, iPad and Mac, a universal purchase, which Apple describes as access to an app and its in-app purchases across platforms “with a single purchase” (Apple).

The steps, checked against the app’s code:

  1. Install it from the Mac App Store. Setup ends at the subscription screen.

  2. Download a model. Click the brain icon in the toolbar. The list comes from our server, sized to your Mac’s memory, and models that won’t fit are marked “Incompatible with device”. In our checks, an 8 GB Mac gets 13 of the 16 text models (files up to 2.5 GB), 16 to 24 GB gets all but Qwen3.6-27B, and 32 GB or more gets all 16, with Qwen3-4B (2.5 GB) marked Recommended. Files download from Hugging Face.

  3. Chat. Start a new chat, pick a model and type. Each chat keeps its model, and chats are saved on the Mac.

  4. Keep every message on the Mac by turning off Automatic Tools in Settings (the gear icon). It’s on by default and lets a capable model choose a tool such as web search, which sends that search to a search engine.

It runs 4-bit GGUF models with llama.cpp: Alibaba’s Qwen (0.5B to 27B), Microsoft’s Phi-4-mini, Mistral’s Ministral 3 3B and Hugging Face’s SmolLM. On macOS 26 or later with Apple Intelligence turned on, it can also use Apple’s own on-device model.

The limits: it’s a subscription, its list tops out at 27 billion parameters, it doesn’t serve other apps the way Ollama and LM Studio do, and it needs Apple silicon. Its App Store privacy label is “Data Not Collected”.

Which LLM should you run on a Mac?

The largest recent model that fits your memory with room to spare. In October 2026 that means Qwen3.5 4B on an 8 GB Mac, Qwen3.5 9B or Gemma 4 12B on 16 GB, gpt-oss-20b or a 24B model on 24 GB, Qwen3.8 27B or Gemma 4 31B on 32 GB, Llama 3.3 70B on 64 GB, and gpt-oss-120b on 128 GB.

These open families are worth starting with, checked on each maker’s Hugging Face pages, with sizes from Ollama’s library:

Family Maker, licence Sizes Try in Ollama
Qwen3.5, 3.6 and 3.8 Alibaba, Apache 2.0 3.5: 0.8B to 397B (February 2026); 3.6: 27B, 35B-A3B (April); 3.8: 27B (August) qwen3.5:4b (3.3 GB), qwen3.5:9b (6.6 GB), qwen3.8:27b (18 GB)
Gemma 4 Google, Apache 2.0 E2B, E4B, 12B, 26B A4B, 31B gemma4:e2b (4.6 GB), gemma4:12b (7.7 GB), gemma4:31b (19 GB)
gpt-oss OpenAI, Apache 2.0 20b (21B, 3.6B active), 120b (117B, 5.1B active) gpt-oss:20b (14 GB), gpt-oss:120b (65 GB)
Ministral 3, Mistral Small Mistral AI, Apache 2.0 3B, 8B, 14B; 24B ministral-3:8b (6.0 GB), mistral-small3.2 (15 GB)
Phi-4 Microsoft, MIT Phi-4-mini 3.8B; Phi-4 and Phi-4-reasoning 14B phi4-mini (2.5 GB), phi4 (9.1 GB)
Llama Meta, Llama licences 3.2: 1B, 3B; 3.1: 8B; 3.3: 70B; Llama 4 Scout: 109B llama3.2 (2.0 GB), llama3.1:8b (4.9 GB), llama3.3 (43 GB)
DeepSeek R1 distilled DeepSeek, MIT 1.5B to 70B, built on Qwen and Llama deepseek-r1:8b (5.2 GB)

What the table doesn’t show:

  • Meta’s newest open models are the Llama 4 releases of April 2025. Downloading Llama from Hugging Face means requesting access first.
  • DeepSeek’s own models are too big (DeepSeek-V4-Flash has 291 billion parameters). The deepseek-r1 models people run locally are smaller ones trained on its reasoning: the 8B default is Qwen3 8B, post-trained on “the chain-of-thought from DeepSeek-R1-0528” (DeepSeek).
  • Reasoning takes time. Qwen3.8 has thinking “on by default” (Qwen), and gpt-oss lets you set reasoning effort to low, medium or high (OpenAI). You wait for those tokens before the answer starts.
  • Small models get things wrong. Asked to name three planets, the same 0.8-billion-parameter model answered “Jupiter, Saturn, and Uranus” on one run and “The four largest planets of our solar system are Jupiter, Saturn, Mars, and Earth” on another. Use the biggest model your Mac runs comfortably, and check anything that matters.

How fast is a local LLM on a Mac?

Fast enough to read along with, and faster the more memory bandwidth the chip has: a 7-billion-parameter model at 4-bit writes about 22 tokens per second on a base M2 and 83 on a 40-core M4 Max, in llama.cpp’s benchmarks.

To write each token, the chip reads the model’s weights from memory, so generation speed tracks memory bandwidth. These rows come from the llama.cpp project’s benchmark thread for Apple silicon, all running the same 7B Llama model at 4-bit (Q4_0):

Chip (GPU cores) Memory bandwidth Tokens per second
M1 (8) 68 GB/s 14.2
M2 (10) 100 GB/s 21.9
M4 (10) 120 GB/s 24.1
M5 (10) 154 GB/s 31.9
M4 Pro (20) 273 GB/s 50.7
M5 Pro (20) 307 GB/s 66.3
M4 Max (40) 546 GB/s 83.1
M5 Max (40) 614 GB/s 119.9
M3 Ultra (80) 800 GB/s 92.1
M5 Ultra (80) 1,228 GB/s 179.1

On one chip, bigger models are proportionally slower. On our M2, Ollama wrote 18 tokens per second with Qwen2.5 7B (4.7 GB) and 7 with Qwen2.5-Coder 14B (9.0 GB); Qwen3.5 0.8B in llama.cpp managed about 50, and Llama 3.2 1B in MLX about 70. Reading your prompt is quicker: llama-bench processed 158 prompt tokens per second with the 7B model. Simulators shared the GPU during some runs, when the same MLX command fell to 36 tokens per second, so treat our figures as a floor.

Can you run an LLM on an Intel Mac?

Yes, on the CPU only: Ollama and llama.cpp support Intel processors, while LM Studio, MLX and Haplo AI need Apple silicon.

Ollama supports “Apple M series (CPU and GPU support) or x86 (CPU only)” (docs), and llama.cpp supports x86 instruction sets such as AVX2 and AVX-512, while its Metal backend targets Apple silicon (README). LM Studio says “Intel-based Macs are currently not supported,” MLX needs Apple silicon, and Haplo AI an M1 or later. Without the GPU, start with a small model.

What stays on your Mac when you run an LLM locally?

Your prompts and the model’s answers are processed on the Mac. What crosses the network is model downloads, update checks, and any cloud or web feature you turn on.

  • Prompts and replies. “Ollama runs locally. We don’t see your prompts or data when you run locally,” says Ollama’s FAQ. LM Studio says that once a model is on your machine, “nothing you enter into LM Studio when chatting with LLMs leaves your device” (LM Studio).
  • Model downloads come from Hugging Face (LM Studio, llama.cpp, MLX and Haplo AI) or Ollama’s registry.
  • Background checks. LM Studio checks for app updates when it opens and goes online to search for models and download runtimes. MLX checks Hugging Face for a newer model file on each load unless HF_HUB_OFFLINE=1 is set; llama.cpp’s --offline does the same job.
  • Cloud features run on someone else’s servers: Ollama’s cloud models and web search, and LM Studio’s Bionic+ models. OLLAMA_NO_CLOUD=1 turns Ollama’s off.
  • Local servers listen only on your own Mac (127.0.0.1) unless you change that: Ollama and LM Studio default to it, and so did our llama serve. LM Studio recommends turning on authentication if you open its server to the network.
  • Haplo AI sends our server the device type, the app’s version numbers and your Mac’s memory to get its model list, and its tools, such as web search, send their queries online.

Local LLMs on a Mac: questions

Can a Mac with 8 GB of memory run an LLM?

Yes, small ones: models up to about 4 billion parameters at 4-bit, 2 to 3.5 GB downloads such as Qwen3.5 4B or Phi-4-mini. LM Studio tells 8 GB Mac owners to stick to smaller models and modest context sizes, and Haplo AI offers files up to 2.5 GB on them.

What is the best local LLM for a Mac?

The largest recent open model that fits your memory with room to spare. In October 2026: Qwen3.5 4B on 8 GB, Qwen3.5 9B or Gemma 4 12B on 16 GB, gpt-oss-20b on 24 GB, Qwen3.8 27B on 32 GB, and Llama 3.3 70B on 64 GB.

Is Ollama or LM Studio better on a Mac?

They run the same open models, so pick by how you work. Ollama is quickest from Terminal, serves other apps on port 11434 and also runs on Intel Macs (CPU only). LM Studio gives you a model browser and a chat window, needs Apple silicon, and has its own command-line tool and server.

Does a local LLM work without internet?

Yes, once the model is downloaded: all five ways run downloaded models on the Mac itself. For llama.cpp and MLX, --offline and HF_HUB_OFFLINE=1 also stop them checking Hugging Face for updates.

How much disk space do local LLMs need?

As much as the download: about half a gigabyte for a 0.8B model at 4-bit, 43 GB for Llama 3.3 70B. Ollama keeps models in ~/.ollama/models; our llama.cpp and MLX downloads went to ~/.cache/huggingface/hub.

How we tested this

We ran everything on an M2 MacBook Pro (10-core GPU, 24 GB) with macOS 27.0 on October 4, 2026:

  • Ollama 0.32.12 (Homebrew): the server with OLLAMA_NO_CLOUD=1, ollama run --verbose with Qwen2.5 3B, 7B and Coder 14B (already on the Mac), ollama ps, ollama ls and the README’s /api/chat request.
  • llama.cpp b9730 (Homebrew): llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF:Q4_0 (with -p and -st to script it), --reasoning off, --offline, llama serve with two API requests, llama-bench on four models, and verbose llama-completion logs for the KV cache sizes.
  • mlx-lm 0.32.0 in a throwaway virtual environment: mlx_lm.generate and mlx_lm.chat with Llama 3.2 1B 4-bit, again with HF_HUB_OFFLINE=1, and mx.device_info() for the GPU limit. We then deleted the environment and every model we had downloaded.

Quoted from the docs, not run: the install scripts, model tags we didn’t have, ollama rm, every LM Studio step (we didn’t install it), mlx_lm.server and the sysctl setting.

For Haplo AI, we read the source code (macOS build settings, model list requests, subscription screen, tools), requested the live model list as the Mac app does for ten memory sizes from 8 to 128 GB, and took the price, requirements and privacy label from the App Store listing. We didn’t open the app for this article. Model facts come from the makers’ Hugging Face pages and Ollama’s library, chip speeds from llama.cpp’s benchmark thread, and the photos from Wikimedia Commons.

References

Image credits