How to Run AI Offline on Your iPhone
What it takes to run a language model on the phone itself: real model sizes, which iPhones can hold them, speed and battery, what stays private, and where Apple's own model fits.
Most AI chat apps send what you type to a server and send the answer back. You can also run a language model on the iPhone itself, so it answers in Airplane Mode and your prompts never leave the phone. There are two ways to do that today: Apple’s own on-device model, which is part of Apple Intelligence, or an app that downloads an open model and runs it locally. We make Haplo AI, an app that downloads open models such as Qwen and Phi to your iPhone and runs them there, so it’s the second way (and it can use Apple’s model too).
The model sizes and memory rules below come straight from the catalog in Haplo AI’s code.
What does running AI offline on an iPhone mean?
A language model is a large file of numbers, called parameters or weights, plus software that uses them to write. It writes one token at a time, and each new token depends on everything before it, as Hugging Face’s documentation explains. A token is a small chunk of text: Apple puts it at roughly three to four characters of English.
Running it offline means both parts live on your phone. You download the model once. After that the phone’s own chip does the arithmetic and no server is involved, so it works in Airplane Mode, on a plane or in a basement.
The trade-off is size. The models that fit on a phone are small, and small models know less and get more wrong. Apple says as much about its own: its developer documentation points apps that need more reasoning or a longer context to Private Cloud Compute or a server model. For drafting, rewriting, summarising and general questions, a local model is often enough. For anything that matters, check the answer.
How much RAM and storage does an AI model need?
Two numbers decide whether a model fits: how many parameters it has, and how many bits each one takes. Multiply them and divide by eight to get bytes.
Models are usually published at 16 or 32 bits per parameter and compressed, or quantized, to fewer bits for phones. Every text model in Haplo AI uses a format called Q4_K_M, which averaged 4.89 bits per weight in llama.cpp’s own measurements on Llama 3.1 8B, against 16.0 for the uncompressed F16 version. That’s under a third of the size.
A worked example: Qwen3-4B has 4.0 billion parameters, according to its model card. At 4.89 bits each, that’s 4.0 × 4.89 ÷ 8 ≈ 2.4 GB, and the file Haplo AI downloads is 2.50 GB. At 16 bits the same model would be about 8 GB, the entire memory of an iPhone 16.

Storage and memory are separate limits. The download sits in storage, but to chat, the model is loaded into RAM. Apple’s researchers put it plainly: traditional language models “require all weights to reside in active memory (DRAM)”. And an app gets less RAM than the spec sheet says:
- iOS caps each app. The system terminates an app that crosses its per-process memory limit, according to Apple. Apps can ask for more with the increased memory limit entitlement, which Apple says only some device models offer. Haplo AI requests it, but no app gets the whole phone.
- The conversation takes memory too. The model keeps a cache of the conversation so far, and it grows with every token, as Hugging Face notes.
So Haplo AI’s catalog leaves headroom. In its code, models up to about 1.2 GB are offered on any iPhone, up to 2 GB need 6 GB of RAM, up to 2.5 GB need 8 GB, and up to 6 GB need 12 GB. A comment in the code explains why models just over 2.5 GB need a 12 GB phone: a 2.7 GB model plus its working memory got the app shut down on an 8 GB iPhone.
Here is every text model the current version offers, with its size as the app lists it and the smallest iPhone RAM it’s offered on:
| Model | Maker | Parameters | Size | iPhone RAM |
|---|---|---|---|---|
| SmolLM2-135M-Instruct | Hugging Face | 135 million | 101 MB | Any |
| SmolLM2-360M-Instruct | Hugging Face | 360 million | 258 MB | Any |
| Qwen2.5-0.5B-Instruct | Alibaba | 490 million | 379 MB | Any |
| Qwen3.5-0.8B | Alibaba | 0.8 billion | 530 MB | Any |
| SmolLM2-1.7B-Instruct | Hugging Face | 1.7 billion | 1.0 GB | Any |
| Qwen2.5-1.5B-Instruct | Alibaba | 1.54 billion | 1.0 GB | Any |
| Qwen3.5-2B | Alibaba | 2 billion | 1.3 GB | 6 GB |
| SmolLM3-3B-Instruct | Hugging Face | 3 billion | 2.0 GB | 6 GB |
| Ministral-3 3B | Mistral AI | 3 billion | 2.0 GB | 6 GB |
| Phi-4-mini, -instruct and -reasoning | Microsoft | 3.8 billion | 2.4 to 2.5 GB | 8 GB |
| Qwen3-4B | Alibaba | 4 billion | 2.5 GB | 8 GB |
| Qwen3.5-4B | Alibaba | 4 billion | 2.6 GB | 12 GB |
| Qwen3.5-9B | Alibaba | 9 billion | 5.3 GB | 12 GB |
| Qwen3.6-27B | Alibaba | 27 billion | 16 GB | Not on iPhone (needs 32 GB) |
The sizes are the app’s rounded labels, and a few files run larger (Qwen3.5-4B is 2.74 GB, Qwen3.5-9B 5.68 GB), so leave some spare space.
Which iPhones can run AI offline?
Apple doesn’t publish iPhone RAM, so the memory figures below come from MacRumors’ buyer’s guides. The last two columns show what Haplo AI offers at each tier and which model it recommends.
| RAM | iPhones | Text models offered | Recommended |
|---|---|---|---|
| 4 GB | iPhone 13, 13 mini, iPhone SE (2022) | The six models up to 1 GB | Qwen3.5-0.8B |
| 6 GB | iPhone 14, 14 Plus, 14 Pro, 14 Pro Max, iPhone 15, 15 Plus | Adds Qwen3.5-2B, SmolLM3-3B and Ministral-3 3B | Qwen3.5-2B |
| 8 GB | iPhone 15 Pro, 15 Pro Max, iPhone 16, 16 Plus, 16 Pro, 16 Pro Max, 16e, iPhone 17, 17e | Adds the Phi-4-mini models and Qwen3-4B | Qwen3-4B |
| 12 GB | iPhone 17 Pro, 17 Pro Max, iPhone Air | Adds Qwen3.5-4B and Qwen3.5-9B | Qwen3.5-4B |
A few things the table doesn’t show:
- You don’t need to look up your RAM. Haplo AI reads it and sends it with the catalog request. Models that won’t fit are dimmed, marked “Incompatible with device”, and have no download button.
- Images need less than you might expect. Stable Diffusion 2.1 Base (1.14 GB) and 1.5 (1.4 GB) are offered on any iPhone with 4 GB or more. The larger Z-Image-Turbo (6.5 GB) and ERNIE-Image-Turbo (5.9 GB) appear only from 8 GB, which MacRumors also calls the minimum for Apple Intelligence.
How fast is offline AI on an iPhone?
A reply comes in two phases. The model reads your prompt, which is quick, then writes the answer one token at a time, which is slower and limited mostly by memory. A study presented at MobiCom 2024, MELTing Point, found on-device inference “largely memory-bound”, and in its tests bigger models were slower and used more battery per token.
Some measured speeds:
- Apple’s 2024 on-device model (3 billion parameters, compressed to an average of 3.7 bits per weight) took about 0.6 milliseconds per prompt token before the first token appeared on an iPhone 15 Pro, then wrote 30 tokens per second, according to Apple.
- llama.cpp on an iPhone 14 Pro’s GPU, in the MELT tests: 24.7 tokens per second for a 1.1-billion-parameter model at 4 bits, 14.8 for a 3-billion-parameter model at 4 bits, and 6.0 for a 7-billion-parameter model at 3 bits.
At 15 tokens per second, a 300-token answer (roughly 900 to 1,200 characters of English, by Apple’s rule of thumb) takes about 20 seconds. Those tests used 2022 and 2023 phones, and we haven’t seen independent measurements for the newest ones.
What about battery and heat?
Running a model works a phone hard. In the MELT tests, iPhones drew up to 13.8 watts sustained, with peaks above 18 watts, and an iPhone 14 Pro’s surface reached 47.9°C after one conversation with a 3-billion-parameter model. Over 50 back-to-back prompts its speed started dropping straight away, with further drops around the 20th and 32nd prompts, which the authors think came from the phone switching power modes.
Each answer is cheap, though. The study measured 0.16 to 0.20 mWh per generated token on the iPhone 14 Pro and estimated that a full battery would cover about 490 to 590 short exchanges (40 tokens in, 135 out) with the phone doing nothing else, or about a fifth of a percent of the battery each. Short sessions barely register; long ones make the phone warm and slower. The smallest model that does the job is also the fastest and the coolest.
Is offline AI private?
With a cloud assistant, your prompt has to travel to someone else’s computer to be processed. With a local model it doesn’t: the prompt, the reply and any photo you ask about are processed on the phone.

Here is what we checked in Haplo AI’s code:
- Chats are saved in the app, on the phone.
- The app’s only calls to our server fetch the lists of available models. They send the iPhone model, the app’s version numbers and how much RAM the phone has. Your messages aren’t part of them.
- Model files download from Hugging Face, where the models are hosted.
- The app has no third-party analytics or advertising SDKs.
- Tools that reach the internet, such as Web Search and Browser, are off until you pick them from the Tools menu. Web Search sends your search query to a search engine (DuckDuckGo, with Bing as a fallback).
Apple Intelligence works the same way when it uses the on-device model. When a request is too much for the phone, Apple can send it to Private Cloud Compute, where Apple says user data is never stored or shared with anyone, including Apple. Its most capable server model now runs there on NVIDIA GPUs in Google Cloud, under the same guarantees, Apple says. Either way those requests go to a server, so they need a connection.
How does Apple’s on-device model fit in?
The third generation of Apple’s models for Apple Intelligence, announced in June 2026, includes two that run on the device: AFM 3 Core, the next version of Apple’s 3-billion-parameter dense model, and AFM 3 Core Advanced, a 20-billion-parameter model that activates only 1 to 4 billion parameters at a time. Core Advanced gets around the RAM limit by keeping the full model in flash storage and moving the parts a prompt needs into memory, and Apple says it’s “unlocked by and optimized for our most capable Apple silicon systems”. Three more models run on Private Cloud Compute, and Apple says it built the family in collaboration with Google.
To get it you need an Apple Intelligence iPhone on iOS 27: iPhone 15 Pro, 15 Pro Max, iPhone 16 models or later, or iPhone Air. It takes up to 8 GB of storage, or up to 14 GB on iPhone 17 Pro, 17 Pro Max and iPhone Air.
Other apps can use Apple’s model through the Foundation Models framework. Its on-device text model, SystemLanguageModel, changes with the operating system (Apple lists separate versions for iOS 26.0 to 26.3, 26.4 and 27.0), and Apple’s technote gives it a context window of 4,096 tokens per session, roughly 12,000 to 16,000 characters of English.
Which to use? Apple’s model comes with the system, so apps don’t have to download anything. Open models also run on 4 GB and 6 GB iPhones, let you choose the maker and size, and only change when you download a new one. You can have both: on an Apple Intelligence iPhone running iOS 26 or later, Haplo AI lists Apple’s model as “Apple Intelligence” at the top of its model picker, ready once Apple Intelligence is turned on.
How to run AI offline on iPhone with Haplo AI
Here’s the whole process in Haplo AI, checked against the app’s code. One thing to know first: it’s a subscription app, and the sign-up screen comes at the end of onboarding.
1. Download a model while you’re online
The first launch walks you through downloading a model. To add more later, tap the brain icon on the home screen for the full list, where one model is marked Recommended for your phone’s memory (the last column of the table above). Downloads come from Hugging Face and carry on in the background. Mobile data works, but Wi-Fi is kinder to your plan for a 2.5 GB file. Check your free space first in Settings > General > iPhone Storage, as Apple describes.
2. Go offline
Open Control Center and tap the Airplane Mode button, following Apple’s steps. Check that Wi-Fi is off too: you can still turn Wi-Fi on while Airplane Mode is on, and then the phone isn’t really offline.
3. Chat
Start a new chat and pick your model. Type as usual, and the reply streams in token by token as the phone writes it. Each conversation keeps its own model, and you can switch by tapping the model name under the chat’s title, so a small model can take quick questions and a bigger one the ones that matter.
4. Generate images
Tap Tools, choose Image Generation, and describe the picture. When you send it, the app shows its image models: Apple’s Core ML versions of Stable Diffusion 2.1 Base and 1.5, plus the two larger models on 8 GB phones. Speed depends on the phone. In Apple’s benchmarks of Stable Diffusion 2.1 Base at 512 × 512 pixels and 20 steps, one image took 7.9 seconds on an iPhone 14 Pro Max and 18.5 seconds on an iPhone 12 mini.

5. Ask about a photo
Tap the + button, then Add Images, pick a photo and ask your question. If the current model only reads text, the app opens the model list so you can switch to the vision model, SmolVLM2-500M from Hugging Face: a 437 MB model plus a 199 MB image encoder. The photo is read on the phone. Hugging Face has shown the 500M version of SmolVLM2 running completely locally in an iPhone app.
6. Tidy up
Models are big, so delete the ones you don’t use: swipe left on a downloaded model in the full list, or open Settings (the gear icon on the home screen) and go to Downloaded Models.
Offline AI on iPhone: quick answers
Can you run an LLM on an iPhone without internet?
Yes, once the model is downloaded. In Haplo AI, any iPhone with 4 GB of RAM can run models up to about 1 GB, 8 GB phones can run 4-billion-parameter models, and 12 GB phones go up to 9 billion.
Does Apple Intelligence work offline?
The on-device models do. Requests Apple sends to Private Cloud Compute run on Apple’s servers, so they need a connection.
How much storage does offline AI need?
In Haplo AI, from about 100 MB for the smallest text model to about 5.7 GB for the largest one offered on iPhone, plus 1.1 to 6.5 GB for each image model you add. Apple Intelligence takes up to 8 GB, or 14 GB on iPhone 17 Pro, 17 Pro Max and iPhone Air.
How we made this
The model names, parameter counts, sizes and memory tiers come from Haplo AI’s model catalog, the list the app downloads from our server. We called it the way the app does, once each for a 4, 6, 8 and 12 GB iPhone, and checked the download sizes against the files on Hugging Face. The steps and privacy notes come from the app’s source code, and the iPhone RAM figures from MacRumors, because Apple doesn’t publish them. Apple’s model details come from Apple’s research posts, developer documentation and support pages. The speed and battery figures come from Apple and the MELT study; we didn’t run our own speed or battery tests for this article. The photos are from Wikimedia Commons.
References
- Apple Machine Learning Research. Introducing the Third Generation of Apple’s Foundation Models. June 8, 2026.
- Apple Machine Learning Research. Introducing Apple’s On-Device and Server Foundation Models. June 10, 2024.
- Apple Developer Documentation. Foundation Models and SystemLanguageModel.
- Apple Developer. TN3193: Managing the on-device foundation model’s context window.
- Apple Support. How to get the next generation of Apple Intelligence.
- Apple Security Research. Private Cloud Compute: A new frontier for AI privacy in the cloud. June 10, 2024.
- Apple Developer Documentation. Identifying high-memory use with jetsam event reports and Increased Memory Limit entitlement.
- Apple Support. Use Airplane Mode and How to check the storage on your iPhone and iPad.
- Apple. ml-stable-diffusion: Stable Diffusion with Core ML on Apple Silicon.
- Laskaridis S, Katevas K, Minto L, Haddadi H. MELTing Point: Mobile Evaluation of Language Transformers. MobiCom 2024.
- llama.cpp. Quantize tool README.
- Hugging Face. How caching works, Transformers documentation.
- Qwen Team. Qwen3-4B model card.
- Hugging Face. SmolVLM2: Bringing Video Understanding to Every Device.
- MacRumors buyer’s guides for the iPhone 13, iPhone SE, iPhone 14, iPhone 14 Pro, iPhone 15, iPhone 15 Pro, iPhone 16, iPhone 16 Pro, iPhone 16e, iPhone 17, iPhone 17e, iPhone 17 Pro and iPhone Air.
Image credits
- A wing tip of an airplane · Photo: U.S. Department of Agriculture, Public domain (Resized)
- Pile of laptop and desktop RAM memory modules · Photo: Wilbysuffolk, CC BY-SA 4.0 (Resized)
- A view of the server room at The National Archives · Photo: The National Archives (UK), CC BY 3.0 (Resized)