Answers that stream, offline
Tokens arrive as they are generated, exactly like a hosted assistant — except the generation is happening on your phone, in aeroplane mode if you like.
Llama, Qwen, Phi and more, running on the phone in your hand. No account, and the conversation never leaves it.
Llama, Qwen, Phi, Gemma, Mistral and SmolLM run on the device in your hand. Download a model once, then no account and no connection after that.
Tokens arrive as they are generated, exactly like a hosted assistant — except the generation is happening on your phone, in aeroplane mode if you like.
Download the ones you want and switch between them per conversation. A small model for quick answers, a larger one when it matters.
Models download once, then the work happens on your phone. Prompts and replies are never sent anywhere — there is no server to send them to.
Stable Diffusion runs on the device too — prompts and results stay in your camera roll.
Point it at a photo and ask. The vision model reads the image without uploading it.
Useful on a plane, in a basement, or anywhere you would rather not paste a work codebase into someone else’s server.
Native on iOS and macOS, sized to the silicon it finds — the Mac build will happily load models an iPhone would not attempt.
“Great offline interface for a wide range of LLMs, very slick and easy to use!”
“Loving this app, so convenient to get so many answers to a variety of topics at your fingertips!”
“Love being able to run these models locally. Pretty cool to use when I was on a flight with no wifi!”
“It’s so cool having AI that works offline and keeps everything private. This is just what I was looking for!”
“Genuinely so much better than the other apps. Love how the developer is already adding so much with each update!”
“Love how easy and accessible this app is to use! Definitely recommend for everyone in need of quick and convenient answers!!”