CappedAI runs open-source AI privately on your phone or drive — no internet required, no subscription, no server that sees your data.
Free app to start. Expand with knowledge packs that accumulate live data into permanent local knowledge. Or get a drive shipped pre-loaded and ready to go.
The open-source models are free. CappedAI gives you a private, offline way to run them — on whatever hardware you already have.
Knowledge packs are portable expertise modules — each with its own
agent, curated knowledge base, and pre-computed embeddings stored
on your drive as a .cap file.
Some packs can fetch live data when you're online — a stock pack pulls real-time quotes, a fantasy football pack gets this week's injury report, a weather pack fetches a current forecast.
But here's what makes CappedAI different: that live data doesn't disappear when the session ends. It gets embedded and written permanently into the pack's local knowledge base on your drive. The pack remembers what it learned.
On your next conversation — online or offline — the agent already knows what it fetched last time. The longer you use a pack, the smarter and more current it becomes.
Every other AI tool fetches live data and forgets it.
CappedAI packs remember.
Qwen 3, Llama 4, Phi-4, Mistral — open-source models now match GPT on most real-world tasks. The weights are free. You've been paying for someone else's server.
Drive owners browse and download from HuggingFace's full model catalog directly inside the app — chat models, image generators, embedding models. Downloaded to the drive, owned permanently. No API key. No billing relationship. No usage cap.
A 128 GB drive holds 15–20 models comfortably. When a better model ships, download it. The old one stays until you choose to remove it.
CappedAI uses Apple's Metal GPU framework for hardware-accelerated inference — the same GPU that powers your device handles the AI. No cloud. No external GPU. Just the silicon already in your hand.
| Platform | Chip | GPU / Inference | Framework | Status |
|---|---|---|---|---|
|
iPhone
iOS 16+
|
A16, A17 Pro, A18 |
Apple GPU · 5–6 core
Unified memory shared with CPU
|
Metal | ✓ Supported |
|
iPad
iPadOS 16+
|
M1, M2, M4 · A14+ |
Apple GPU · up to 10 core
M-series iPads match MacBook performance
|
Metal | ✓ Supported |
|
MacBook Air
macOS 13+
|
M1, M2, M3, M4 |
Apple GPU · 7–10 core
Fast enough for 7B–13B models comfortably
|
Metal | ✓ Supported |
|
MacBook Pro
macOS 13+
|
M1 Pro/Max, M2 Pro/Max, M3 Pro/Max, M4 Pro/Max |
Apple GPU · up to 40 core
70B+ models viable on Max chips
|
Metal | ✓ Supported |
|
Mac mini / Mac Studio / iMac
macOS 13+
|
M1–M4 · M2 Ultra / M3 Ultra |
Apple GPU · up to 76 core
Studio Ultra handles 70B+ at full quality
|
Metal | ✓ Supported |
|
Windows / Android
—
|
Various | No unified GPU framework | — | Not planned |
Apple's unified memory architecture means the GPU and CPU share the same memory pool — no data transfer bottleneck. This is why Apple Silicon runs large models faster than discrete GPU setups with the same theoretical FLOPS.
Every other AI tool processes your queries on a server somewhere. Policies change. Servers get breached. Companies get acquired.
CappedAI runs entirely on your own hardware. The models and your data live on your phone or drive — never on a server. Inference runs on your device's CPU or GPU. There is nothing in the cloud to breach, subpoena, or shut down.
When a pack fetches live data, that's an explicit network call made by the pack — and the result is written to your drive, not to our servers. Your conversation context never leaves your device.
Join the list for launch updates. iOS app and drives coming soon.
Leave your email and we'll reach out when the iOS app and drives are ready.
Interested in a drive? Just mention it in a reply — we'll be in touch.
Built in Austin, TX · CappedAI LLC