Edge0 Runs 35B AI Models on iPhone With Just 2.9 GiB Memory
Edge0 has demonstrated the ability to run 35B-parameter AI models on iPhones using only 2.9 GiB of memory, alongside similar breakthroughs streaming 8B models from SSDs with just 1 GiB. This represents a significant shift in on-device AI inference, enabling large language models to run locally without cloud dependency or massive resource requirements. The achievement challenges traditional assumptions about model size constraints on consumer hardware.
Why it matters
💻 Developer · Direct impact: you can now deploy 35B models to iOS without cloud backends. This means offline-capable AI apps, lower latency, and no user data leaving the device. The memory efficiency here—running 35B in 2.9 GiB—opens new architectures for local-first apps.
📦 Product · This unlocks new product categories: offline-capable AI assistants, privacy-first note-taking, on-device RAG systems. Users get faster responses and privacy guarantees. The constraint isn't model size anymore—it's UX design around local inference.
🎨 Design · Edge inference changes interaction design. No loading spinners waiting for API responses, instant offline access. Design for local-first workflows: emphasize privacy, speed, and persistent availability. Consider how offline-capable features reduce friction.
📈 Business · Lower infrastructure costs per user, better differentiation on privacy (Apple's narrative), new margins on edge-deployed models. Reduces cloud API dependency—meaningful for cost optimization and feature velocity.
🤔 Just Curious · This is the bridge between cloud and device intelligence. 35B on iPhone means sophisticated AI reasoning offline. It raises questions about future AI: Does everything eventually move to edge? What does centralized AI become?
Sources: Edge0 Runs a 35B AI Model on an iPhone With Just 2.9 GiB, Edge0 Streams an 8B AI Model From SSD Using Only 1 GiB