Part 1 - I Wanted the Model to Stay on the Phone
Beyond the Model: Building a Private AI Playground for Android & iOS
Part 1 - I Wanted the Model to Stay on the Phone
The engineering journey of integrating established on-device AI runtimes into one multimodal mobile product - and everything it takes to make the pieces behave like a coherent system
There was one constraint behind AI Playground that shaped almost every architectural decision that came after it:
I did not want the core AI workflow to depend on a server.
I want to be precise about what that means.
I did not build a new mobile inference engine. I did not invent a new way to execute transformer models on phones. The project stands on top of established work from the open-source and platform ecosystems - llama.cpp, GGML, Google AI Edge/LiteRT-LM, MediaPipe, Apple MLX, and other native technologies.
My job was different.
I wanted to take those building blocks and make them behave like one application.
That meant answering the questions that only appear once you move from “model demo” to “mobile product”:
Which runtime should handle this model?
Does this model actually support the requested input?
Can this device safely load it?
What happens when acceleration fails?
How do I keep native work away from the UI thread?
What happens if the user switches models mid-operation?
How do I make Android and iOS feel like themselves
without duplicating the entire application?
That is where the project became interesting.
1. The promise is simple. The engineering is not
From a user’s point of view, the experience sounds almost boring:
flowchart TD
A["Install application"] --> B["Download model"] --> C["Turn off network"] --> D["Start conversation"] --> E["Receive answer"]
The repository makes the offline-first boundary explicit.
The inference path is local:
flowchart TD
Prompt["Prompt"] --> App["Shared application logic"]
App --> Runtime["Selected local runtime"]
Runtime --> Native["Native inference"]
Native --> Response["Local response / history"]
The internet-enabled pieces are separated from inference:
flowchart TD
Network["Network"] --> Discovery["Model discovery"]
Network --> Download["Model download"]
Network --> Search["Live Hugging Face search"]
Network --> Other["Other explicitly allowed surfaces"]
That separation is more important than the phrase “offline”.
A product can call itself local while still quietly sending data through a logging service, a fallback endpoint, a remote parser, or a convenience API.
I wanted the architecture to make accidental data movement harder.
2. I needed Android and iOS to share the right things
The project uses Kotlin Multiplatform and Compose Multiplatform, but I never wanted “multiplatform” to mean “pretend both platforms are identical”.
The architecture is closer to:
flowchart TD
subgraph KMP["Shared KMP"]
Domain["Domain"]
Data["Data"]
end
KMP --> Boundaries["Runtime boundaries"]
Boundaries --> Android["Android
Kotlin / C++"]
Boundaries --> iOS["iOS
Swift / Metal"]
The shared layer is responsible for:
models
use cases
repositories
policies
state
feature contracts
Platform code owns things that actually depend on the operating system or native acceleration stack.
That is the line I keep coming back to:
Share the decision. Keep the hardware-specific execution native.
3. The first abstraction I cared about was the runtime boundary
Once the app had more than one way to execute a model, the UI could no longer be allowed to know all of them.
I did not want:
if (model.isLlamaCpp()) {
// llama-specific UI logic
} else if (model.isMediaPipe()) {
// MediaPipe-specific UI logic
} else if (model.isMLX()) {
// MLX-specific UI logic
}
That kind of condition spreads everywhere.
Instead, the project converged on a common runtime contract.
A simplified version is:
interface LocalModelRuntime {
val loadedModel: StateFlow<LoadedModel?>
suspend fun load(
model: LocalModelRef
): Result<ModelLoadInfo>
fun generate(
request: GenerationRequest
): Flow<GenerationEvent>
suspend fun cancel()
suspend fun unload()
}
The UI asks for:
load
generate
cancel
unload
The runtime implementation decides how those operations are actually performed.
4. That does not mean the runtimes are interchangeable
The abstraction hides implementation details.
It does not erase real capability differences.
The current project integrates runtime families including:
Android
-------
GGUF -> llama.cpp / GGML
.task / .litertlm -> LiteRT-LM / MediaPipe
LiteRT multi-graph -> image generation
stable-diffusion.cpp -> GGUF image models
iOS
---
llama.cpp / Metal path
Apple MLX
Core ML
The application has to know which path is appropriate.
That led to one of the most useful concepts in the project:
A model is not “supported” just because the file downloaded successfully.
5. Model support became a capability problem
I started thinking of a model as a tuple of properties:
flowchart TD
Model["Model"]
Model --> F["Format"]
Model --> A["Architecture"]
Model --> M["Modality"]
Model --> R["Runtime"]
Model --> P["Platform"]
Model --> C["Companion artifacts"]
Model --> Req["Minimum runtime requirements"]
Model --> Res["Resource profile"]
A model can be:
GGUF = yes
Android = yes
llama.cpp = yes
vision = no
or:
.task = yes
MediaPipe runtime = yes
text = yes
image input = no
Or even:
architecture supported
runtime available
companion file missing
The last case is particularly important.
The UI needs to know before it sends the user into a native error path.
6. Capability detection belongs below the UI
The application has Android-side inspectors for GGUF and LiteRT-related artifacts.
The broader architecture looks like:
flowchart TD
Artifact["Model artifact"] --> Inspection["Binary / metadata inspection"]
Inspection --> Report["Capability report"]
Report --> Resolver["Runtime resolver"]
Resolver --> UI["UI"]
A simplified data structure:
data class ModelCapabilityReport(
val textInput: Boolean,
val imageInput: Boolean,
val audioInput: Boolean,
val supportedRuntimes: Set<Backend>,
val platformSupported: Boolean,
val missingArtifacts: List<String>
)
The actual project goes deeper than this example, but the principle is the important part.
Capability is data.
It should not be duplicated as assumptions in five screens.
7. Memory turned out to be the next abstraction
The next surprise was how misleading model file size can be.
If a model artifact is:
2 GB on disk
the application should not make a decision from:
if (fileSize < availableRam)
The runtime also needs:
weights
+ KV cache
+ native overhead
+ temporary allocations
+ accelerator buffers
+ multimodal components
+ application pressure
So the application has a model-memory estimation layer.
Conceptually:
estimatedPeakBytes =
weightBytes
+ kvCacheBytes
+ runtimeOverheadBytes
+ modalityOverheadBytes
+ safetyMarginBytes
This is an admission heuristic, not a promise that the native allocator will use exactly that amount.
The job of the admission layer is:
Do I have enough evidence to start this expensive operation safely?
8. mmap changed how I thought about model files
For large native model artifacts, memory mapping is useful because it changes the relationship between:
storage
virtual address space
resident memory
The conceptual difference is:
flowchart LR
subgraph ReadAll["Read-all approach"]
direction TB
F1["file"] --> B1["large memory buffer"]
end
subgraph MMap["mmap approach"]
direction TB
F2["file"] --> R2["mapped address range"]
R2 --> P2["pages become resident as needed"]
end
This does not make memory pressure disappear.
It does make it possible to reason more accurately about file size versus actual resident pages.
9. GPU acceleration is another integration layer
The Android GGML integration uses a modular backend layout.
The project packages a baseline CPU/native stack and loads acceleration backends dynamically.
The intended backend order is:
flowchart TD
Vulkan["Vulkan
Preferred accelerated path"]
OpenCL["OpenCL
Fallback accelerated path"]
CPU["CPU
Final stable fallback"]
Vulkan -->|fallback| OpenCL
OpenCL -->|fallback| CPU
That is a project-level decision.
It is not an attempt to replace the GPU runtimes themselves.
The useful engineering work is in deciding:
when to use
when to fall back
how to detect failure
how to keep the application alive
10. Why I classified failures
This became important very quickly.
Consider these two errors:
GPU initialization failed
and:
not enough memory
They both mean:
load failed
but they are not the same problem.
The first may reasonably trigger:
Vulkan -> OpenCL -> CPU
The second often should not.
If RAM is the limiting factor, trying a different GPU backend does not magically create RAM.
So the routing logic is intentionally reason-aware:
for (backend in candidateBackends) {
val result = runtimeFor(backend).load(model)
if (result.isSuccess) {
return result
}
if (result.isInsufficientMemory()) {
return result
}
}
return lastFailure
That small distinction eliminates a surprising amount of pointless retry behavior.
11. Native code is a boundary, not a magic escape hatch
The Android app has native C++ integration around llama.cpp and a separate native image-generation module.
The conceptual stack is:
flowchart TD
Kotlin["Kotlin"] --> JNI["JNI"]
JNI --> CPP["C / C++"]
CPP --> Runtime["Third-party inference runtime"]
CPP --> Backend["Accelerator backend"]
Once a feature crosses JNI, I need to think about:
ownership
lifecycle
threading
cancellation
error mapping
resource cleanup
process crashes
I learned to treat every native context like a resource with an explicit lifecycle:
flowchart TD
Load["load"] --> Resident["resident"]
Resident --> Generate["generate"]
Resident --> Cancel["cancel"]
Resident --> Unload["unload"]
The UI should never own that lifecycle directly.
12. Why I kept heavy work away from the main thread
The project has an explicit concurrency model:
flowchart TD
UI["UI"] --> Store["MVI Store"]
Store --> UseCase["Use Case"]
UseCase --> Repo["Repository / Runtime"]
Repo --> IO["Dispatchers.IO"]
Repo --> Default["Dispatchers.Default"]
Heavy work includes:
file reads
SHA-256 verification
model inspection
database queries
document parsing
native inference dispatch
image decoding
generation orchestration
A simple example:
viewModelScope.launch {
val result = withContext(Dispatchers.IO) {
repository.inspectModel(modelPath)
}
state.update {
it.copy(model = result)
}
}
The point is not “always use IO”.
The point is:
Do not make the UI thread your accidental systems-processing thread.
13. Streaming needs pacing too
A native model can produce a lot of tiny generation events.
The UI does not need to render each event as a separate expensive state transition.
The project keeps a dedicated generation coordinator that coalesces updates at roughly a 65 ms UI cadence and uses longer watchdog windows for stalled generation.
Conceptually:
flowchart TD
Events["Native events"] --> Coalesce["Coalesce (~65ms)"]
Coalesce --> StateFlow["StateFlow"]
StateFlow --> Compose["Compose"]
That is a small detail with a large effect on perceived performance.
14. I wanted failure to be understandable
One of the first product lessons was that:
"Load failed"
is almost useless.
A better message tells the user what happened:
This model is estimated to exceed the device's safe
memory budget at the selected context length.
Try:
- a smaller quantization
- a shorter context
- unloading the current model
The engineering policy becomes part of the UX.
That is something I want AI applications to do more often.
15. What the first version changed in my thinking
I started with:
chat screen
+
local model
I ended up with:
chat
+
runtime contracts
+
model capability inspection
+
memory admission
+
hardware diagnostics
+
native boundaries
+
persistence
+
MVI
+
KMP
The model was still the important dependency.
But the application around it had become the real system.
And that led to the next question:
What happens when the input is not text?
A PDF.
A spreadsheet.
A screenshot.
A photo.
That is where the next phase started.
Build this yourself
Project 1 - Offline Pocket LLM
Do not start with a giant model.
Start with:
one GGUF
one runtime
one chat screen
local persistence
streaming
cancel
unload
Instrument:
time to first token
tokens / second
load time
peak process memory
UI frame stability
Run the same test on CPU and accelerated paths.
The exercise is not to beat a server.
It is to understand what your device is actually doing.
Project 2 - Capability-aware model catalog
Build a model details screen that answers:
Can this model:
run on this device?
use this runtime?
accept this input?
fit the current memory budget?
find all required artifacts?
Make the UI derive its attachment controls from this report.
That single project teaches an important product lesson:
The UI should reflect runtime truth, not just model metadata.
What I want to measure next
I still want better answers for:
How should thermal state affect runtime choice?
How much should battery state affect model selection?
How accurate can memory admission become without being expensive?
At what point is a smaller model with smarter retrieval
better than a larger model with a giant context?
Those questions belong in the next layers of the platform.
And that is exactly what happened.
Android iOS Kotlin Multiplatform On-Device AI LLM llama.cpp MediaPipe
1991 Words
2026-10-05 04:30 (Last updated: 2026-10-05 10:56)
e100978 @ 2026-10-05