Part 3 - The Model Was Fine. The Phone Wasn’t.
The Model Was Fine. The Phone Wasn’t.
Part 3 - The difficult bugs were memory, native lifecycle, GPU drivers and operating-system behavior.
There is a point in a local AI project where the model stops being the hardest part.
For me, that point was when I started testing the application on actual phones instead of treating a successful development run as proof of correctness.
A model would load.
A prompt would work.
Then I would increase the context.
Add a document.
Switch the model.
Turn the screen off.
Run image generation.
And suddenly the “AI bug” turned out to be:
memory accounting
GPU initialization
native allocation
process lifecycle
stale state
background execution
That changed the question I was asking.
Not:
“Does this model work?”
But:
“When is it responsible to start this operation on this device?”
That distinction became one of the foundations of AI Playground.
1. Model size is not peak memory
Suppose a model artifact is:
3 GB
That number describes the artifact.
It does not describe the full runtime footprint.
The operation may need:
weights
+
KV cache
+
runtime state
+
temporary tensors
+
GPU buffers
+
image/vision components
+
application memory
+
allocator overhead
The project therefore maintains explicit memory estimation code and tests.
The conceptual model is:
estimated peak
=
weights
+
KV cache
+
runtime overhead
+
modality overhead
+
safety margin
This is deliberately a policy input, not a fake claim of perfect precision.
2. KV cache is why context settings matter
A user sees:
Context: 4096
as a text setting.
The runtime sees more work and more memory.
A simplified attention KV-cache relationship is roughly proportional to:
2
× layers
× KV heads
× head dimension
× context tokens
× bytes per element
The exact value depends on architecture and implementation.
The practical lesson is enough:
flowchart TD
Context["larger context"] --> KV["larger KV cache"]
Context --> Work["more prefill work"]
Context --> Latency["potentially higher latency"]
This is why context budgeting from Part 2 is also a memory feature.
3. I added an admission step before model loading
Instead of:
loadModel()
the application thinks in stages:
flowchart TD
Req["request"] --> Est["estimate"] --> Adm["admission"]
Adm --> Allow["ALLOW"]
Adm --> Warn["WARN"]
Adm --> Reclaim["RECLAIM"]
Adm --> Unload["UNLOAD_OTHER"]
Adm --> Deny["DENY"]
The admission result can explain itself:
data class AdmissionResult(
val decision: AdmissionDecision,
val estimatedPeakBytes: Long,
val safeAvailableBytes: Long,
val reason: String
)
That last field is important.
The user should not have to reverse-engineer a native exception.
4. Physical RAM is not the same as usable RAM
An “8 GB phone” does not mean my process gets 8 GB.
The runtime cares about:
physical RAM
available RAM
process RSS
native memory
current model residency
system pressure
The project therefore exposes a platform-neutral telemetry contract:
interface RuntimeMemoryTelemetry {
fun snapshot(): RuntimeMemorySnapshot
}
Android and iOS can then gather their own native signals without leaking platform details into shared domain code.
5. Memory pressure changes over time
A static preflight check is not enough.
Imagine:
T0
model fits
T1
user opens another app
T2
system pressure rises
T3
image preprocessing allocates buffers
T4
native backend reaches another allocation peak
So the project watches memory during long-running work as well.
The conceptual path is:
flowchart TD
Preflight["preflight"] --> Load["load"] --> Gen["generation"]
Gen --> Monitor["memory pressure monitor"]
Monitor --> Normal["normal"]
Monitor --> Warning["warning"]
Monitor --> Critical["critical"]
At higher pressure, the application can stop starting new heavy operations and reclaim disposable resources.
6. The stale-pressure bug was a good lesson
At one point, a critical memory-pressure signal could outlive the condition that caused it.
That meant the application could later observe:
live available memory: healthy
historical pressure: critical
and still refuse an operation.
The fix was a decaying memory-pressure monitor.
Conceptually:
flowchart TD
A["platform event"] --> B["remember severity"] --> C["time-based decay"] --> D["combine with live memory"] --> E["decision"]
I prefer that approach to simply ignoring pressure callbacks.
The historical signal still matters.
It just should not own the decision forever.
7. Another bug: counting the same memory twice
Image generation exposed a different class of accounting problem.
The application already knew the current image runtime had a resident footprint.
Another part of the decision logic counted the same weight footprint again.
The result was:
available memory
<
reported requirement
even though the resident model was already loaded and the request did not need to pay the full cost again.
The fix was to distinguish:
resident footprint
from:
incremental peak
That distinction is essential whenever multiple features share a runtime.
8. Switching models became a resource decision
Suppose:
chat model loaded
and then the user asks for image generation.
The application should not simply say:
memory too low
without considering whether the chat model can be unloaded safely.
A better decision tree is:
flowchart TD
Req["new request"] --> Est["estimate incremental cost"]
Est -->|fits with current residency| Run["run"]
Est -->|doesn't fit| Unload["unload reclaimable model"]
Unload --> Recheck["re-check"]
Recheck -->|fit| Run
Recheck -->|fail| Fail["fail"]
This is why the model lifecycle coordinator and memory admission policy are separate concepts.
One owns lifecycle.
The other decides whether the lifecycle transition is responsible.
9. GPU fallback is only useful if it is safe
The Android GGML integration uses a hierarchical backend approach:
flowchart TD
Vulkan["Vulkan"] -->|fallback| OpenCL["OpenCL"] -->|fallback| CPU["CPU"]
But fallback should not happen blindly.
The runtime distinguishes things like:
GPU initialization failure
unsupported backend
missing binary
native runtime error
insufficient memory
The first few may justify trying another backend.
The last one may not.
That is the difference between:
fallback
and:
retry until the app dies
10. Why the Vulkan probe runs in another process
This is one of the most unusual decisions in the project, and one of the most practical.
A GPU driver is native software outside my control.
A buggy vendor implementation can crash a process during probing.
If the probe happens in the main application process:
GPU probe crash
|
v
whole app dies
So the project isolates Vulkan probing in a child process:
flowchart TD
Parent["parent process"] -->|fork| Child["child process"]
Child --> Probe["Vulkan probe"]
Probe -->|success| OK["report OK"]
Probe -->|crash| Survives["parent survives"]
The point is not that Vulkan is “bad”.
The point is:
An optional acceleration path should not be able to take down the product while you are deciding whether it is available.
That is a systems-engineering decision around someone else’s runtime.
11. Native libraries need build-time guarantees too
The Android application packages native pieces for the selected runtime paths, including the llama/GGML stack and GPU backends, along with the separate image-generation JNI module.
The repository also contains scripts for engine setup, native builds, QA and invariant checks.
That matters because:
source code says backend exists
does not mean:
APK actually contains backend
Packaging is part of runtime correctness.
12. Background execution is another resource problem
A local generation can run longer than a user expects.
The user starts it.
Then:
screen off
or:
navigate away
The operating system is now part of the runtime.
The production Android architecture uses a foreground inference path for long-running inference, with the appropriate service type and a visible user-facing notification. The image-generation feature also supports session continuation while navigating between application surfaces.
The lifecycle looks conceptually like:
flowchart TD
Start["start"] --> Session["foreground inference session"] --> Work["native work"]
Work --> Cancel["cancel"]
Work --> Finish["finish"]
Cancel --> Release["release"]
Finish --> Release
That is not an implementation detail.
It is how I tell the operating system:
This work is important, visible and currently active.
13. Cancellation is a real contract
A button labeled:
Cancel
does not prove cancellation.
Real cancellation has to travel through the stack:
flowchart TD
UI["UI"] --> Store["Store"] --> Coord["Coordinator"] --> Cancel["runtime.cancel()"] --> Native["native cancellation"] --> Exit["native loop exits"]
The project has explicit tests around true cancellation and race conditions because stale cancellation state can otherwise poison the next generation.
A subtle bug looked like:
Run A
|
cancel
|
Run A fails
Run B
|
stale isCancelled = true
|
Run B immediately stops
The fix was to reset cancellation state at the start of every new run.
Simple.
Easy to miss.
Exactly the kind of thing that deserves a regression test.
14. Screen-awake state needs ownership too
The project uses a reference-counted screen-awake controller.
The logic is:
acquire
-> count + 1
release
-> count - 1
count > 0
-> keep awake
count == 0
-> restore normal behavior
Why?
Because multiple operations can overlap.
If operation A turns the state off when it completes while operation B is still active, B loses its protection.
Reference counting makes ownership explicit.
15. Trimming should target disposable memory first
When the system reports pressure, I do not want the first response to be:
unload the model
Model loading is expensive.
The better sequence is to reclaim things that are genuinely disposable:
image caches
temporary buffers
native allocator caches
short-lived objects
and preserve expensive resident state where the resource envelope still permits it.
This leads to a general principle:
Reclaim what is cheap to recreate before destroying what is expensive to rebuild.
16. Resource-management code needs real tests
The repository contains tests for things such as:
KV-cache estimation
model-memory estimation
memory-pressure classification
resident model footprint
model-load decisions
model switching
image memory admission
runtime memory snapshots
That is not accidental.
The memory layer can produce incorrect user-visible behavior even when the inference engine itself is perfectly healthy.
So it deserves unit-level coverage independent of native inference.
17. The hardware test matrix matters
I think about testing in dimensions:
flowchart TD
Hardware["Hardware Test Matrix"]
Hardware --> RAM["RAM Tier
(low / mid / high)"]
Hardware --> Model["Model
(text / vision / image)"]
Hardware --> Op["Operation
(load / generate / switch / background / cancel)"]
The question is:
Where are the boundaries?
Not:
“Does my phone work?”
That is how I want to collect real device knowledge over time.
Build this yourself
Project 1 - Mobile model admission library
Create:
data class DeviceMemory(
val physicalBytes: Long,
val availableBytes: Long,
val processRssBytes: Long
)
data class ModelRequirement(
val weightsBytes: Long,
val kvCacheBytes: Long,
val runtimeBytes: Long
)
and:
fun decide(
memory: DeviceMemory,
requirement: ModelRequirement
): AdmissionDecision
Then add:
resident model state
+
memory pressure
+
safety margin
+
reclaim strategy
Test the decision engine before connecting a real model.
Project 2 - GPU crash-isolated probe
Build:
flowchart TD
Main["main process"] -->|fork| Child["child process"]
Child --> Probe["GPU probe"]
Probe --> Report["report success / failure"]
Intentionally make the child fail.
Your parent should survive and choose a safe fallback.
Project 3 - Long-running local inference
Build a local inference session that survives:
screen off
activity recreation
navigation
cancellation
and records:
start
finish
duration
cancel latency
peak memory
completion state
That is a much more realistic mobile-AI exercise than another chat screen.
The systems lesson
By now, I had learned something I wish I had written down on day one:
Mobile AI reliability is mostly about making fewer bad assumptions.
Do not assume:
8 GB RAM = 8 GB available to me
Do not assume:
GPU available = GPU reliable
Do not assume:
model file size = peak memory
Do not assume:
cancel button = cancellation
Do not assume:
background process = process that will remain alive
And do not assume:
runtime failure = model failure
Once I started treating those as explicit contracts, the application became much easier to reason about.
Then I did something that added a completely different class of pressure to the system.
I tried to generate images locally.
Android Performance Vulkan Memory Management C++ On-Device AI
1910 Words
2026-10-05 06:30 (Last updated: 2026-10-05 10:56)
e100978 @ 2026-10-05