Paging the Experts: A Reproducible Characterization of Flash-Backed MoE Inference on iPhone
Organizations: Independent Researcher
Abstract
Sparse activation reduces mixture-of-experts computation without eliminating the need to store all experts. We present Routide, a Swift/MLX runtime that executes the text path of a pinned public Qwen3.6-35B-A3B quantized checkpoint while keeping expert weights in iPhone storage and a byte-budgeted subset in memory. We characterize cache-policy sensitivity, numerical comparison boundaries, and measurement limits. Across five recorded 128-token workloads, fixed-route replay gives 0.00% demand hits with a 512 MiB LRU cache, 18.80% with seeded random eviction at the same budget, and 38.58% with 576 MiB LRU. The apparent capacity cliff is therefore a policy/workload interaction, not a universal memory requirement. Same-runtime Mac controls preserve generated sequences across eviction and asynchronous prefetch, including 2,560 exact token comparisons and 10,334 speculative loads. In contrast, complete resident-Python versus recorded-phone sequences disagree on all five tested cases, precluding a general numerical equivalence claim. Two separately scoped iOS 27 memory protocols observe sampled process-footprint peaks of 1.87-2.32 GiB on short prompts and 2.39-2.73 GiB on one longer prompt. We retain a thermal stopping event, negative timing comparisons, and a single qualified whole-device power estimate. These results establish bounded feasibility and identify limitations that a deployment claim must not hide.
Figures & tables
| Quantity | Bytes | Interpretation |
|---|---|---|
| Checkpoint tensor payload | 20,401,929,952 | Includes the vision tower |
| Text-only resident-reference parameters | 19,508,787,456 | Retained language tensors, not process RAM |
| Excluded vision tensors | 893,142,496 | Explicitly excluded by the text loader |
| Routed-expert payload | 18,119,393,280 | All layers and experts |
| Non-expert checkpoint payload | 2,282,536,672 | Not an always-resident runtime set |
| Complete packed files | 20,403,825,982 | Includes manifest and alignment |
| Cohort | Workload and purpose | Platform | Scope |
|---|---|---|---|
| Initial capacity study | One-character prompt ( . ); 23 prepared tokens, 8 outputs; five cold/warm repetitions per budget | iPhone 17 Pro Max, iOS 26.5.2 | Narrow LRU calibration |
| Held-out routes | Five original -001 prompts, 128 outputs each | Same phone, iOS 26.6.1 | Measured routes; offline policy/prefetch replay |
| Counting diagnostics | Six balanced 128-token pairs; separately, four 512-token ABBA runs | Same phone, iOS 26.6.1 | Exploratory timing, traffic, sustained feasibility |
| Resident reference | Exact phone input IDs; five -001 cases | M1 Ultra, 128 GiB, macOS 26.6.2; Python MLX 0.31.2 | Cross-runtime diagnostic, not speed |
| Same-runtime controls | Five -001 cases; five additional -003 cases | Actual Swift model on the same high-memory Mac | Eviction/prefetch transparency |
| Process-memory parent | Three -001 prompts, two budget pairs each; stopped on request 12 | Phone, iOS 27.0; executable 9da46f6e... | Twelve retained rows, eleven fully valid |
| Budget (MiB) | Cold median elapsed (s) | Cold median decode (tokens/s) | Cold logical reads (GB) |
|---|---|---|---|
| 512 | 13.585 | 2.246 | 16.987 |
| 576 | 10.810 | 3.052 | 9.709 |
| 640 | 11.185 | 2.957 | 9.709 |
| 768 | 11.289 | 2.949 | 9.564 |
| 1024 | 11.137 | 3.173 | 8.025 |
| Policy | 512 MiB hits (%) | 576 MiB hits (%) | 1,024 MiB hits (%) |
|---|---|---|---|
| LRU | 0.00 | 38.58 | 47.59 |
| FIFO | 0.00 | 25.66 | 40.38 |
| LFU (per residency) | 0.00 | 1.82 | 7.84 |
| Recency-frequency hybrid | 0.00 | 14.65 | 49.83 |
| Random (five-seed mean) | 18.80 | 21.80 | 37.47 |
| Pinned farthest-next-use | 54.49 | 56.93 | 68.23 |
| Expert budget (MiB) | Prefetch | Exact sequences | Token comparisons | New speculative loads | Logical reads (GB) |
|---|---|---|---|---|---|
| 512 | none | 5/5 | 640 | 0 | 556.605 |
| 576 | none | 5/5 | 640 | 0 | 348.807 |
| 512 | preserve-resident, prefill 0.20 | 5/5 | 640 | 10,334 | 563.879 |
| 576 | preserve-resident, prefill 0.20 | 5/5 | 640 | 0 | 348.807 |
| Prompt | First different output position (1-based) | Fixed-history next-token matches |
|---|---|---|
| conversation-001 | 3 | 123/128 |
| code-001 | 83 | 127/128 |
| mathematics-001 | 73 | 126/128 |
| reasoning-001 | 2 | 127/128 |
| expository-001 | 1 | 121/128 |
| Order | Policy | Elapsed (s) | Decode (tokens/s) | Logical reads (GB) | Peak thermal |
|---|---|---|---|---|---|
| A1 | None | 314.105 | 1.795 | 319.354 | nominal |
| B1 | Prefill 0.20 | 272.711 | 2.085 | 320.297 | nominal |
| B2 | Prefill 0.20 | 301.553 | 1.864 | 320.297 | nominal |
| A2 | None | 304.148 | 1.852 | 319.354 | nominal |
| Prompt / pair | 512 MiB elapsed (s) | 576 MiB elapsed (s) | 576 vs 512 change (%) | Read reduction (%) | Both nominal? |
|---|---|---|---|---|---|
| conversation-001 / 1 | 83.782 | 74.811 | -10.71 | 40.10 | yes |
| conversation-001 / 2 | 75.045 | 62.730 | -16.41 | 40.10 | yes |
| code-001 / 1 | 84.173 | 70.838 | -15.84 | 38.16 | yes |
| code-001 / 2 | 76.458 | 81.946 | +7.18 | 38.16 | yes |
| mathematics-001 / 1 | 89.128 | 82.120 | -7.86 | 39.01 | yes |
| mathematics-001 / 2 | 86.120 | 76.707 | -10.93 | 39.01 | NO; retained stop row |
| Cohort / cache budget | Prompt/output tokens | Requests | Sampled footprint peaks (GiB) | Peak thermals |
|---|---|---|---|---|
| Stopped parent / 512 and 576 MiB | 49-69/128 | 12 | 1.87-2.32 | 11 nominal; 1 fair (stop) |
| Separate follow-up / 576 MiB | 340/32 | 1 | 2.39 | nominal |
| Separate follow-up / 512 MiB | 340/32 | 1 | 2.73 | nominal |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Statement supported here | Stronger statement not established |
|---|---|
| Text generation from the pinned flash-paged model on one iPhone | Universal device/model feasibility or state-of-the-art speed |
| Policy-dependent hits on five recorded traces | An unavoidable minimum RAM threshold or globally optimal 576 MiB budget |
| Exact tested outputs with the same Swift implementation under cache changes | General cross-framework or cross-device numerical equivalence |
| A fixed predictor’s selected-set overlap and read accounting | General latency hiding or energy savings |
| Sampled process footprint on separately identified protocols | Continuous device-wide memory maxima or context-length scaling |
| One whole-device power-estimator integral | Calibrated joules/token or app-only energy benefit |
| Protocol/run | Cache (MiB) | Prompt/output tokens | Start footprint (GiB) | Peak footprint (GiB) | Peak RSS (GiB) | Peak thermal |
|---|---|---|---|---|---|---|
| Parent/1 | 512 | 49/128 | 0.027 | 1.874 | 1.770 | nominal |
| Parent/2 | 576 | 49/128 | 1.854 | 2.152 | 2.070 | nominal |
| Parent/3 | 576 | 49/128 | 1.935 | 2.305 | 2.104 | nominal |
| Parent/4 | 512 | 49/128 | 1.918 | 2.226 | 2.102 | nominal |
| Parent/5 | 576 | 52/128 | 1.905 | 2.212 | 2.100 | nominal |
| Parent/6 | 512 | 52/128 | 1.913 | 2.221 | 2.102 | nominal |