Preprint Characterizes Sustained Fine-Tuning of 3B LLM on Smartphone
arXiv preprint dated October 5 2026 presents the first systematic measurements of complete training runs for a 3B-parameter model on a mobile device, covering memory, time…

- An iPhone 17 Pro can fine-tune a 3B-parameter LLM to a typical user within one battery charge according to the abstract.
- Sustained training throttles throughput to about half the initial rate; tested pausing and burst schedules do not recover performance.
- Nearly all training time is spent in the frozen base model's backward pass, which most audited mobile runtimes do not accelerate.
Research Summary
The preprint by Geyko, Mosbach and Brinkmann provides the first systematic characterization of sustained fine-tuning of a multi-billion-parameter LLM directly on a phone. It measures memory consumption, per-step execution time, thermal behavior and energy use across full training runs rather than isolated steps.
The work also evaluates whether adapters trained on-device improve personalization and compares results to server-trained equivalents. Evidence is limited to the supplied abstract; the full paper is not available.
Key Performance Observations
According to the abstract an iPhone 17 Pro can complete the fine-tuning within one battery charge. Sustained training causes thermal throttling that reduces throughput to roughly half the starting rate. None of the pausing or burst schedules examined restored the original performance.
The authors report that nearly all computation occurs in the frozen base model's backward pass. They audited ten other mobile ML runtimes and found nine do not accelerate this operation.
Runtime Improvements
Analysis of Apple's MLX framework identified an unused kernel for the backward pass that was also incorrect. The researchers repaired the kernel; the fix has been merged upstream and is reported to train an adapter 1.47 times faster while consuming one-third less energy.
The preprint concludes that on-device fine-tuning is feasible on current phones but that significant further efficiency gains require treating training as a first-class workload in both runtimes and operating systems.
Evidence Limitations
This article is based solely on the arXiv abstract dated October 5 2026. No full paper, exact metrics, hardware configuration details or independent replication is available. Claims of exact equivalence in personalization quality or specific battery-charge completion therefore cannot be independently verified from the provided evidence.
Readers should await peer review and the complete manuscript before drawing firm conclusions about applicability across different models or devices.


