cs.OS · 2026-09-02 · No. 103
Operating Systems, 2026-09-02.
1 new papers in cs.OS. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
1 entries-
01
mzCache: On-Device LLM Memory Management under Multitasking
Hongseung Yu, Minsung Kim, Jongseok Park, Kyunghan Lee
cs.OS · cs.DC · cs.LG
On-device mobile Large Language Model (LLM) inference is gaining significant attention. However, mobile devices operate in highly dynamic multitasking environments where users frequently switch between applications. This creates memory pressure, forcing LLM memory (model weights and KV cache) to be evicted by the operating system. When a new inference request arrives, the inference system must restore the evicted memory through slow storage...
This edition is part of The Daily Abstract — cs.OS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.