cs.PF · 2026-08-31 · No. 101

Performance, 2026-08-31.

1 new papers in cs.PF. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

1 entries
  1. 01

    Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms

    Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton, John Paul Shen

    cs.PF · cs.DC · cs.LG

    Large language model (LLM) inference serving is priced by tokens, but GPU energy is consumed over inference windows. This accounting mismatch makes token-normalized metrics incomplete, since average output-token energy can decrease even when total request energy increases. We characterize this behavior with a decomposed energy model: a fixed one-time prefill with a fixed generation setup cost, while each output-token generation step adds...

    arxiv.org/abs/2608.28044 · PDF

This edition is part of The Daily Abstract — cs.PF archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.