cs.AR · 2026-07-27 · No. 66

Hardware Architecture, 2026-07-27.

2 new papers in cs.AR. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

2 entries
  1. 01

    HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

    Chao Fang, Jun Yin, Man Shi, Marian Verhelst

    cs.AR · cs.AI · cs.LG

    With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that exploits KV cache redundancy through hierarchical importance awareness. Algorithmically, HiKV compresses the KV cache at two granularities: Stage I evicts unimportant tokens within a fixed budget,...

    arxiv.org/abs/2607.22389 · PDF

  2. 02

    Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators

    Afzal Ahmad, Gaoyu Mao, Shoubo Hu, Hui-Ling Zhen, Mingxuan Yuan, Xinyu Chen, Wei Zhang

    cs.AR · cs.AI

    Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations. We observe that the task command, typically available before inference begins, provides a free signal that can be exploited to skip unnecessary computation at the hardware level. We present a HW/SW co-designed approach in which a lightweight...

    arxiv.org/abs/2607.22038 · PDF

This edition is part of The Daily Abstract — cs.AR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.