cs.IR · 2026-08-25 · No. 95

Information Retrieval, 2026-08-25.

6 new papers in cs.IR. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Adaptive Item-based Collaborative Structures via Noise Rescheduling in Diffusion for Generative Recommendation

    Jiaqi Wang, Tianying Liu, Heng Chang, Jihong Guan, Wengen Li, Shuigeng Zhou

    cs.IR · cs.AI

    Discrete Diffusion Models (DDMs) have recently been introduced to recommendation systems, modeling user history as a token generation process via iterative denoising. However, while effective at capturing user-level sequential patterns, these methods often fail to explicitly integrate item-based collaborative filtering information, a critical component for accurate recommendation. This deficiency manifests in two key aspects: (1) the item...

    arxiv.org/abs/2608.23400 · PDF

  2. 02

    Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

    Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang, Letian Gong, Baokun Wang, Huan Li, Yu Cheng, Weiqiang Wang

    cs.IR · cs.AI

    User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of...

    arxiv.org/abs/2608.23392 · PDF

  3. 03

    Hierarchical Exponential-Gaussian Mixtures for Watch-Time Distribution Prediction

    Sofia Gulevskaia, Mikhail Trapeznikov, Aleksandr Poslavsky, Alexander D'yakonov

    cs.IR · cs.LG · stat.ML

    Accurate watch-time (WT) prediction is an important requirement for short-video recommendations. Yet WT distributions are near-zero-inflated, long-tailed and multimodal. The recent Exponential-Gaussian Mixture Network (EGMN) models the full conditional WT distribution rather than a single point estimate and achieves state-of-the-art performance. Our large-scale reproduction study reveals that EGMN is vulnerable to variance collapse, component...

    arxiv.org/abs/2608.23356 · PDF

  4. 04

    Retrieval-Augmented Classification of Environmental Mitigations in Hydropower Licensing Documents

    Hong-Jun Yoon, Tom Ruggles, Joanna Lee, Debjani Singh

    cs.IR · cs.AI

    Identifying and classifying environmental mitigation obligations in Federal Energy Regulatory Commission hydropower licensing documents is a labor-intensive task requiring deep domain expertise. We formulate this as a multi-label classification problem over a structured 135-category taxonomy and address the central challenge of severe label scarcity: 40 of 135 categories have no training examples, and 26 have fewer than five. A supervised...

    arxiv.org/abs/2608.23241 · PDF

  5. 05

    Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision

    Yong-Hoon Choi, Youngjin Cho

    cs.IR · cs.LG

    Historical retrieval for time-series prediction commonly treats past similarity as a proxy for usefulness. We ask a different question: which historical examples should be expected to matter for a query? We define predictive relevance as expected future utility conditioned on inference-time information, using realized futures only during training as privileged supervision. A normalized-pattern retriever first forms a coarse candidate set, and...

    arxiv.org/abs/2608.23221 · PDF

  6. 06

    Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

    Kishore Konda

    cs.IR · cs.AI

    Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat an embedding as an indivisible point in a high-dimensional space. In this work, we propose the Hypergraph Embedding Index (HEI), a framework that instead organizes documents according to combinations of highly activated latent embedding dimensions. This formulation enables inverted-index style candidate...

    arxiv.org/abs/2608.22980 · PDF

This edition is part of The Daily Abstract — cs.IR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.