cs.IR · 2026-06-18 · No. 27

Information Retrieval, 2026-06-18.

3 new papers in cs.IR. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    SAERec: Constructing Fine-grained Interpretable Intents Priors via Sparse Autoencoders for Recommendation

    Jiangnan Xia, Xuansheng Wu, Yu Yang, Xin Wang, Ninghao Liu

    cs.IR · cs.AI

    Intent-based recommender systems have gained significant attention for improving accuracy and interpretability by modeling the underlying motivations behind user behaviors. Most existing models derive intents directly from user sequences via clustering or prototype learning. However, they are sensitive to sequence quality, require presetting the number of intents, and lack explicit semantic grounding. These issues lead to an incomplete and...

    arxiv.org/abs/2606.18897 · PDF

  2. 02

    Rescaling MLM-Head for Neural Sparse Retrieval

    Youngjoon Jang, Seongtae Hong, Jonah Turner, Heuiseok Lim

    cs.IR · cs.AI

    Learned sparse retrieval (LSR) models such as SPLADE have traditionally used BERT-style masked language models as backbone encoders. A natural expectation is that replacing BERT with stronger pretrained encoders should improve retrieval effectiveness. However, we find that under standard SPLADE training recipes, backbones with large MLM-head L2 norms can suffer performance degradation and even training collapse under standard SPLADE training...

    arxiv.org/abs/2606.18811 · PDF

  3. 03

    SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval

    Youngjoon Jang, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim

    cs.IR · cs.AI

    With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology for global information access. MLIR enables users to retrieve semantically relevant documents from multilingual text collections using a single-language query. However, recent multilingual dense retrieval models often exhibit a strong preference for documents in the same language as the query. This leads to...

    arxiv.org/abs/2606.18801 · PDF

This edition is part of The Daily Abstract — cs.IR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.