cs.IR · 2026-08-03 · No. 73

Information Retrieval, 2026-08-03.

6 new papers in cs.IR. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    QASP: Query-Adaptive Robust Vector Search Policy

    Hakan Ferhatosmanoglu, Kushal Kumar, Tal Wagner, Andy Warfield

    cs.IR · cs.LG

    A fundamental challenge of vector search is achieving consistently high recall while minimizing computational costs. Fixed search parameters cause significant performance variance across queries, and conventional evaluation on average recall masks these per-query disparities. We introduce QASP (Query-Adaptive robust vector Search Policy), which predicts the complete recall progression curve per query via a single upfront supervised...

    arxiv.org/abs/2607.29606 · PDF

  2. 02

    RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

    Haoran Ling, Yuecheng Li, Zeyu Song, Jing Yao, Shuwen Kang, Chi Lu, Wenjin Wu, Peng Jiang

    cs.IR · cs.AI · cs.CL

    Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to both select modification directions and generate concrete hypotheses often leads to unstable search under limited experiment budgets. Inspired by the above challenge, we propose RecHarness, a Bandit-Routed...

    arxiv.org/abs/2607.29241 · PDF

  3. 03

    GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

    Jiping Liu, Zhongmin Zhang, Zisen Sang, Zhijia Fang, Tao Ouyang, Ma Jiang, Shaopeng Liang, Zeyang Hou, Guodong Cao, Jia Jia

    cs.IR · cs.LG

    Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text...

    arxiv.org/abs/2607.29213 · PDF

  4. 04

    PaletteID: Prototype-Composed Semantic Identifiers for Multimodal CTR Prediction

    Huanyu Liu, Baining Chen, Hui Liu, Zengyang Li, Ziyi Huang

    cs.IR · cs.LG

    Multimodal information can improve the accuracy of click-through rate (CTR) prediction and effectively alleviate item cold-start and long-tail problems. Recent studies commonly discretize pretrained multimodal embeddings into semantic identifiers (SIDs), allowing the model to learn task-specific semantic representations for recommendation. However, existing methods still provide limited gains due to two major limitations. First, codebook...

    arxiv.org/abs/2607.29000 · PDF

  5. 05

    Don't Contrast the Impossible: Region-Constrained Batching for Contrastive User Modeling on a Local Community Platform

    Seungho Han, Byeongchang Kim, Jin Yu

    cs.IR · cs.LG

    Contrastive learning is widely used for user modeling in large-scale recommender systems, where standard in-batch negatives implicitly assume universal exposure that any user can be shown any item. On local community platforms such as Karrot, however, exposure is geographically constrained; many user-item pairs are impossible by design yet still treated as negatives during training, diluting the contrastive learning signal. We address this...

    arxiv.org/abs/2607.28971 · PDF

  6. 06

    RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data

    Sidahmed Benabderrahmane, Talal Rahwan

    cs.IR · cs.AI · cs.LG

    Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jaccard, cosine, and Hamming compare objects through atomic overlap. IDF (Inverse document frequency) weighting partially reduces this effect but remains atom-wise and cannot explicitly represent informative higher-order co-occurrences. We introduce RareSense, a rarity-aware similarity framework for sparse...

    arxiv.org/abs/2607.28879 · PDF

This edition is part of The Daily Abstract — cs.IR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.