cs.IR · 2026-07-08 · No. 47
Information Retrieval, 2026-07-08.
3 new papers in cs.IR. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
3 entries-
01
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
Riccardo Terrenzi, Serkan Ayvaz
cs.IR · cs.AI
Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. We study six metadata-generation settings for RDF datasets, ranging from simple rewriting to profile-grounded and agentic graph-based generation, and evaluate them jointly for retrieval effectiveness and faithfulness. Unconstrained metadata rewriting delivers the strongest retrieval gains over the original...
-
02
CMDR: Contextual Multimodal Document Retrieval
Ryota Tanaka, Taku Hasegawa, Kyosuke Nishida
cs.IR · cs.AI · cs.CL · cs.CV
Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document. However, existing benchmarks primarily evaluate simple lexical or semantic matching, and most methods encode pages independently. Consequently, they overlook the contextual information in the document required to resolve queries that aggregate information across multiple pages. In this paper, we introduce...
-
03
SCOReD: Student-Aware CoT Optimization for Recommendation Distillation
Haz Sameen Shahgir, Yufei Li, Frank Shyu, Luke Simon, Sandeep Pandey, Xi Liu, Yue Dong
cs.IR · cs.AI
Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approach the recommendation task with unusually high reasoning uncertainty, repeatedly rechecking their answers without revising them; supervised fine-tuning on such traces produces verbose students that never revise their initial guess. Furthermore, due to the novelty of...
This edition is part of The Daily Abstract — cs.IR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.