cs.DB · 2026-07-27 · No. 66
Databases, 2026-07-27.
5 new papers in cs.DB. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
5 entries-
01
Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG
Chuangtao Ma, Arijit Khan
cs.DB · cs.AI
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG)...
-
02
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng
cs.DB · cs.AI · cs.CL · cs.LG
LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space scale and complexity (causal diagnosis across thousands of time series, business logs, and concurrent activity); solution-space openness (multiple remediations with...
-
03
Benchmarking Text-to-SQL under Role-Based Access Control
Yang Fei, Yangfan Jiang, Yin Yang, Xiaokui Xiao
cs.DB · cs.AI
Given a database S and a natural language question Q, text-to-SQL systems aim to generate an SQL query that correctly answers Q when executed against S. Currently, popular text-to-SQL benchmarks mostly assume unrestricted access to S; in practice, however, user access is often restricted, e.g., through role-based access control (RBAC) policies. This leads to a potential disconnect between benchmarking results and real-world performance: an...
-
04
MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery
Grace Fan, Eden Wu, Majid Daliri, Juliana Freire
cs.DB · cs.AI
Join discovery is a core task in dataset search, enabling users to find columns that can be joined with a given query column. Early approaches focused on equi-joins, but data lakes and open-data repositories often contain columns whose values refer to the same entity but use different syntactic representations. To address this challenge, recent approaches discover semantically joinable columns but face a fundamental trade-off: methods that...
-
05
Prompt as a Data Type: In-Database LLM Prompt Management and Rewriting
Denis Mayr Lima Martins, Gottfried Vossen
cs.DB · cs.LG
Large Language Models (LLMs) are increasingly used in database-backed applications to classify tuples, filter records using semantic predicates, extract structured attributes, and enrich query results. Yet the prompt that start these computations are typically stored outside the DBMS in unstructured formats, making them invisible to query execution, metadata management, and optimization. Drawing on Stonebraker's QUEL as a Data Type and the...
This edition is part of The Daily Abstract — cs.DB archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.