cs.DB · 2026-07-27 · No. 66

Databases, 2026-07-27.

5 new papers in cs.DB. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG

    Chuangtao Ma, Arijit Khan

    cs.DB · cs.AI

    Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG)...

    arxiv.org/abs/2607.22319 · PDF

  2. 02

    DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

    Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng

    cs.DB · cs.AI · cs.CL · cs.LG

    LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space scale and complexity (causal diagnosis across thousands of time series, business logs, and concurrent activity); solution-space openness (multiple remediations with...

    arxiv.org/abs/2607.22165 · PDF

  3. 03

    Benchmarking Text-to-SQL under Role-Based Access Control

    Yang Fei, Yangfan Jiang, Yin Yang, Xiaokui Xiao

    cs.DB · cs.AI

    Given a database S and a natural language question Q, text-to-SQL systems aim to generate an SQL query that correctly answers Q when executed against S. Currently, popular text-to-SQL benchmarks mostly assume unrestricted access to S; in practice, however, user access is often restricted, e.g., through role-based access control (RBAC) policies. This leads to a potential disconnect between benchmarking results and real-world performance: an...

    arxiv.org/abs/2607.22115 · PDF

  4. 04

    MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery

    Grace Fan, Eden Wu, Majid Daliri, Juliana Freire

    cs.DB · cs.AI

    Join discovery is a core task in dataset search, enabling users to find columns that can be joined with a given query column. Early approaches focused on equi-joins, but data lakes and open-data repositories often contain columns whose values refer to the same entity but use different syntactic representations. To address this challenge, recent approaches discover semantically joinable columns but face a fundamental trade-off: methods that...

    arxiv.org/abs/2607.21781 · PDF

  5. 05

    Prompt as a Data Type: In-Database LLM Prompt Management and Rewriting

    Denis Mayr Lima Martins, Gottfried Vossen

    cs.DB · cs.LG

    Large Language Models (LLMs) are increasingly used in database-backed applications to classify tuples, filter records using semantic predicates, extract structured attributes, and enrich query results. Yet the prompt that start these computations are typically stored outside the DBMS in unstructured formats, making them invisible to query execution, metadata management, and optimization. Drawing on Stonebraker's QUEL as a Data Type and the...

    arxiv.org/abs/2607.21756 · PDF

This edition is part of The Daily Abstract — cs.DB archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.