cs.DB · 2026-09-03 · No. 104

Databases, 2026-09-03.

3 new papers in cs.DB. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Poisoning Attacks on the PGM-index

    Atsuki Sato, Martin Aumüller, Yusuke Matsui

    cs.DB · cs.CR · cs.LG

    The PGM-index (Ferragina and Vinciguerra, VLDB'20) is one of the most practical learned indexes, owing to its theoretical elegance and consistently strong empirical performance. It is built on optimal piecewise linear approximations (PLAs) that minimize the number of segments. In this paper, we ask how sensitive this optimal PLA itself is to poisoning attacks. We propose PGM-attack, an efficient poisoning attack that sequentially inserts...

    arxiv.org/abs/2609.02328 · PDF

  2. 02

    A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search

    Sajad Faghfoor Maghrebi, Navid Eslami, Niv Dayan

    cs.DB · cs.AI · cs.IR · cs.LG

    Most vector databases rely on graph-based indexes, notably HNSW and Vamana, for approximate nearest neighbor search. With embedding models widely adopted, the datasets these databases store grow rapidly. At a fixed accuracy, how does search cost scale with dataset size? The prevailing answer is poly-logarithmic growth. Yet the claim is proven only under special conditions and asserted without proof for the indexes used in practice. It is also...

    arxiv.org/abs/2609.02143 · PDF

  3. 03

    Git4Data: Database-Native Version Control for AI Agents

    Hongshen Gou, Zuyu Zhang, Yuze Sun, Peng Xu, Feng Tian, Long Wang, Jianguo Wang

    cs.DB · cs.AI

    Large Language Model (LLM) agents increasingly explore many candidate states of relational data in parallel, each of which should remain isolated, reproducible, and auditable, preferably through the same SQL interface used for ordinary data work. Existing tools support this requirement only partially: source-code version control does not scale to large datasets, whereas relational databases manage large data efficiently but rarely expose...

    arxiv.org/abs/2609.02106 · PDF

This edition is part of The Daily Abstract — cs.DB archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.