cs.CR · 2026-06-22 · No. 31

Cryptography and Security, 2026-06-22.

7 new papers in cs.CR. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

7 entries
  1. 01

    Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

    Jun He, Deying Yu

    cs.CR · cs.AI · cs.DC · cs.LG

    Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. Existing access-control mechanisms authorize identities, while assurance layers certify proposed actions; neither alone provides a mandatory enforcement point for certified authority at the moment of mutation. This paper introduces the Sovereign...

    arxiv.org/abs/2606.20520 · PDF

  2. 02

    Efficient and Sound Probabilistic Verification for AI Agents

    Alaia Solko-Breslin, Pramod Kaushik Mudrakarta, Mihai Christodorescu, Somesh Jha, Krishnamurthy Dj Dvijotham

    cs.CR · cs.AI

    Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulate and enforce policies expressed in a formal language like Datalog offer a promising solution. However, existing approaches are restricted to deterministic policies. In many practical applications of AI agents, there is a need to enforce security policies in the face of ambiguity, leading to probabilistic...

    arxiv.org/abs/2606.20510 · PDF

  3. 03

    Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

    Arastoo Zibaeirad, Marco Vieira

    cs.CR · cs.AI · cs.SE

    Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved. We present CWE-Trace, a framework for LLM vulnerability detection built from 834 manually curated Linux kernel samples spanning 74 CWEs. The framework enforces a strict temporal split (pre-2025 historical set / post-cutoff leakage-free set), preserves context-aware vulnerable--patched pairs,...

    arxiv.org/abs/2606.20502 · PDF

  4. 04

    Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

    Reza Soosahabi, Vivek Namsani

    cs.CR · cs.AI

    Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic...

    arxiv.org/abs/2606.20470 · PDF

  5. 05

    Multi-View Decompilation for LLM-Based Malware Classification

    Bercan Turkmen, Vyas Raina

    cs.CR · cs.AI

    Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that large language models (LLMs) can assist this process by classifying decompiled code as benign or malicious, but existing pipelines typically rely on a single decompiler view. We argue that this assumption is fragile: decompilers are lossy heuristic tools, and different decompilers can expose different...

    arxiv.org/abs/2606.20436 · PDF

  6. 06

    LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

    Hanwool Lee, Dasol Choi, Bokyeong Kim, Seung Geun Kim, Haon Park

    cs.CR · cs.AI

    Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a...

    arxiv.org/abs/2606.20408 · PDF

  7. 07

    FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

    Chaeyun Kim, Daeyoung Park, Junghwan Kim, Jinyoung Jeong, Eunji Song, Yongtaek Lim, Minwoo Kim

    cs.CR · cs.AI

    Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU...

    arxiv.org/abs/2606.19887 · PDF

This edition is part of The Daily Abstract — cs.CR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.