cs.CR · 2026-07-06 · No. 45
Cryptography and Security, 2026-07-06.
2 new papers in cs.CR. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
2 entries-
01
Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring
William Hackett, Peter Garraghan
cs.CR · cs.AI
As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical. Guardrail systems that detect and block malicious instructions sent to and from an LLM are an essential component of AI security. However, researchers conducting black-box adversarial emulation against production AI systems often struggle to determine whether a guardrail block or an LLM rejection...
-
02
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map
Gabriel Hurtado
cs.CR · cs.AI
Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: they score generations, not the artifact. We combine two cheap internal signals, a reference-anchored activation refusal-gap and a weight-recovery energy of the base-to-candidate weight difference, into a threshold-free checkpoint audit. The two are negatively correlated and label-complementary: the gap...
This edition is part of The Daily Abstract — cs.CR archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.