cs.AI · 2026-09-21 · No. 120

Artificial Intelligence, 2026-09-21.

42 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

42 entries
  1. 01

    Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

    Hongyang Du, Lan Yan, Christian Flores, Asim Kadav

    cs.AI · cs.CV

    Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design...

    arxiv.org/abs/2609.22086 · PDF

  2. 02

    CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

    Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li,...

    cs.AI

    Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases...

    arxiv.org/abs/2609.22068 · PDF

  3. 03

    A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

    Hiskias Dingeto

    cs.AI

    Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic method that identifies guilty knowledge by presenting a suspect with the true detail among plausible decoys and measuring a stronger response to...

    arxiv.org/abs/2609.21996 · PDF

  4. 04

    Learning Cardiac Features: ECG Biometrics Across Time and~Exercise

    Luca Thiebaud, Paul Chauchat, Mustapha Ouladsine, Stéphane Delliaux

    cs.AI · q-bio.TO

    Electrocardiograms (ECGs) carry subject-specific patterns enabling reliable individual discrimination, forming the basis of ECG biometrics. Beyond authentication, this paradigm holds significant potential to secure sensitive cardiac data and to serve as a pretext task in self-supervised learning. Yet, most studies remain confined to singlesession, resting data, leaving robustness to temporal and physiological variations largely untested. We...

    arxiv.org/abs/2609.21962 · PDF

  5. 05

    AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

    Zijie Cao, Xijun Qu, Zhicheng Gu, Xiaoshu Chen, Duanyang Yuan, Yanning Hou, Sihang Zhou, Jianxing Gong, Jian Huang, Yang Mei

    cs.AI

    Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions. Existing memory systems typically rely on fixed granularities or static schemas, but these designs struggle when heterogeneous information, such as preferences, events, constraints, and temporal updates, is embedded in a single mixed representation. The resulting semantic interference makes top-K...

    arxiv.org/abs/2609.21940 · PDF

  6. 06

    What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

    Lyucheng Qian, John Yuehan Zhang, Pingyu Wang

    cs.AI

    Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs, although question value is determined by the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and...

    arxiv.org/abs/2609.21924 · PDF

  7. 07

    AutoRecLab: Describe the Experiment, Get the Code!

    Moritz Baumgart, Philipp Meister, Justus Krell, Michael Schmidt, Bela Gipp, Joeran Beel

    cs.AI · cs.IR · cs.LG

    Empirical evaluation is central to recommender-systems (RecSys) research, but turning experimental designs into executable code remains a manual and error-prone task. We present AutoRecLab, a Python-based autonomous RecSys lab that automates RecSys experiments from natural-language prompts. Given a research idea, AutoRecLab derives explicit experiment requirements, builds and validates a prototype, and iteratively expands it into the...

    arxiv.org/abs/2609.21863 · PDF

  8. 08

    EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

    Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid

    cs.AI · stat.ML

    Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment...

    arxiv.org/abs/2609.21841 · PDF

  9. 09

    MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention

    Muhammet Sami Yavuz, Sabri Mustafa Kahya, Richard R. Chen, Jana Lipkova, Benedikt Wiestler

    cs.AI · cs.CV

    Multimodal survival models can combine complementary prognostic information from whole-slide images and genomic profiles, but effective fusion remains challenging amid external cohort shift and computational complexity. To address these challenges, we propose MIST, multimodal survival prediction with genomic-guided histology attention. MIST represents genomic features as tokens and allows them to query compact foundation-model-derived...

    arxiv.org/abs/2609.21811 · PDF

  10. 10

    LLM-Generated Feature Pools for Time Series Anomaly Detection

    Youssef Attia El Hili, Malik Tiomoko, Corinne Ancourt

    cs.AI

    We study how far a simple statistical pipeline can go on univariate time series anomaly detection under a strict selection protocol. The method extracts a small pool of statistics over sliding windows, scores each window with a transductive robust (MAD) model, and selects a feature subset per domain on a held-out tuning split. On TSB-AD-U it reaches $0.529$ per-series VUS-PR, above the best neural ($0.45$) and statistical ($0.44$) entries on...

    arxiv.org/abs/2609.21801 · PDF

  11. 11

    ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction

    Jinning Liang, Mingcheng Zhu, Tingting Zhu

    cs.AI

    Emergency department (ED) decision-making relies on heterogeneous clinical information, including patient history, vital signs, laboratory results, and electrocardiograms (ECGs). Vision--language models (VLMs) can jointly process these modalities, but strong predictive performance does not necessarily imply meaningful use of the correct patient's ECG. We term this failure mode ECG Mirage: apparent multimodal capability without useful...

    arxiv.org/abs/2609.21755 · PDF

  12. 12

    World Modeling in Transformers

    Pierre Beckmann, Matthieu Queloz, Andre Freitas

    cs.AI · cs.CL

    Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position,...

    arxiv.org/abs/2609.21748 · PDF

  13. 13

    Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

    Yunji Chu

    cs.AI · cs.CL · cs.CV · cs.SD

    Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains...

    arxiv.org/abs/2609.21683 · PDF

  14. 14

    GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

    Zeyu Yan, Guanghao Zhou, Minghui Qiu, Ming Gao, Cen Chen

    cs.AI

    Recent advances in large reasoning models (LRMs) have made machine unlearning more challenging, as protected facts or unsafe rationales may surface in intermediate chain-of-thought (CoT) traces before the final answer is produced. Existing unlearning objectives typically suppress the target content or redirect internal representations, but they never specify how the post-forgetting trajectory should continue, which can lead to hallucinated...

    arxiv.org/abs/2609.21677 · PDF

  15. 15

    Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

    Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen

    cs.AI · cs.CL

    Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion...

    arxiv.org/abs/2609.21672 · PDF

  16. 16

    One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction

    Hongliang Li, Lu Wang, Yong Xu, Hanyang Chen, Zhitao Hou, Xiaoting Qin, Song Ge, Qingwei Lin, Dongmei Zhang

    cs.AI

    Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and...

    arxiv.org/abs/2609.21626 · PDF

  17. 17

    Calibrating Teacher--Student Discrepancy for On-Policy Distillation

    Qiangqiang He, Jin Li, MingCai Chen

    cs.AI

    On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during...

    arxiv.org/abs/2609.21619 · PDF

  18. 18

    Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education

    Andy Gray, Jake Hobbs

    cs.AI · cs.CY · cs.HC

    Access to academic support is a key determinant of student success, yet students experience it unequally: some readily seek help from lecturers or tutors, while others hesitate due to anxiety, fear of judgement, uncertainty about expectations, or low confidence in their understanding. This may be especially evident in computing education, where programming tasks are cumulative and cognitively demanding. Although students increasingly turn to...

    arxiv.org/abs/2609.21600 · PDF

  19. 19

    Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection

    Syed Ali Ahmed, Malaika Raza, Muhammad Shoaib Siddiqui, Muhammad Rafi

    cs.AI

    Fraudulent job posting detection aims to identify job advertisements that are corrupted either through fake content, misleading information, or negative intent, disrupting the online eco-system of job-seekers and employers. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent-space representations that capture subtleties among fake posts. To this end, we propose...

    arxiv.org/abs/2609.21599 · PDF

  20. 20

    Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

    Shuiying Liao, P. Y. Mok

    cs.AI · cs.LG

    Sequential recommendation aims to predict the next item a user will interact with based on their historical behavior. Advances in Transformers have significantly improved sequential recommendation but are still limited by cost efficiency. Although State Space Models (SSMs) have recently enabled efficient long-range modeling, most existing methods encode each item with a single static contextual role, overlooking the phenomenon of item...

    arxiv.org/abs/2609.21548 · PDF

  21. 21

    Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks

    Giambattista Amati, Federica Mangiatordi, Pierpaolo Salvo, Emiliano Pallotti, Simone Angelini

    cs.AI

    Artificial Intelligence (AI) is becoming a fundamental design principle of future AI-native communication networks, enabling autonomous resource management, adaptive control, and zero-touch network operation. While current AI-native architectures increasingly embed intelligence across network functions, they provide little guidance on how optimisation knowledge should be systematically generated, transferred, and exploited by AI models. This...

    arxiv.org/abs/2609.21519 · PDF

  22. 22

    The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

    Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella, Samy Bengio

    cs.AI

    When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the...

    arxiv.org/abs/2609.21509 · PDF

  23. 23

    PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

    Zicheng Zhao, Dongyin Chen, Rui Xu, Yinghui Xu

    cs.AI

    Multimodal large language models, or MLLMs, perform well at visual understanding and structured generation, yet these capabilities do not establish whether an engineering design will work when executed. Existing benchmarks assess spatial reasoning, structural validity, or physics-grounded construction, but they do not determine whether MLLMs can synthesize complete load-bearing structures and repair them after simulator execution exposes a...

    arxiv.org/abs/2609.21493 · PDF

  24. 24

    LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers

    Jingyu Hu, Shu Yang, Weiru Liu, Di Wang

    cs.AI · cs.LO · cs.SC

    Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely rely on outcome-based feedback, leaving the logical validity of intermediate reasoning steps largely unverified. To address the gap whereby LLMs arrive at correct final answers through logically flawed intermediate reasoning chains, we propose LogicTrack, a neuro-symbolic framework that audits...

    arxiv.org/abs/2609.21492 · PDF

  25. 25

    Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving

    Jiaxing Chen, Hengduo Zou, YuKai Qin, Yiren Zhao, Lidong Yu, Bolin Gao

    cs.AI

    Multimodal trajectory prediction improves behavioral coverage in end-to-end autonomous driving, but existing methods remain limited by sparse scene representations. Incomplete evidence leads to low-quality candidate generation and unreliable ranking among geometrically similar trajectories. On a register-based baseline, bad and poor candidates constitute 19.74% of the candidate set, while the oracle-best candidate ranks only 33.9th on...

    arxiv.org/abs/2609.21486 · PDF

  26. 26

    Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving

    Jiaxing Chen, Hengduo Zou, Yiren Zhao, Bolin Gao

    cs.AI

    Sparse representation formulates the environment perception for the end-to-end driving system as a set of discrete elements like objects and lane lines. This formulation meets safety risks in crowded, occluded scenes dealing with unstructured obstacles, uncertain regions, and intricate interactions. In this paper, we propose a dense representation, risk-aware occupancy, to characterize planning-relevant risks in an explicit and uniform...

    arxiv.org/abs/2609.21470 · PDF

  27. 27

    GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

    Kaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang, Yang Song, Dingqian Hong, Hui Xiong

    cs.AI · cs.LG

    Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that...

    arxiv.org/abs/2609.21432 · PDF

  28. 28

    DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

    Siyuan Liu, Fan Yu, Dongyu Ru, Yizhu Liu, Yifan Yang, Xuezhi Cao, Xunliang Cai, Yixin Cao

    cs.AI

    Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to distill these traces into reusable feedback without post-hoc outcome labels, drawing on their evidence of local progress, recovery, and unfinished requirements. We introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes this evidence into evidence-grounded nested...

    arxiv.org/abs/2609.21423 · PDF

  29. 29

    Offline Multimodal Large Language Models for Decision Support in Air Operations

    Joao P. A. Dantas, Jelton A. Cunha, Gabriel Dietzsch

    cs.AI · cs.CL

    Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal...

    arxiv.org/abs/2609.21390 · PDF

  30. 30

    LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

    Steve Drew, Jiayu Zhou

    cs.AI

    Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and...

    arxiv.org/abs/2609.21325 · PDF

  31. 31

    GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

    Xiuhui Zhang, Yi Chen, Shusheng Xu, Fan Li, Huan Wang, Tongkai Yang, Binhang Yuan

    cs.AI · cs.SE

    Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements. We introduce GameASG-Bench, a benchmark that makes behavioral testability part of the generation task for game development. Our design declares an evaluation interface specification before generation,...

    arxiv.org/abs/2609.21293 · PDF

  32. 32

    Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

    Yining She, Lei Lin

    cs.AI · cs.SE

    Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historical caching,...

    arxiv.org/abs/2609.21267 · PDF

  33. 33

    PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking

    Qiufeng Li, Chengxuan Wang, Rongqian Chen, Quan Cheng, Yihui Ren, Chia-Tung Ho, David Z. Pan, Tian Lan, Weidong Cao

    cs.AI

    Automated macro placement remains a fundamental challenge in VLSI physical design. Despite decades of research, existing approaches predominantly optimize hand-crafted proxy objectives, such as estimated wirelength, and typically produce placements through one-shot numerical optimization, limiting their ability to incorporate visual layout context, codified design expertise, and downstream physical-design feedback in a unified loop. We...

    arxiv.org/abs/2609.21263 · PDF

  34. 34

    CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

    Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M....

    cs.AI

    Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce...

    arxiv.org/abs/2609.21259 · PDF

  35. 35

    A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning

    Hongyan Wei, Wael AbdAlmageed

    cs.AI · cs.RO

    Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical rules. Conventional methods convert perception into discrete symbolic facts and then plan, discarding perceptual uncertainty and severing task-level feedback to perception. We introduce a generic, fully differentiable neuro-soft-symbolic framework that connects visual perception and task planning...

    arxiv.org/abs/2609.21221 · PDF

  36. 36

    Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis

    Boyuan Zhao, Meng Ye

    cs.AI

    Cognitive diagnosis infers students' concept mastery from response logs. However, students' responses are not determined by mastery alone: non-cognitive factors such as emotion, engagement, and fatigue can also affect performance. Affective cognitive diagnosis therefore extends conventional cognitive diagnosis by incorporating affective states. Existing methods often assume that the cognitive diagnosis backbone has already explained ability,...

    arxiv.org/abs/2609.21214 · PDF

  37. 37

    Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation

    Ana Nunez, Peyman Najafirad

    cs.AI

    Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but they suffer from two coupled pathologies: permissiveness collapse, where pass-rate rewards are maximised by trivial, non-discriminative tests, and concentration bias, where i.i.d. sampled tests cluster on modal inputs and inflate estimator variance. We introduce CoVer (Co-trained Coder and...

    arxiv.org/abs/2609.21208 · PDF

  38. 38

    AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture

    John Cuneo, David Chun, Gaurav Khanna

    cs.AI · cs.CY · cs.MA

    Organizations deploying agentic artificial intelligence must determine more than whether a model is trustworthy; they must establish what to validate, control, and observe for a use case to deliver its intended outcome while meeting applicable obligations. This paper proposes AI-GRACE (Agentic Intelligence-Governance, Risk, Assurance, Controls, and Evidence) as a use-case operationalization framework connecting organizational governance with...

    arxiv.org/abs/2609.21192 · PDF

  39. 39

    Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

    Adrien Deliège, Claas Beger, Marc Van Droogenbroeck, Melanie Mitchell

    cs.AI · cs.LG

    The Abstraction and Reasoning Corpus and related benchmarks evaluate whether AI models can solve novel reasoning tasks, but often leave unclear whether success reflects inference of the intended underlying rule or reliance on shortcuts. We address this gap by studying test-time task embeddings in Vision ARC (VARC), a model in which a pre-trained backbone is complemented by a trainable embedding representing the transformation rule. In the...

    arxiv.org/abs/2609.21181 · PDF

  40. 40

    SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity

    Thao Nguyen, Heng Ji

    cs.AI

    Off-target protein binding is a major source of adverse effects for small-molecule drugs, yet most structure-based molecular design methods focus on generating selective compounds de novo rather than improving the selectivity of existing, well- characterized drugs. We introduce specificity optimization (SpecOpt), a molecular design task that seeks constrained structural modifications to an existing compound that increase its binding...

    arxiv.org/abs/2609.21165 · PDF

  41. 41

    Can Agents Design Better Chips with a Higher Level Abstraction?

    Zijian Ding, Yang Zou, Yizhou Sun, Jason Cong

    cs.AI · cs.AR

    Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL. We ask whether agents can design better chips by leveraging higher-level abstractions. We compare Direct RTL Design, Agent-based HLS Design, Post-Compiler HLS Refinement, and Post-HLS RTL Refinement, and combine Agent-based HLS Design with Post-HLS RTL Refinement as Agent-based HLS with RTL Refinement...

    arxiv.org/abs/2609.21157 · PDF

  42. 42

    Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

    King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim, Peter Zandi, Jason Straub, Jacob...

    cs.AI · cs.CL

    Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these...

    arxiv.org/abs/2609.21149 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.