← Back to Index
Daily Research Digest

arXiv Papers

2026-09-04
347
Papers
8
Categories
79
Translated
收藏清单 0
精选 · Favorites
79
cs.AI / 1 / 2609.03402
A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant
一种提示工程方法,用于在通用人工智能教学助手中实现可扩展、灵活且实时的混合微观个性化
Saptarshi Basu, Sandeep Kakar, Ashok Goel
cs.AI
large language model
大语言模型相关
Abstract
Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This study presents a prompt-engineering-based framework for personalizing general-purpose LLM/RAG-based AI teaching assistants such as Jill Watson across academic disciplines and courses. The framework adapts responses using six learner-specific dimensions: self-assessment, abstraction preference, verbosity preference, perceptual orientation, information processing style, and level of understanding, yielding 96 distinct learner profiles. Student queries are additionally analyzed using Bloom's Taxonomy to estimate cognitive complexity at the interaction level. Learner attributes and cognitive assessments are encoded in structured prompts that condition the LLM without requiring model retraining. The framework is evaluated through experiments using NLP metrics and a human study with five participants. Results show perceived differences in response style and structure across personalization conditions, with statistical analyses identifying learner attributes associated with measurable response changes. These findings provide preliminary evidence that prompt-based personalization can support adaptive behavior in LLM-powered educational agents.
Chinese Translation
由大语言模型(LLM)驱动的人工智能(AI)教学助手提供了可扩展的教育支持,但通常所提供的个性化程度有限。本研究提出了一种基于提示工程的框架,用于对基于通用大语言模型/检索增强生成(RAG)的人工智能教学助手(如Jill Watson)进行跨学科和课程的个性化定制。该框架使用六个学习者特定维度来调整响应:自我评估、抽象偏好、简洁性偏好、感知取向、信息处理风格和理解水平,从而产生96种不同的学习者画像。此外,还使用布鲁姆分类法对学生查询进行分析,以在交互层面估计认知复杂度。学习者属性和认知评估被编码为结构化提示,在不要求模型重新训练的情况下对大语言模型进行条件约束。该框架通过使用自然语言处理指标的实验以及一项包含五名参与者的人体研究进行评估。结果显示,在不同个性化条件下,响应风格和结构存在可感知的差异,统计分析识别出了与可测量的响应变化相关联的学习者属性。这些发现提供了初步证据,表明基于提示的个性化能够支持由大语言模型驱动的教育代理中的自适应行为。
cs.AI / 2 / 2609.03407
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
困于故事:多轮大语言模型对话中的叙事俘获
Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang
cs.AI
large language model
大语言模型相关
Abstract
People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversation often elicits one party's self-justifying account, which can unfold over multiple turns and create information asymmetry. We introduce \textbf{narrative captivity}, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation without seeking missing perspectives. To measure this phenomenon, we build a benchmark of $5{,}078$ interpersonal-conflict scenarios spanning six moral dimensions. Across 17 LLMs, narrative captivity is widespread: end-state judgments under multi-turn narration shift by 25 percentage points on average beyond the matched single-turn baseline. Stage-level analysis identifies preference optimization as a major contributor, while four inference-time strategies provide only partial mitigation. We hope our project fosters LLM advisors that preserve independent judgment in real-world consultation.
Chinese Translation
人们越来越多地转向大型语言模型(LLM)寻求日常建议,这使得具有伦理争议的人际问题成为一种实际的道德咨询情境。大多数先前工作通过单轮判断或施加压力的反驳来研究这种情境,而这些假设与现实世界中人们寻求指导的方式并不相符。这些假设留下了一个尚无答案的问题:在没有明确对立立场的情况下,仅凭叙述本身是否能在多轮道德咨询中改变模型判断。然而,现实中的道德冲突对话常常会引出某一方自我辩护的叙述,这种叙述可在多个回合中展开并造成信息不对称。我们引入叙事俘获(narrative captivity):一种失败模式,其中模型将未经反驳的单方叙述视为完整,并在没有寻求缺失视角的情况下与叙述者的解读保持一致。为度量该现象,我们构建了一个包含 $5{,}078$ 个跨越六个道德维度的人际冲突场景的基准。在17个LLM中,叙事俘获广泛存在:与匹配的单轮基线相比,多轮叙述下的最终状态判断平均偏移了25个百分点。阶段级分析指出偏好优化是主要促成因素,而四种推理时策略只能提供部分缓解。我们希望我们的项目能培养出在真实咨询中保持独立判断的LLM顾问。
cs.AI / 3 / 2609.03527
NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis
NeoRed:一种面向新生儿呼吸系统疾病诊断的知识-逻辑-对齐多模态大语言模型
Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Jingkuan Song, Ye Luo
cs.AI
large language model
大语言模型相关
Abstract
Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain gap arising from predominantly adult training data; (2) insufficient integration of multidimensional clinical context for accurate diagnosis. To address these challenges, we collect two real-world clinical datasets (NeoCXR and NeoCXR-EV) and propose NeoRed, to the best of our knowledge, the first MLLM tailored for neonatal respiratory disease, filling the gap in neonatal diagnostic reports generation. To enhance joint diagnosis from heterogeneous clinical context and chest X-rays, we design a novel Knowledge-Logic-Alignment (KLA) framework which constrains model behavior from three perspectives: 1) Knowledge Prior Injection (KPI) incorporates neonatologist-inspired diagnostic priors into multimodal representations, guiding disease-specific attention across modalities; 2) Diagnostic Logic Constraint (DLC) aligns the semantics of generated reports with multimodal diagnostic logic; and 3) Visual Semantic Alignment (VSA) establishes semantic correspondence between visual features and imaging conclusions. Extensive experiments demonstrate that NeoRed enables accurate neonatal diagnostic reports generation, achieving ROUGE-L of 53.29% and Clinical Efficacy F1 score of 65.19% on NeoCXR, outperforming existing MLLMs. NeoRed also preserves competitive report generation performance on adult benchmarks (MIMIC-CXR and IU-Xray). Datasets will be available upon application.
Chinese Translation
新生儿呼吸系统疾病是新生儿发病和死亡的主要原因,在临床实践中构成重大挑战。尽管近期取得了进展,但现有的多模态大语言模型(MLLMs)在新生儿诊断方面面临两个关键局限性:(1)由于训练数据以成人为主要来源而产生的领域差距;(2)对多维临床背景信息整合不足,难以实现准确诊断。为解决这些挑战,我们收集了两个真实世界临床数据集(NeoCXR 和 NeoCXR-EV),并提出了 NeoRed,据我们所知,这是首个专门针对新生儿呼吸系统疾病的多模态大语言模型,填补了新生儿诊断报告生成方面的空白。为了增强异质临床背景与胸部X光片的联合诊断能力,我们设计了一种新颖的知识-逻辑-对齐(KLA)框架,从三个角度约束模型行为:1)知识先验注入(KPI)将受新生儿科医生启发的诊断先验融入多模态表示中,引导跨模态的疾病特定注意力;2)诊断逻辑约束(DLC)将生成报告的语义与多模态诊断逻辑对齐;3)视觉语义对齐(VSA)在视觉特征与影像结论之间建立语义对应关系。大量实验表明,NeoRed 能够实现准确的新生儿诊断报告生成,在 NeoCXR 上达到了 53.29% 的 ROUGE-L 和 65.19% 的临床效能 F1 分数,优于现有 MLLMs。NeoRed 在成人基准(MIMIC-CXR 和 IU-Xray)上也保持了具有竞争力的报告生成性能。数据集将在申请后提供。
cs.AI / 4 / 2609.03546
Dalek: A Constructive Agent Machine
Dalek:一种构造性智能体机器
Wanpeng Xie
cs.AI
large language model
大语言模型相关
Abstract
We present Dalek, a closed machine designed for agents that realizes self-maintenance, self-evolution, self-reproduction, and self-organization on any substrate satisfying a general host contract. The machine is built from three primitives---actors, messages, and channels. Four obligations---a host boundary, a construction language, admissible transitions, and rule heredity---give its boundary, identity, and closure a structural basis. Von Neumann's 1948 self-reproducing automaton supplies a hereditary constructional core: a self-description together with a constructor, a copier, and a controller. Dalek combines this core with the four obligations and rederives its medium for a text-and-message agent substrate, adding explicit structures for boundary, identity, history, and growth. A large language model and a compiler occupy the payload position and form a general capability producer. New capabilities are authored, compiled, installed into the description, and inherited by descendants. The same path produces the machine's own organs and even its runtime, closing heredity and evolution within the machine.
Chinese Translation
我们提出 Dalek,一种为智能体设计的封闭机器,它在满足通用宿主契约的任何基质上实现自维护、自进化、自复制和自组织。该机器由三个原语构建——行动者、消息和通道。四个义务——宿主边界、构造语言、可容许转移和规则遗传——为其边界、身份和闭合提供了结构基础。冯·诺依曼 1948 年的自复制自动机提供了一个可遗传的构造核心:一个自描述连同构造器、复制器和控制器。Dalek 将此核心与四个义务相结合,并针对一个文本与消息的智能体基质重新推导其媒介,添加了用于边界、身份、历史和生长的显式结构。大语言模型和编译器占据载荷位置,形成一个通用能力产生器。新能力被编写、编译、安装到描述中,并由后代继承。同样的路径产生机器自身的器官,甚至其运行时,从而在机器内部实现遗传和进化的闭合。
cs.AI / 5 / 2609.03580
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews
HalluPeer:面向科学同行评审中幻觉检测的基于分类法的基准
Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review, where verification requires grounding claims in long, technical papers. We introduce HalluPeer, a benchmark for detecting hallucinations in scientific peer reviews, providing aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Our pipeline induces a peer-review-specific hallucination taxonomy, identifies review contexts, and injects hallucinations with automated filtering. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, while evaluation on authentic reviews demonstrates that HalluPeer-defined hallucination patterns occur in real peer reviews, highlighting the critical need for source-aware verification. Our project page can be found in https://github.com/Lin-TzuLing/HalluPeer.git
Chinese Translation
学术同行评审规模的不断增长推动了将大语言模型(LLMs)用作评审助手的做法,然而LLMs可能生成流畅却缺乏依据的陈述,从而损害评审的可靠性。现有的幻觉基准并非为同行评审而设计,因为在同行评审中,验证要求将陈述锚定于篇幅较长且技术性强的论文中。我们提出了HalluPeer,一个用于检测科学同行评审中幻觉的基准,其提供了由论文内容、人工撰写的评审以及注入幻觉的评审所构成的对齐三元组,并针对检测、分类和定位进行了标注。我们的流程引入了一种面向同行评审的幻觉分类法,识别评审上下文,并通过自动化过滤注入幻觉。在12,000篇论文和38,000条评审上进行的实验表明,现有检测器难以将幻觉与合理的批评区分开来,而对真实评审的评估则证明,HalluPeer所定义的幻觉模式确实出现在真实的同行评审中,这凸显了基于来源感知的验证的迫切需求。我们的项目页面可在 https://github.com/Lin-TzuLing/HalluPeer.git 找到。
cs.AI / 6 / 2609.03586
The Attention Triangle in Audio-Video Models
音频-视频模型中的注意力三角
Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
cs.AI
diffusion
扩散模型相关
Abstract
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.
Chinese Translation
音频-视频扩散模型依赖跨模态注意力来协调文本、声音和视觉内容,然而这一相同的机制也可能引入微妙且系统性的语义泄漏。我们通过探测和分析“注意力三角”来研究这些模型,该三角由连接文本、音频和视频流的三个交叉注意力边组成,并考察在生成过程中语义信息如何跨模态进行路由。我们的分析揭示,沿音频-视频边的路由是双向的:音频可以影响视频生成,而视频也可以影响音频生成。这条边受到模型参数中编码的偏差影响,并成为泄漏的主要来源:当提示与习得的先验相冲突时,跨模态交互可能覆盖预期的条件设定,并将语义重新路由到视觉上典型但错误的结果。这些效应表明,语义伪影不仅源于注意力扩展到预定目标之外,更源于沿特定路径的结构性、偏差驱动的交互。基于这一视角,我们提取了源自注意力的信号,以揭示语义如何在各模态间分布和锚定,并将其用作诊断工具,既用于分析,也用于在受控条件下故意引发泄漏。这使我们能够探测跨模态路由的内部动态,并隔离单个交互的作用。我们进一步利用这些信号来指导推理期间的干预,以促进更一致的跨模态对齐。大量实验支持了我们的分析,并证明了在保持生成质量的同时改善了语义锚定。
cs.AI / 7 / 2609.03635
Analysis of Prompt Engineering for Drug Toxicity Prediction
药物毒性预测的提示工程分析
Mia MacGregor, Aakash Welgamage Don, Mark Bartlett
cs.AI
large language model
大语言模型相关
Abstract
Clinical trials in the UK can cost up to £1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly explored to aid in the prediction of drug toxicity, with extensive use of large language models (LLMs). However, LLMs can show considerable variation when minor changes are made to prompts, which raises concerns about their sensitivity to prompt engineering. Prompt engineering is used to optimise a prompt given to an LLM to generate the desired output. This paper proposes a method to analyse prompt engineering for drug toxicity prediction. The aim of the paper is to investigate the importance of prompt phrasing for drug toxicity prediction. LLMs were prompted to identify chemical properties of significance when predicting drug toxicity. Prompts were constructed to investigate; job role, prompt structuring, and rule interpretation. LLMs were then used to generate datasets, using the identified features from initial prompting, which were then passed to machine learning algorithms. The experiments show that the natural variance which occurs in LLMs outweighs any fine-tuning of prompts. There were, however, substantial improvements in model performance when using chemoinformatic code to extract features instead of using LLM-generated values. The proposed analysis methodology is applicable to a wide range of prompt types across different areas of bioinformatics.
Chinese Translation
英国临床试验的费用可高达130万英镑,约90%的药物失败率。毒性是导致药物失败的一个主要因素。测试既耗时又费钱。近年来,人们越来越多地探索利用人工智能来辅助预测药物毒性,其中广泛使用了大型语言模型(LLMs)。然而,当对提示做微小更改时,LLMs可能表现出相当大的变异,这引发了人们对其对提示工程敏感性的担忧。提示工程用于优化给LLM的提示,以生成期望的输出。本文提出了一种分析药物毒性预测中提示工程的方法。本文旨在研究提示措辞对药物毒性预测的重要性。LLMs被提示在预测药物毒性时识别具有重要意义的化学性质。构建的提示用于研究:工作角色、提示结构化以及规则解释。随后,利用LLMs以初始提示中识别出的特征来生成数据集,这些数据集再传给机器学习算法。实验表明,LLMs中出现的自然变异超过了任何对提示的微调效果。然而,当使用化学信息学代码来提取特征而非使用LLM生成的值时,模型性能有了显著提升。所提出的分析方法适用于生物信息学不同领域的多种提示类型。
cs.AI / 8 / 2609.03727
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
主动服务智能体:一个统一决策框架、方法与评估
Yan Tang, Tingyu Cao, Yuanbo Tang, Huaze Tang, Keer Hu
cs.AI
large language model
大语言模型相关
Abstract
Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete environmental and user signals, choose among remaining silent, asking, assisting, and acting, and account for interruption, misunderstanding, overreach, and privacy costs. This survey gives an operational definition centered on initiative and formulates the problem as a partially observable sequential decision process constrained by authorization and risk. The formulation represents timing, content, and delivery within one structured action, while making explicit the option value of waiting, the decision value of questions, and feedback-induced state changes. On this basis, we organize existing methods along one decision pipeline (state and need estimation, intervention gating, action construction, and feedback adaptation) and describe prescribed, predictive, model based, and return optimizing mechanisms as nonexclusive policy-construction components. We further normalize decision units and three-axis evidence descriptors across streaming dialogue, screen, video, software-engineering, and human-agent collaboration resources, and formalize metrics for triggering, timing, calibration, user burden, safety, and policy value. The synthesis shows why offline classification performance alone does not predict deployment benefit and why long-term memory is not a defining condition of proactivity. Reliable proactive service instead requires calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.
Chinese Translation
大规模语言模型智能体能够规划、调用工具并修改外部状态,但大多数系统仍将明确的用户指令作为固定起点。主动服务将决策上移:智能体必须从不完全的环境信号和用户信号中推断服务机会,在保持沉默、询问、协助和行动之间进行选择,并考虑中断、误解、越权及隐私成本。本综述给出了以主动性为核心的操作性定义,并将该问题形式化为一个受授权和风险约束的部分可观测序贯决策过程。该形式化在一个结构化动作中表示时机、内容和传递方式,同时显式说明等待的期权价值、询问的决策价值以及反馈带来的状态变化。在此基础上,我们沿一条决策流水线(状态与需求估计、干预门控、动作构建和反馈适应)组织现有方法,并将规定式、预测式、基于模型和回报优化机制描述为非排他性的策略构建组件。我们还对跨流式对话、屏幕、视频、软件工程和人机协作资源的决策单元和三轴证据描述符进行了标准化,并正式定义了触发、时机、校准、用户负担、安全性和策略价值的度量指标。综合结论表明,仅靠离线分类性能无法预测部署收益,长期记忆也不是主动性的决定性条件。可靠的主动服务反而需要校准后的增量干预价值、可验证的授权、可恢复的执行以及反事实证据。
cs.AI / 9 / 2609.03753
SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
SimSkill:一个用于自主掌握交通模拟的终身学习AI智能体
Qi Liu, Qinzheng Wang, Yiming Bie
cs.AI · cs.MA
large language model
大语言模型相关
Abstract
As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-evolving agent built around the Simulation of Urban MObility (SUMO) traffic simulator. SimSkill identifies capability gaps, generates and solves environment-grounded tasks, verifies solutions through an action--critic loop, and consolidates experience into episodic, procedural, and semantic memory without updating the backbone model. Through autonomous exploration, it builds a reusable library spanning the traffic-simulation workflow. We evaluate SimSkill on two held-out benchmarks with three backbone LLMs and independent artifact-based verification. SimSkill improves verified completion by up to 25 percentage points, while ablations show complementary contributions from procedural and semantic memory. Its benefits remain backbone- and budget-dependent: memory does not improve every model or uniformly reduce inference cost. More broadly, SimSkill illustrates a design paradigm in which natural language preserves and composes computational capabilities, while executable tools and code provide precise and reproducible execution. All code and experimental data are publicly available at https://github.com/qiliuchn/SimSkill-V1.
Chinese Translation
随着大语言模型(LLMs)能力日益增强,AI系统的长期价值不仅取决于解决单个请求,还在于将经验和积累的知识转化为持久、可复用的能力。我们引入SimSkill,一个围绕城市交通模拟器SUMO构建的自我进化智能体。SimSkill识别能力差距,生成并解决与环境紧密关联的任务,通过行动-评论家循环验证解决方案,并将经验整合为情景记忆、程序性记忆和语义记忆,而无需更新骨干模型。通过自主探索,它构建了一个覆盖交通模拟工作流的可复用库。我们在两个保留基准上,使用三种骨干LLM和基于工件的独立验证来评估SimSkill。SimSkill将验证完成率最多提高了25个百分点,同时消融实验表明程序性记忆和语义记忆具有互补的贡献。其收益仍然依赖于骨干模型和预算:记忆并不能改善每个模型,也不能统一降低推理成本。更广泛而言,SimSkill展示了一种设计范式,其中自然语言保留并组合计算能力,而可执行工具和代码则提供精确且可复现的执行。所有代码和实验数据均可公开获取:https://github.com/qiliuchn/SimSkill-V1。
cs.AI / 10 / 2609.03871
Bioinfoysis Technical Report
Bioinfoysis 技术报告
Qingyang Shao, Xin Zhang, Zhouyang Yuan, Xianying Chen, Yujia Xiang, Zihao Yang, Tong Ye, Yangqi Zhang, Jiakang Xu, Xiaoqing Yan, Xuan Luo, Keyi Li, Enci Fan, Kai Kang, Zhuohan Liu, Xingyu Jin, Chunran Teng, Tao Li, Xinyu Lv, Minghui Wang, Wenfeng Li, Yidan Gao, Siyu Liu, Mingrui Luo, Zhu Liang, Guanren Qiao, Zhiping Xu
cs.AI · cs.MA
large language model
大语言模型相关
Abstract
Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness that represents each request as a persistent, artifact-grounded analysis run. Bioinfoysis combines global planning with step-wise, evidence-driven replanning: the planner maintains an executable checklist and revises pending steps using structured handoffs returned after each worker execution. These handoffs bind intermediate results to their responsible agent, checklist step, and plan generation, preventing stale evidence from being silently reused after replanning. A controlled runtime validates generated scripts, tables, and figures before they are used in downstream analysis or reporting, while role-specific context, persistent memory, and governed bioinformatics skills support reliable execution over long analysis trajectories. We evaluate Bioinfoysis on BixBench and two question-answering tracks of LAB-Bench 2. On BixBench, Bioinfoysis achieves state-of-the-art accuracy of 82.4\%. Across four underlying language models, Bioinfoysis increases average accuracy from 27.81\% to 64.13\% on SeqQA2 and from 3.13\% to 31.25\% on DbQA2. These results demonstrate that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow. We hope that the emergence of Bioinfoysis will play a driving and leading role in the development of the bioinformatics community. Our demo website can be seen in https://report.bioinfoysis.com/.
Chinese Translation
大型语言模型智能体已在生物信息学中展现出潜力,但大多数现有系统主要侧重于生成最终答案,将规划、工具使用和代码执行视为瞬时交互。这种设计难以适用于长期生物信息学任务,在这些任务中,结论必须始终与支持它们的数据、计算和中间证据保持关联。我们引入了 \textbf{Bioinfoysis},这是一种多智能体框架,可将每个请求表示为一个持续的、基于产物的分析运行。Bioinfoysis 将全局规划与逐步的、证据驱动的重新规划相结合:规划器维护一个可执行的检查清单,并利用每次工作代理执行后返回的结构化交接信息来修订待执行的步骤。这些交接信息将中间结果绑定到其负责的智能体、检查清单步骤和计划生成过程,从而防止过时证据在重新规划后被静默复用。受控运行时会在生成的脚本、表格和图用于下游分析或报告之前对其进行验证,同时,特定角色的上下文、持久记忆以及受管控的生物信息学技能可支持在长分析轨迹上的可靠执行。我们在 BixBench 以及 LAB-Bench 2 的两个问答赛道上对 Bioinfoysis 进行了评估。在 BixBench 上,Bioinfoysis 达到了 82.4\% 的当前最优准确率。在四种底层语言模型上,Bioinfoysis 将 SeqQA2 上的平均准确率从 27.81\% 提高到 64.13\%,并将 DbQA2 上的平均准确率从 3.13\% 提高到 31.25\%。这些结果表明,可靠的生物信息学自动化不仅取决于模型能力,还取决于管理规划、执行、记忆和证据流的框架。我们希望 Bioinfoysis 的出现将对生物信息学社区的发展起到推动和引领作用。我们的演示网站可在 https://report.bioinfoysis.com/ 查看。
cs.AI / 11 / 2609.03874
STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation
STAIR(结构感知信息检索器):一种面向文档结构增强的新型数据集与基于LLM的检索器
Vineet Kumar, Meghanadh Pulivarthi, vishwajeet kumar, Jaydeep Sen, Riyaz Ahmad Bhat, Sachindra Joshi
cs.AI
large language model
大语言模型相关
Abstract
Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from "lost in the middle" problem. Thus, precise and accurate retrieval is important. Current retrievers chunk long context into length-based manageable chunks - in the process throwing away rich and informative semantic global structure in the corpus. We introduce a novel retrieval system STAIR that empowers an LLM to exploit global structure in a corpus such as a Table of Contents (ToC) to efficiently store and retrieve information from its model parameters. Our thorough and careful ablation studies with a finetuned Differentiable Search Index (DSI) system show that ToC helps build a low hallucination (less than 0.05%) generative Information Retrieval (IR) system and can generalize to examples where very few training samples are available. To further research in this novel direction of ToC based retrieval we release SearchTome - a diverse benchmark created from 18 books across 6 diverse domains to further research in this novel direction. STAIR achieves a high Recall@1 score of 82.6% on SearchTome as compared to DSI (76.9%), where the difference is found to be statistically significant. STAIR easily beats other strong baselines such as BM25 (59.5%), DPR (68.7%) and out-of-the-box Mistral (13.8%).
Chinese Translation
检索增强生成(RAG)是利用大型语言模型(LLMs)生成准确且无幻觉答案的关键组成部分。LLMs在处理长上下文方面不断进步,但仍然存在“迷失在中间”的问题。因此,精确且准确的检索十分重要。当前的检索器将长上下文切分为基于长度的可管理块——在此过程中丢弃了语料库中丰富且具有信息量的语义全局结构。我们提出了一种新型检索系统STAIR,它使LLM能够利用语料库中的全局结构(例如目录(ToC))来高效地在模型参数中存储和检索信息。我们使用微调后的可微搜索索引(DSI)系统进行的细致而全面的消融研究表明,ToC有助于构建低幻觉(小于0.05%)的生成式信息检索(IR)系统,并且能够泛化到训练样本极少的情况。为了进一步研究这一基于ToC的检索新方向,我们发布了SearchTome——一个从6个不同领域的18本书中创建的多领域基准数据集。STAIR在SearchTome上取得了82.6%的高Recall@1分数,而DSI为76.9%,且该差异具有统计学显著性。STAIR还轻松超越了BM25(59.5%)、DPR(68.7%)以及开箱即用的Mistral(13.8%)等其他强基线模型。
cs.AI / 12 / 2609.04013
LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
LLM4CKD:用于早期慢性肾脏病筛查的大型语言模型
Muhammad Ashad Kabir, Sirajam Munira
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the effectiveness of large language models (LLMs) for CKD screening under zero-shot and few-shot in-context learning settings and compares them with traditional ML and DL methods. We propose a framework that uses clinically selected tabular features and structured prompt templates to enable LLM-based inference without task-specific training. LLM performance is evaluated across multiple prompt styles, feature configurations, and data settings, and compared with standard ML, DL, and tabular foundation model (TFM) baselines, and existing CKD screening tools. The results show that LLMs can achieve competitive performance using only a small number of examples, often matching or outperforming traditional approaches in low-data settings. However, their performance remains model-dependent and less stable as input complexity increases. In contrast, ML, DL, and TFM models show more consistent improvement with larger training data. Overall, the findings highlight a trade-off between data efficiency and stability, suggesting that LLMs may serve as a flexible complementary approach for CKD screening when labeled data are limited.
Chinese Translation
慢性肾脏病(CKD)的早期筛查对于及时干预至关重要,然而大多数机器学习(ML)和深度学习(DL)方法需要标注数据和模型训练,限制了它们在实际筛查场景中的应用。本研究评估了大型语言模型(LLMs)在零样本和少样本上下文学习设置下进行CKD筛查的有效性,并将其与传统ML和DL方法进行比较。我们提出了一个框架,利用临床选择的表格特征和结构化提示模板,使基于LLM的推理无需任务特定训练即可进行。LLM性能在多种提示风格、特征配置和数据设置下进行评估,并与标准ML、DL、表格基础模型(TFM)基线以及现有CKD筛查工具进行比较。结果表明,LLMs仅使用少量示例即可达到有竞争力的性能,在低数据设置中通常能与传统方法相当或更优。然而,其性能仍然依赖于具体模型,并且随着输入复杂性的增加而稳定性下降。相比之下,ML、DL和TFM模型在更大训练数据下表现出更一致的改进。总体而言,研究结果凸显了数据效率与稳定性之间的权衡,表明在标注数据有限的情况下,LLMs可作为CKD筛查的一种灵活补充方法。
cs.AI / 13 / 2609.04014
InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models
InSituMeasure:用多模态大语言模型探测工业场景中的情境测量基础
Chao Shen, Xinyuan Li, Yunfan Zhou, Jianguo Yao, Haibing Guan, Zhihai Wang, Xijun Li
cs.AI
large language model
大语言模型相关
Abstract
For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmarks. Existing benchmarks expose this weakness but isolate measurement from realistic, knowledge-grounded settings, with limited situated context, specialized instruments, real-world noise, and matched diagnostic annotations, reducing realism and constraining root-cause analysis. We introduce InSituMeasure to evaluate situated measurement grounding. It contains 2,922 real industrial monitoring scenes across eight functional categories of professional engineering instruments, with dense gauge-attribute annotations and noise tags for failure diagnosis. We define metrics for numerical accuracy under predefined tolerances and unit consistency, rejection of fake or unanswerable tasks, and alignment between model failures and annotated error factors. Across 24 state-of-the-art MLLMs, the best model reaches only 25.7\% joint value-unit accuracy and 51.8\% confidence-diagnosis F1, revealing a substantial gap between general multimodal competence and reliable situated measurement. Further analysis identifies failures from text-induced shortcuts, overconfident responses, and authentic industrial noise, including mixed disturbances, viewpoint deviation, occlusion, and environmental interference.
Chinese Translation
对于训练有素的操作员而言,读取仪表仅需很少的专业知识,认知负担低,且可重复性高。然而,多模态大语言模型(MLLMs)尽管在通用多模态基准上取得了突出成绩,但在连续值测量中仍然不可靠。现有基准暴露了这一弱点,但将测量与那些情境上下文有限、配备专用仪表、真实噪声和匹配诊断标注的现实且基于知识的场景相分离,从而降低了真实性并限制了根因分析。我们引入了 InSituMeasure,用于评估情境测量基础。它包含 2,922 个真实工业监控场景,涵盖八个功能类别的专业工程仪表,并带有密集的仪表属性标注和用于故障诊断的噪声标签。我们定义了多种指标,用于衡量在预定义容差下的数值准确性和单位一致性、对虚假或无法回答任务的拒绝情况,以及模型失败与已标注错误因素之间的一致性。在 24 个最先进的多模态大语言模型(MLLMs)中,最佳模型仅达到 25.7% 的数值与单位联合准确率和 51.8% 的置信度-诊断 F1,揭示了通用多模态能力与可靠情境测量之间的显著差距。进一步分析发现,失败源自文本诱导的捷径、过度自信的响应以及真实工业噪声,包括混合干扰、视角偏差、遮挡和环境干扰。
cs.AI / 14 / 2609.04021
FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
FLY-EVAL++:面向大型语言模型安全约束飞行预测的证据驱动评估协议
Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an evidence-driven evaluation protocol that combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with fixed rubric-guided aggregation into interpretable multi-dimensional scores. We instantiate FLY-EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance is the most discriminative dimension of model behavior: models with comparable predictive performance differ by more than 28 points in safety score, and we observe recurrent failures including safety violations under physically plausible predictions and instability in multi-step rollouts. These results show that evaluation in safety-critical domains should measure constraint satisfaction and structured validity explicitly rather than rely on accuracy-centric reporting alone.
Chinese Translation
在安全关键且受物理规律支配的环境中评估大型语言模型(LLM)时,仅靠基于准确度的指标是不够的,因为数值上接近真实值的预测仍可能违反操作约束、以物理上不一致的方式组合字段,或者无法产生可用的结构化输出。现有评估协议无法可靠地衡量这些失败模式。我们提出FLY-EVAL++,一种证据驱动的评估协议,它将协议合规性、物理可行性和安全约束的确定性验证,与固定评分标准引导的聚合相结合,形成可解释的多维分数。我们通过扩展PilotBench设置,加入基于历史条件的预测任务和多步预测任务,将FLY-EVAL++实例化于飞行轨迹与姿态预测(FTAP)。在66个LLM中,安全合规性是模型行为中最具区分度的维度:预测性能相当的模型在安全分数上相差超过28分,并且我们观察到反复出现的失败,包括在物理上看似合理的预测中出现安全性违规,以及多步展开中的不稳定性。这些结果表明,在安全关键领域的评估应显式地衡量约束满足情况和结构化有效性,而不是仅仅依赖以准确性为中心的报告。
cs.AI / 15 / 2609.04030
IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations
IRWOZ 2.0:面向工业机器人对话的大型语言模型驱动的对话数据集
Chen Li, Dimitrios Chrysostomou
cs.AI
large language model
大语言模型相关
Abstract
IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on dialogue state tracking demonstrate significant improvements, with GPT-2's BLEU-4 score increasing from 0.1651 to 0.5604 compared to original IRWOZ. To support industrial HRI research, we publicly released IRWOZ 2.0 dataset at https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations
Chinese Translation
IRWOZ通过领域特定的标注改进了工业人机交互(HRI)对话系统。然而,其初始版本在对话状态和话语中包含大量噪声,限制了状态跟踪的准确性。我们引入了IRWOZ 2.0,通过大型语言模型(LLM)增强生成(Mistral/Claude-3.5)和质量改进来解决这些限制。我们改进的数据集扩展到4个工业领域(装配、配送、定位、搬迁)中的390个对话,并包含手动修正和自动化错别字去除。对话状态跟踪的基准实验表明,与原始IRWOZ相比,GPT-2的BLEU-4得分从0.1651提高到0.5604,取得了显著改进。为了支持工业HRI研究,我们在https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations公开发布了IRWOZ 2.0数据集。
cs.AI / 16 / 2609.04127
Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
大语言模型建议的认识论保证:在客观真值不可用时刻画依赖的基础
Shai Vardi, João Sedoc
cs.AI
large language model
大语言模型相关
Abstract
Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology, we introduce epistemic warrant, a decision-level construct that characterizes the stability of a model's preference and the scope over which that preference holds. We operationalize this construct through a four-tier reliance certificate for pairwise recommendations, distinguishing among unstable, context-dependent, locally supported, and broadly supported recommendations. We validate the construct using contemporary methodologies: known-groups tests successfully recover expert-prespecified warrant orderings, and stronger warrants systematically align with independent consensus from crowd workers. Furthermore, we demonstrate that epistemic warrant provides information distinct from verbalized confidence and is not readily explained by decision difficulty. Ultimately, this framework offers a theoretically grounded, implementable approach for characterizing the warrant of individual LLM recommendations when objective ground truth is unavailable.
Chinese Translation
大语言模型越来越多地用于支持组织决策,然而用户往往缺乏评估是否应依赖某一具体建议的原则性依据。现有方法通常评估模型的宽泛属性,如可靠性、不确定性或鲁棒性,或关注用户信任,而不是依赖单个建议的深层基础。借鉴认识论的理论基础,我们引入认识论保证(epistemic warrant)——一种决策层面的构念,用于刻画模型偏好的稳定性以及该偏好得以成立的范围。我们通过对成对建议的四个层级依赖证书来操作化该构念,区分不稳定的、依赖上下文的、局部支持的和广泛支持的建议。我们使用当代方法论验证该构念:已知组测试成功恢复了专家预先指定的保证排序,且更强的保证系统性地与来自众包工作者的独立共识一致。此外,我们证明认识论保证提供了不同于口头表达的信心(verbalized confidence)的信息,并且不易仅用决策难度来解释。最终,该框架提供了一种在客观真值不可用时刻画单个大语言模型建议之保证的、有理论根基且可实施的方法。
cs.AI / 17 / 2609.04172
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
大型语言模型同策略蒸馏的再思考 II:单个训练示例
Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.
Chinese Translation
同策略蒸馏(OPD)将学生生成的轨迹与来自教师的密集词元级监督相结合。现有工作主要研究了其算法行为,而训练数据的作用尚不清楚。我们通过单个查询进行训练,在数据最小极限下考查这一作用。单样本 OPD 在数百步内持续改进,并在各任务领域和模型家族中恢复了全数据 OPD 的大部分收益。我们通过训练过程中访问的状态以及学生与教师对齐的速率来解释这一结果。我们度量状态覆盖率,即一个查询集的轨迹所到达的状态占全数据 OPD 所访问状态的比例。单个查询即可达到 \(71.5\%\),其中大部分在前 100 步内达到。加入语义不同的查询会同时提高覆盖率和验证准确率,直到 16 个查询达到 \(98.9\%\) 并与全数据训练相匹配。然而,无论 OPD 是使用单个查询还是整个数据集训练,对齐都以类似的速度放缓;即便是固定的状态集,也需要数百步才能被吸收。因此,OPD 处于数据过喂、算法饥饿的状态。其轨迹很快呈现出广泛的监督,而学生模型吸收这种监督的速度却越来越慢。该状态覆盖率结果可推广至多教师 OPD(MOPD),其中每个领域 16 个语义多样的查询即可与全数据 MOPD 相匹配。作为进一步的压力测试,内容轻量的模板和领域外的 WildChat 查询也接近真实查询基线。因此,任务内容与所引发的状态覆盖率可以相互分离。我们希望这些发现能引导未来的工作关注 OPD 的步进效率,并促使重新审视数据以及 OPD 在前沿模型后训练中近期成功背后的机制。
cs.CL / 18 / 2609.03718
What Do CAE Simulation Agents Really Need Beyond a Generic Harness?
在通用框架之外,CAE 仿真智能体真正需要什么?
Jiasheng Shi, Tianhan Zhang
cs.CE · cs.CL · physics.comp-ph
large language model
大语言模型相关
Abstract
Computer-aided engineering (CAE) simulation is among the largest and most demanding areas of engineering, where setting up a solver such as OpenFOAM, FEniCS, or COMSOL takes real expertise. Large language model (LLM) agents promise to turn a natural-language request into a working simulation, and recent CAE agents add simulation-specific machinery: multi-agent decomposition, domain retrieval, and scripted reflection. That machinery suited weak base models; modern harnesses already supply multi-turn reasoning, tool use, and execution feedback. We ask what a CAE simulation agent still needs beyond a generic harness. With information access and repair budget held fixed, a single-agent harness matches or beats multi-agent specialized systems (FoamBench 96.4\% vs.\ 88.2\%). Ablations trace this to capabilities the harness already provides: execution-feedback repair lifts FoamBench from 71.8\% with no repair round to 96.4\%, while scripted reflection adds nothing. The one input that still helps is domain knowledge supplied as solver tutorials, our largest measured gain (80.9\% to 96.4\%).
Chinese Translation
计算机辅助工程(CAE)仿真是工程领域中规模最大、要求最高的领域之一,配置 OpenFOAM、FEniCS 或 COMSOL 等求解器需要真正的专业知识。大语言模型(LLM)智能体有望将自然语言请求转化为可运行的仿真,近期的 CAE 智能体还加入了仿真专用机制:多智能体分解、领域检索和脚本化反思。这些机制适用于较弱的基座模型;现代框架已经提供了多轮推理、工具使用和执行反馈。我们探讨的是:在通用框架之外,CAE 仿真智能体还需要什么?在信息访问与修复预算固定的条件下,单智能体框架能够匹敌乃至超越多智能体专用系统(FoamBench 96.4% 对 88.2%)。消融实验将这一表现归因于框架已提供的能力:执行反馈修复将 FoamBench 从无修复轮的 71.8% 提升到 96.4%,而脚本化反思没有带来任何提升。唯一仍然有帮助的输入是以求解器教程形式提供的领域知识,这也是我们测得的最大增益(从 80.9% 提升至 96.4%)。
cs.CL / 19 / 2609.03148
Large Language Models in Resolving Contextual Knowledge Conflicts
大语言模型在解决上下文知识冲突中的作用
Xinye Yang, Zhenyang Liu, Ruisi Li, Yuanyuan Lei
cs.CL
large language model
大语言模型相关
Abstract
Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual conflicts (factual, inferential, temporal, granularity, perspective, and ambiguity) and contribute a comprehensive dataset ContextConflict for this setting. The dataset contains 5,781 samples, covers both reasoning and summarization tasks, and includes both explicit contradictions and implicit conflicts that require multi-step reasoning. Experiments on nine LLMs show that current models still fall short in resolving contextual knowledge conflicts. We further provide mechanistic interpretability insights into how LLMs process such conflicts, revealing their latent awareness of conflicts and the representational geometry underlying conflict processing. In addition, our analysis uncovers a consistent model bias towards earlier evidence, and this positional preference serves as a key obstacle to effective conflict resolution. Motivated by these findings, we further propose a simple training-free, label-free steering method that steers activations to encourage a more comprehensive incorporation of evidences for better conflict resolution. On our dataset, the method consistently improves accuracy on reasoning tasks and generates higher-quality, more balanced summaries for summarization tasks.
Chinese Translation
大多数先前的工作集中于大语言模型(LLM)的内部参数知识与外部提供的上下文之间的冲突。相比之下,我们研究LLM如何处理上下文知识本身内部出现的冲突。我们引入了一个包含六种上下文冲突类型的分类体系(事实性、推理性、时间性、粒度、视角和歧义),并为此场景贡献了一个全面的数据集ContextConflict。该数据集包含5,781个样本,涵盖推理和摘要两类任务,并同时包含显式矛盾和需要多步推理的隐式冲突。在九个LLM上的实验表明,当前模型在解决上下文知识冲突方面仍然不足。我们进一步提供了关于LLM如何处理这类冲突的机制可解释性洞见,揭示了它们对冲突的潜在意识以及冲突处理背后的表征几何结构。此外,我们的分析揭示了模型对较早证据的持续偏见,并且这种位置偏好是有效解决冲突的关键障碍。基于这些发现,我们进一步提出了一种简单、无需训练、无需标签的引导方法,通过引导激活来促进对证据的更全面整合,从而实现更好的冲突解决。在我们的数据集上,该方法在推理任务上持续提高准确率,并在摘要任务上生成更高质量、更均衡的摘要。
cs.CL / 20 / 2609.03160
No country for old linguists: LLM-brain alignment underdetermines neural computation
老语言学家无处容身:LLM-大脑对齐不足以确定神经计算
Elliot Murphy
cs.CL
large language model
大语言模型相关
Abstract
Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both rely on distributed, context-sensitive representations shaped by statistical learning. Their rejection of simple cortical "boxology" is persuasive, and they articulate a strong case for the value of LLM-brain alignment research. The key question is what kind of inference LLM-brain alignment licenses. My claim here will be narrow: representational alignment can in principle constrain mechanistic hypotheses, but it does not by itself identify a mechanism. Nastase et al. acknowledge that an encoding model can capture features represented in neural activity without establishing a shared architecture or algorithm. Yet the authors sometime move from alignment to "shared computational principles" and ultimately to LLMs as mechanistic models of natural language. Indeed, their methodological caveat that alignment does not establish a shared architecture or algorithm sits uneasily with their conclusion that LLMs might instantiate the same computational principles as biological brains and provide a "fully mechanistic model" of language. I discuss what I consider to be problems of logical, causal, and computational underdetermination in Nastase et al.'s (2026) proposal.
Chinese Translation
Nastase等人(2026)认为,大语言模型(LLMs)或许能阐明语言处理过程,因为二者都依赖于由统计学习塑造的分布式、上下文敏感的表征。他们对简单的皮层“盒式划分”的驳斥是有说服力的,并且他们为LLM-大脑对齐研究的价值提出了有力的论证。关键问题是,LLM-大脑对齐允许我们作出何种推断。我在此的论断范围较窄:表征对齐原则上能约束机制性假说,但它本身并不能确定一个机制。Nastase等人承认,编码模型无需确立共享的架构或算法也能捕获神经活动中所表征的特征。然而,作者有时会从对齐走向“共享的计算原理”,并最终将LLMs视为自然语言的机制性模型。的确,他们在方法论上提出的警示——即对齐并不能确立共享的架构或算法——与他们所得出的结论,即LLMs可能例示与生物大脑相同的计算原理并提供语言的“完全机制模型”,两者并不契合。我讨论了在我看来Nastase等人(2026)的提议中所存在的逻辑、因果和计算上的欠定问题。
cs.CL / 21 / 2609.03213
LLMs Learn Better In-Context from Rules than from Examples
大语言模型从规则中比从示例中能更好地进行上下文学习
Xiang Fu, Seungmin Cho, Yukyung Lee, Najoung Kim
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from prompt contexts without weight updates. We compare the learning efficacies of two prominent modes of in-context learning: (1) learning from descriptions of rules (instruction following); and (2) learning from examples of input-output demonstrations (few-shot prompting). Through five learning tasks that cover diverse domains (games, arithmetic, linguistic inferences), we compare two modes of learning (rules vs. examples) specifying the same underlying task. We furthermore explore model and task properties that modulate the learning efficacies. We find that models generally learn more reliably from rules than from examples alone, and additional examples on top of rules or simply scaling up the number of examples do not lead to consistent and significant gains. Instruction tuning amplifies the benefit of rule-based learning while keeping example-based learning capacities intact. Surprisingly, we find no privileged effect of example-based learning in base models, and rules still lead to gains in algebraic task domains. Overall, the comparative efficacy of rules over examples is larger when the task recruits algebraic abstractions and computations, and smaller when the task requires distributional sensitivity and/or recruits parametric knowledge.
Chinese Translation
大语言模型(LLMs)展现出上下文学习能力,即它们无需更新权重即可从提示上下文中学习新任务。我们比较了两种显著的上下文学习模式的学习效能:(1)从规则描述中学习(指令跟随);(2)从输入-输出示范示例中学习(少样本提示)。通过涵盖不同领域(游戏、算术、语言推理)的五项学习任务,我们比较了指定相同底层任务的两种学习模式(规则 vs. 示例)。我们进一步探索了调节学习效能的模型与任务属性。我们发现,模型通常从规则中学习比单独从示例中学习更可靠,并且在规则之上额外增加示例或仅扩大示例数量并不会带来一致且显著的收益。指令微调增强了基于规则学习的优势,同时保持了基于示例学习的能力不变。令人惊讶的是,我们发现基础模型中基于示例学习没有特殊优势,并且规则在代数任务领域仍然能带来性能提升。总体而言,当任务涉及代数抽象和计算时,规则相对于示例的比较效能更大;当任务需要分布敏感性和/或调用参数化知识时,这种比较效能更小。
cs.CL / 22 / 2609.03218
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis
提示词中的分析师:LLM金融分析中的角色、检索与记忆偏差
Ahmed Asaad, Amr Mohamed, Yang Zhang, Omneya Abdelsalam
cs.CL · cs.CE · q-fin.PM
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.
Chinese Translation
大语言模型(LLM)越来越多地利用用户上下文(如记忆、用户画像和角色提示词)来个性化其响应。这种个性化会影响基于证据的判断:相同的证据在不同的用户上下文下可能导致不同的结论。金融领域为研究这一问题提供了高风险场景,因为决策往往依赖于对冗长复杂文档的解读。我们使用3,575份美国证券交易委员会(SEC)申报文件对12个大语言模型进行了测试。我们比较了角色条件化检索、中立检索和记忆框架上下文,以区分证据选择的影响与解读的影响。我们发现,大部分用户上下文溢出效应来自模型在不同角色下对相同证据的解读方式,而非来自检索到不同证据。随后,我们测试了两种简单的缓解策略:以用户画像而非助手角色的形式表达相同的投资者心态,以及将基于证据的输出与个性化输出分离。这两种策略都减少了溢出,但都不能完全消除溢出,而且其有效性在不同模型之间存在显著差异。
cs.CL / 23 / 2609.03235
SGD-KV: Summarization Guided KV Cache Compression
SGD-KV:摘要引导的 KV 缓存压缩
Zeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi, Vivek Govindan, Aram Galstyan, Sravan Babu Bodapati, Srikanth Ronanki
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different attention heads. We present SGD-KV (Summarization-Guided KV Cache Compression), a head-aware framework that leverages a novel chunk-summarization diagnostic task to systematically identify and prioritize attention heads specialized in hierarchical information aggregation. Experiments on Qwen2.5-7B-1M and Qwen3-32B across diverse long-context benchmarks demonstrate that SGD-KV achieves state-of-the-art performance with contexts up to 1M tokens, while reducing KV cache memory usage by up to 75%. Our findings show that strategically allocating the KV cache budget based on the summarization score distribution of attention heads yields a superior efficiency-accuracy trade-off for long-context inference.
Chinese Translation
大语言模型(LLMs)在长上下文推理中面临严重的内存瓶颈,因为键值(KV)缓存的大小呈线性增长。现有的 KV 缓存压缩技术通常依赖简单的启发式方法,忽视了不同注意力头的不同功能角色。我们提出了 SGD-KV(摘要引导的 KV 缓存压缩),一种头感知框架,利用新颖的块摘要诊断任务来系统性地识别并优先处理专门负责层级信息聚合的注意力头。在 Qwen2.5-7B-1M 和 Qwen3-32B 上跨越多种长上下文基准的实验表明,SGD-KV 在长达 100 万 token 的上下文中实现了最先进的性能,同时将 KV 缓存内存使用量减少高达 75%。我们的研究结果表明,根据注意力头的摘要得分分布战略性分配 KV 缓存预算,可以在长上下文推理中获得更优的效率-准确性权衡。
cs.CL / 24 / 2609.03254
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
还需要修复什么?探索对话生成工件中修订传播的性价比测试时计算
Daisuke Kikuta
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies and propagate the revision to all affected parts of the artifact. This paper studies this ability of LLMs on conversationally generated artifacts, where the artifact context and its dependencies may be embedded in the conversation history. Toward practical use, we also explore cost-effective test-time compute for this new setting. Specifically, we introduce a new benchmark for this setting, and evaluate nine revision methods, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark. The results show that baselines achieve accuracies of 68.3--93%, and the most cost-effective method is selecting from three parallel samples using either LLM-based or medoid selection, which improves accuracy by 2.2--9.7%. Our code and dataset are available at https://github.com/ntt-dkiku/llm-revision-propagation.
Chinese Translation
大型语言模型(LLM)通常通过对话中迭代的生成与修订循环来帮助用户生成工件。这里的一个挑战是,当用户在修订过程中仅指定局部更改时,LLM 必须识别相关的依赖关系,并将修订传播到工件的所有受影响部分。本文研究了 LLM 在对话式生成工件上的这种能力,其中工件上下文及其依赖关系可能嵌入在对话历史中。为了实际使用,我们还为这一新设置探索了性价比高的测试时计算。具体而言,我们为此设置引入了一个新基准,并使用 gpt-oss-20b/120b、gpt-5.4-mini 和 qwen3.5-9b/27b/122b 在该基准上评估了九种修订方法,包括顺序反思和并行采样变体。结果表明,基线方法的准确率为 68.3%--93%,而最具性价比的方法是从三个并行样本中使用基于 LLM 的选择或 medoid 选择,这可将准确率提高 2.2%--9.7%。我们的代码和数据集可在 https://github.com/ntt-dkiku/llm-revision-propagation 获取。
cs.CL / 25 / 2609.03321
Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue
将话轮转换与语义解耦:一种基于有限状态机的全双工对话的解耦数据方法
Yihang Li, Chenhui Chu
cs.CL
large language model
大语言模型相关
Abstract
The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializing turn-taking control and response generation onto a single causal tape under the standard next-token prediction objective, thereby preserving semantic prowess at a low fine-tuning cost. However, its reliance on synthetic text data fundamentally limits turn-taking naturalness, as Large Language Models (LLMs) cannot faithfully simulate the fine-grained acoustic temporal dynamics of real human dialogues. In this work, we propose a decoupled data approach that learns turn-taking from real Human-Human (HH) spoken dialogues while shaping semantic behavior through configurable Human-Agent (HA) text dialogues. To operationalize this approach, we introduce a rule-based event-guided data transformation method that serializes HH spoken dialogues into FSM tapes by classifying turn-taking events and applying deterministic mapping rules, enabling scalable supervision without LLM-generated annotations. We further propose a Source-Aware Calibrated (SAC) Loss that jointly calibrates the long-tailed distribution of state transition tokens and channels each data source toward the capability it best supervises. Experiments show that our approach substantially improves turn-taking proficiency while recovering the foundation LLM's semantic capability. Our code and model are available at https://github.com/Liyht/def-fsm.
Chinese Translation
神经有限状态机(NFSM)框架通过将话轮转换控制和响应生成序列化到单一因果带上,并在标准的下一个词元预测目标下进行训练,为全双工对话提供了一条实用的路径,从而以较低微调成本保留了语义能力。然而,其对合成文本数据的依赖从根本上限制了话轮转换的自然度,因为大型语言模型(LLMs)无法忠实地模拟真实人类对话中细粒度的声学时间动态。在这项工作中,我们提出了一种解耦数据方法,从真实的人-人(HH)口语对话中学习话轮转换,同时通过可配置的人-智能体(HA)文本对话来塑造语义行为。为了实现这种方法,我们引入了一种基于规则的事件引导数据变换方法,通过分类话轮转换事件并应用确定性映射规则,将HH口语对话序列化为FSM磁带,从而在没有LLM生成的标注的情况下实现可扩展的监督。我们进一步提出一种源感知校准(SAC)损失,它联合校准状态转移词元的长尾分布,并将每个数据源引导到其最擅长监督的能力上。实验表明,我们的方法在恢复基础LLM语义能力的同时,显著提高了话轮转换的熟练度。我们的代码和模型可在 https://github.com/Liyht/def-fsm 获取。
cs.CL / 26 / 2609.03322
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
扰动如何传播:大型语言模型鲁棒性的多层次分析
Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang
cs.CL · stat.ML
large language model
大语言模型相关
Abstract
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints by analyzing layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores are especially associated with activation-patching recovery under token substitution and shuffling. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.
Chinese Translation
语言模型会遇到拼写错误、损坏的文本、被改写的单词以及被打乱的词序,然而鲁棒性通常仅通过输出行为来评估。我们研究了六种自然主义和合成输入扰动如何在三个层次上通过仅解码器语言模型传播:输出行为、隐藏状态几何以及注意力头功能。我们通过使用中心核对齐和内在本征维数分析逐层几何结构,评估了跨越四个GPT-2和两个Qwen2.5检查点的行为效应,并检查了GPT-2中注意力头的响应。扰动类型会产生可区分的度量轮廓,这些轮廓无法完全由输出度量捕捉,并且在所测试的检查点之间仅部分一致。复制分数尤其与令牌替换和打乱下的激活修补恢复相关。在GPT-2中,梯度引导的HotFlip扰动也比速率匹配的随机令牌替换引起更强的行为和表征破坏;它们的行为效应在所有六个测试的检查点上是一致的。我们的结果表明,基于单一行为或表征度量的鲁棒性声明可能会产生误导,并促使我们采用多层次评估来理解扰动如何改变语言模型的计算。
cs.CL / 27 / 2609.03330
Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour
少即是道德:一个用于认可行为中道德基础检测的迷人CHARM框架
Huixiang Fu, Marian-Andrei Rizoiu
cs.CL · cs.CY · cs.SI
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MA\textbf{C}- and \textbf{H}ate-speech-\textbf{A}ware \textbf{R}ationale-aligned \textbf{M}oral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\% subsample of the MFTC, MFRC, and News training pools together with the richer supervision in MFTCXplain, CHARM improves AUC by up to 15.3\% in-domain, surpasses the supervised baselines on every out-of-domain dataset in both AUC and F1, and offers a scalable, low-cost alternative to prompting-based LLM detectors. We further apply CHARM to large-scale COVID-19 discourse on Twitter and show that moral value alignment is strongly associated with online endorsement behavior. By making moral framing measurable at scale, CHARM offers a practical tool for studying the spread of morally charged misinformation.
Chinese Translation
道德语言在塑造在线认可和信息传播方面发挥着核心作用,然而现有的道德基础检测系统往往存在跨领域泛化能力差、理由依据薄弱以及依赖成本高昂的基于提示的大语言模型(LLMs)等问题。我们提出了CHARM——一个基于轻量级微调LLM的、结合了MAC和仇恨言语感知的、理由对齐的道德基础检测框架,它整合了互补的道德依据、理由对齐以及极性感知的仇恨言语信号,以支持更稳健、更忠实的道德预测。与以往基于词典、微调或提示的检测器不同(这些检测器将计算与心理学理论相分离),CHARM的构建使得每个组件——MAC跨注意力、理由对齐和仇恨言语调制——都操作化了一个不同的心理构念。利用MFTC、MFRC和News训练池的30%子样本以及MFTCXplain中更丰富的监督信息,CHARM在领域内将AUC最多提高了15.3%,在每一个领域外数据集上的AUC和F1均超过监督基线,并为基于提示的LLM检测器提供了一种可扩展、低成本的替代方案。我们进一步将CHARM应用于Twitter上大规模的COVID-19话语,并表明道德价值对齐与在线认可行为密切相关。通过使道德框架可在大规模上被测量,CHARM为研究带有道德色彩的虚假信息的传播提供了实用的工具。
cs.CL / 28 / 2609.03370
FrameBench:A Language Understanding Benchmark Based on Frame Semantics
FrameBench:基于框架语义的语言理解基准
Chihiro Yano, Ryohei Sasano
cs.CL
large language model
大语言模型相关
Abstract
In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieved strong performance across a wide range of downstream tasks. However, it remains unclear whether they can reproduce the kinds of implicit enrichment that humans naturally make during comprehension. To address this question, we introduce FrameBench, a benchmark grounded in frame semantics. FrameBench consists of multiple-choice questions that test whether models distinguish the frames evoked by the same verb across contexts. We construct the benchmark for English and Japanese using FrameNet-style resources and a generation-and-verification pipeline with native-speaker judgments. Our experiments on a diverse set of models reveal challenges for small models, while several large models surpass the human reference scores. We release the constructed FrameBench dataset and the code for dataset construction and evaluation at https://github.com/SasanoLab/FrameBench.
Chinese Translation
在框架语义学中,句子理解被认为是通过将词汇意义与称为语义框架的背景知识联系起来而进行的,从而使读者能够在文本中隐式地补充未明确陈述的信息。近年来的大型语言模型(LLM)在广泛的下游任务中取得了强劲的表现。然而,它们是否能够复现人类在理解过程中自然进行的那种隐式补充,仍不清楚。为了解决这个问题,我们引入了FrameBench,一个基于框架语义学的基准。FrameBench由多项选择题组成,测试模型是否能区分同一动词在不同语境中所唤起的框架。我们使用FrameNet风格的资源以及包含母语者判断的生成与验证流程,为英语和日语构建了该基准。我们在多种模型上进行的实验揭示了小模型面临的挑战,而几个大型模型的得分超过了人类参考分数。我们在https://github.com/SasanoLab/FrameBench 发布了所构建的FrameBench数据集以及用于数据集构建和评估的代码。
cs.CL / 29 / 2609.03395
TabScope: Question-Adaptive Scope Selection for Table Question Answering
TabScope:面向表格问答的问题自适应范围选择
Yuxiang Wang, Junhao Gan, Jianzhong Qi
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affected by irrelevant table content, while questions requiring broader evidence may still benefit from full-table reasoning. Based on this observation, we propose a question-adaptive framework that dynamically selects between localized and full-table reasoning. The framework constructs question-specific sub-tables through operation-aware table decomposition and uses the predicted question type to determine the appropriate reasoning mode. We further introduce silver reference sub-tables for evaluating evidence selection and construct SLQA, a benchmark based on real-world long tables. Experiments on WikiTQ and SLQA show that localization is particularly effective for lookup and local reasoning questions, while adaptive selection between localized and full-table reasoning achieves the best overall performance. These results highlight that long-table QA requires deciding not only how to localize, but also when to localize. Our code and datasets will be made available upon publication of the paper.
Chinese Translation
大型语言模型(LLMs)在表格问答上表现出强劲性能,但随着表格尺寸的增大,其准确性往往会下降。我们发现,这种下降在不同问题类型之间并非均匀分布。对局部化敏感的问题尤其受到无关表格内容的影响,而需要更广泛证据的问题仍可能受益于全表推理。基于这一观察,我们提出了一个问题自适应框架,动态地在局部化推理与全表推理之间进行选择。该框架通过操作感知的表分解来构建问题特定的子表,并使用预测的问题类型来确定适当的推理模式。我们进一步引入银标准参考子表来评估证据选择,并构建了SLQA——一个基于真实世界长表格的基准。在WikiTQ和SLQA上的实验表明,局部化对于查找和局部推理类问题尤其有效,而在局部化与全表推理之间进行自适应选择则能达到最佳的整体性能。这些结果强调,长表问答不仅需要决定如何局部化,还需要决定何时局部化。我们的代码和数据集将在论文发表后公开提供。
cs.CL / 30 / 2609.03410
To What Extent Do Large Language Models Understand Bangla Idioms?
大型语言模型在多大程度上理解孟加拉语习语?
Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi, Swakkhar Shatabda
cs.CL
large language model
大语言模型相关
Abstract
Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms, complemented by a synthetic multiple-choice question (MCQ) dataset for idiom meaning identification. We conduct a comprehensive evaluation of recent large language models (LLMs) across three idiom-related tasks: paraphrasing, idiom span detection, and meaning identification, leveraging zero-shot and few-shot prompting strategies. Our results reveal substantial variability in model performance, with no single LLM consistently outperforming others across all tasks. Notably, Phi-4-mini-instruct excels in paraphrasing, Kimi-K2-32b-instruct in span detection, and Gemini-2.5-flash in meaning identification. We believe that our datasets and analyses will provide valuable resources to guide future research in improving LLM comprehension of idiomatic expressions, particularly in Bangla and other low-resource languages.
Chinese Translation
习语是自然语言中不可或缺的一部分,反映了文化细微差别,并对计算模型提出了独特的挑战,尤其是在低资源语言中。在本文中,我们提出了首个大规模孟加拉语习语基准数据集,并辅以一个用于习语意义识别的合成多项选择题(MCQ)数据集。我们对近期的大型语言模型(LLM)进行了全面评估,涵盖三个与习语相关的任务:释义、习语跨度检测和意义识别,并利用了零样本和少样本提示策略。我们的结果揭示了模型性能的巨大差异,没有哪个单一的LLM能够在所有任务上持续优于其他模型。值得注意的是,Phi-4-mini-instruct 在释义方面表现出色,Kimi-K2-32b-instruct 在跨度检测方面表现优异,而 Gemini-2.5-flash 在意义识别方面表现突出。我们相信,我们的数据集和分析将为未来改进LLM对习语表达理解的研究提供宝贵资源,尤其是在孟加拉语和其他低资源语言方面。
cs.CL / 31 / 2609.03430
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
随机注意力:重新思考KV缓存驱逐以实现高效推理
Heng Wang, Jielin Qiu, Wenting Zhao, Cheng Qian, Liangwei Yang, Jiawei Han, Heng Ji, Silvio Savarese, Shelby Heinecke, Huan Wang
cs.CL
large language model
大语言模型相关
Abstract
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks it matches the strongest prior evictor while serving 32-43% higher throughput than it in vLLM deployment. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them. Our code is publicly available at https://github.com/SalesforceAIResearch/Random-Attention.
Chinese Translation
大语言模型在需要扩展推理的任务上取得了优越的性能,但长思维链使得KV缓存成为一个严重的存储瓶颈。现有的KV缓存压缩方法共享一个范式:通过某种对每个缓存token未来重要性的估计来为其打分,并保留得分最高的那些。我们表明,这种选择信号几乎没有任何贡献。随机注意力保留提示,并在每个注意力头内均匀随机地驱逐,不计算任何分数;在四个模型和六个推理任务中,它与最强的先前驱逐器性能相当,同时在vLLM部署中,其吞吐量比后者高32-43%。受控实验揭示了这一点,表明:1)提示是缓存中脆弱的部分,而选择器之间的大部分差距仅仅在于它们的选择信号是否恰好保留了提示;2)推理轨迹通过两个层面的冗余保护自身免受驱逐——在文本层面(模型在推理过程中会重述其仍然需要的内容)和跨注意力头层面(每个头保留自己的轨迹副本)——因此,一旦提示得以安全保存,随机抽取就能保留模型仍然需要内容的足够副本,并且无需任何分数就能选出它们。我们的代码公开可用,网址为 https://github.com/SalesforceAIResearch/Random-Attention。
cs.CL / 32 / 2609.03454
When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA
何时检索有帮助:单轮心理健康问答的选择性检索
Hyunseo Oh, Chong-Kwon Kim, Yoonhyuk Choi
cs.CL · cs.IR
large language model
大语言模型相关
Abstract
Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model responses, but its effect is not uniformly beneficial in single-turn mental-health question answering, where user queries often combine emotional distress, treatment concerns, and safety-sensitive needs. We study when retrieval helps or hurts mental-health QA, and whether a lightweight selective retrieval policy can better control this trade-off. We operationalize retrieval need using three draft-conditioned utility dimensions: psychoeducational need, coping need, and response specificity, together with a rule-based safety trigger. Following psychotherapy-grounded RAG systems such as coTherapist, we construct a compact and controllable guideline corpus comprising coping-strategy, psychoeducational, and safety resources. We fine-tune an instruction-tuned generator on MentalChat16K using QLoRA and compare Closed-book, Always Retrieval, and Selective Retrieval settings on CounselBench-Eval and CounselBench-Adv. Experiments show that retrieval is not uniformly beneficial in this domain. Always Retrieval improves specificity but lowers overall quality and introduces additional safety-sensitive failures. Selective Retrieval preserves closed-book behavior for low-need cases while avoiding the additional degradation caused by unconditional retrieval, supporting the view that retrieval activation is a safety-sensitive control decision.
Chinese Translation
检索增强生成(RAG)可以提高大语言模型回答的具体性和依据性,但其在单轮心理健康问答中的效果并非总是有益,在此类场景中,用户查询往往融合了情绪困扰、治疗关切和安全敏感需求。我们研究检索何时对心理健康问答有帮助或有损害,以及轻量级的选择性检索策略能否更好地控制这一权衡。我们使用三个以草稿为条件的效用维度来操作化检索需求:心理教育需求、应对需求和回复具体性,并辅以基于规则的安全触发机制。借鉴诸如coTherapist等基于心理治疗的RAG系统,我们构建了一个紧凑且可控的指南语料库,包含应对策略、心理教育和安全资源。我们使用QLoRA在MentalChat16K上对指令微调的生成器进行微调,并在CounselBench-Eval和CounselBench-Adv上比较闭卷、始终检索和选择性检索设置。实验表明,检索在该领域并非始终有益。始终检索提高了具体性,但降低了整体质量,并引入了额外的安全敏感失误。选择性检索在低需求情况下保持了闭卷行为,同时避免了无条件检索带来的额外性能下降,支持了检索激活是一项安全敏感的控制决策这一观点。
cs.CL / 33 / 2609.03467
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
当用户不提问时:评测对话智能体中的上下文驱动记忆检索
Wen-Yu Chang, Yun-Nung Chen
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory benchmark derived from Lo- CoMo with four query styles: dialog, implicit, counterfactual, and composed. Across five rep- resentative memory systems, we evaluate both retrieval recall and end-to-end response qual- ity. Our experiments show that conversational framing exposes substantial retrieval gaps over- looked by QA benchmarks, especially on im- plicit and composed queries, which multi-facet query rewriting narrows for raw-turn mem- ory but not abstractive memory. We further find that strong retrieval does not fully trans- late into response quality, and that implicit queries exhibit silent grounding, where mem- ory improves contextual grounding without ex- plicitly surfacing the gold fact. These results point to reasoning-based memory elaboration as a promising direction, and we release aux- iliary supportive_memory annotations captur- ing conversationally useful context beyond the original gold evidence.
Chinese Translation
大型语言模型(LLM)正越来越多地被部署为长时程对话智能体,这推动了对记忆系统日益增长的兴趣。然而,现有基准主要通过问答式(QA)探测来评估记忆,而非在实际对话场景中进行评估。我们提出LOCOMO-CONV,一个源自LoCoMo的对话记忆基准,包含四种查询风格:对话式、隐式、反事实和组合式。在五个具有代表性的记忆系统上,我们同时评估了检索召回率和端到端的回复质量。我们的实验表明,对话式框架暴露了QA基准所忽视的实质性检索差距,尤其是在隐式和组合查询上,而多方面的查询改写可以缩小原始轮次记忆的此类差距,但对抽象记忆无效。我们进一步发现,强大的检索并不能完全转化为回复质量,并且隐式查询表现出“静默接地”——即记忆在不明确呈现金标事实的情况下提升了上下文接地性。这些结果表明,基于推理的记忆细化是一个有前景的方向,同时我们发布了辅助性的supportive_memory标注,以捕捉超出原始金标证据的对话上有用的上下文。
cs.CL / 34 / 2609.03511
Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations
在重排中迷失:语义保持扰动下多语言大语言模型的结构敏感性
Karthika Nhayakkat, Rajat Verma, Maharaj Brahma, Vetcha Gnana Mahesh, Maunendra Sankar Desarkar, Ganesh Ramakrishnan, Rohit Saluja
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the structural sensitivity of multilingual LLMs using two linguistically grounded perturbation settings in Hindi and Malayalam: constrained constituent reordering and active-passive voice transformation. We introduce a benchmark dataset IndicReStruct, with two variants, GSM8K-Reordered and GSM8K-Voice, constructed from GSM8K while preserving semantic meaning. Across six state-of-the-art LLMs and multiple prompting strategies, we observe consistent and significant degradation in mathematical reasoning performance under structurally perturbed inputs. To further understand these failures, we perform qualitative error analysis and mechanistic interpretability experiments using residual-stream activation patching. Our analyses show that reasoning failures frequently arise from disruptions in entity-quantity alignment and that intermediate transformer layers contribute most strongly toward reasoning restoration. Overall, our findings suggest that current multilingual LLMs remain highly sensitive to surface syntactic realization and lack robust compositional invariance under structurally different but semantically equivalent inputs.
Chinese Translation
大型语言模型(LLMs)展现出强大的多语言推理能力,然而它们对保持语义的结构变化的鲁棒性仍未得到充分探索,尤其是对于语序相对自由的语言。我们使用印地语和马拉雅拉姆语中两种具有语言学基础的扰动设置来研究多语言LLMs的结构敏感性:受限成分重排和主动-被动语态转换。我们引入了一个基准数据集IndicReStruct,包含两个变体GSM8K-Reordered和GSM8K-Voice,这两个变体在保持语义的同时基于GSM8K构建。在六个最先进的LLMs和多种提示策略上,我们观察到在结构扰动输入下,数学推理性能出现了一致且显著的下降。为了进一步理解这些失败,我们使用残差流激活修补进行了定性错误分析和机制可解释性实验。我们的分析表明,推理失败通常源于实体-数量对齐的破坏,并且中间Transformer层对推理恢复的贡献最为显著。总体而言,我们的研究结果表明,当前的多语言LLMs对表层句法实现仍然高度敏感,并且在结构不同但语义等价的输入下缺乏鲁棒的组合不变性。
cs.CL / 35 / 2609.03619
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation
记住并重新加权:利用经验记忆与置信度估计增强多智能体辩论
Xuanfa Jin, Zhijian Ma, Yongcheng Zeng, Xinyu Cui, Haifeng Zhang, Jun Wang
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the agents' inherently biased concept priors unaddressed. To mitigate this systematic weakness, we propose R$^2$-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates. R$^2$-MAD intervenes on both failure modes through two complementary mechanisms: A debate-state-aware retrieval policy dynamically calibrates the concept prior by retrieving relevant historical evidence based on the current consensus level. Then these retrieved experiences provide a basis for estimating per-agent reliability, yielding confidence weights to modulate peer influence. Experiments on various benchmarks show that R$^2$-MAD achieves consistent improvements over existing single-agent and MAD baselines.
Chinese Translation
多智能体辩论(MAD)通过让多个智能体在讨论中迭代地优化其回答,提升了大语言模型的推理能力。然而,MAD存在一个被称为“共享误解”的关键弱点:当大多数智能体最初收敛于一个错误答案时,辩论过程往往会放大而非纠正该错误。现有方法主要处理同伴偏斜,却未解决智能体自身固有的有偏概念先验。为缓解这一系统性缺陷,我们提出了R$^2$-MAD(多智能体辩论中的记住与重新加权),一个为智能体配备从过往辩论中积累的经验记忆的框架。R$^2$-MAD通过两种互补机制同时干预上述两类失败模式:一种基于辩论状态的检索策略,根据当前共识水平检索相关历史证据,从而动态校准概念先验;随后,这些检索到的经验为估计每个智能体的可靠性提供了基础,产生用于调节同伴影响的置信权重。在多个基准上的实验表明,R$^2$-MAD相较于现有的单智能体和MAD基线取得了持续的改进。
cs.CL / 36 / 2609.03775
Typological Feature Prediction with Large Language Models: An In-Context Learning Approach
基于大语言模型的类型学特征预测:一种上下文学习方法
Qianwen Wang, York Hay Ng, Aditya Khan, En-Shiun Annie Lee
cs.CL
large language model
大语言模型相关
Abstract
Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels and feature types remains underexplored. Given LLMs' abilities in meta-linguistic reasoning and in providing rationales, we investigate LLMs' performance in typological feature prediction via an in-context learning approach with linguistic data from URIEL+ and Glottolog. We find that zero-shot prompting is insufficient, but when given phylogenetic and geographic neighbour evidence, LLMs substantially outperform all baselines without disadvantaging low-resource languages. We further find that most LLM rationales are consistent with the provided evidence, offering a step toward explainable typological feature prediction.
Chinese Translation
类型学特征被广泛用于多语言自然语言处理中,对此类特征的预测具有下游实用价值。然而,现有预测缺失值的方法缺乏对预测结果的可解释性论证,同时它们在资源层级和特征类型上的表现仍未被充分探索。鉴于大语言模型在元语言推理及提供推理依据方面的能力,我们通过一种结合 URIEL+ 和 Glottolog 语言学数据的上下文学习方法,考察了大语言模型在类型学特征预测中的表现。我们发现,零样本提示是不够的,但当提供系统发育和地理邻近证据时,大语言模型在不损害低资源语言的情况下显著优于所有基线方法。我们还发现,大多数大语言模型的推理依据与所提供的证据一致,这为迈向可解释的类型学特征预测迈出了一步。
cs.CL / 37 / 2609.03781
IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
IndicSafeEval:多语言说服性越狱攻击下大语言模型的安全鲁棒性
Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and risk categories. Our analysis shows that the model does not behave equally safely across all languages and prompt styles. Instead, safety performance depends strongly on both the languages used and the way a request is phrased using persuasive cues. We further observe that different risk categories exhibit different levels of vulnerability, with some types of harmful content being significantly more susceptible to persuasion-based jailbreaks than others. These findings reveal important limitations of current safety evaluations, which are largely English-centric, and underscore the need for multilingual and persuasion-aware benchmarking frameworks to more accurately assess real-world LLM safety. Our implementation is available at https://github.com/MonSaikat/IndicSafeEval. Warning: this paper contains example data that may be offensive or harmful.
Chinese Translation
大语言模型(LLMs)越来越多地被用于多语言环境中,然而其安全性仍然主要在英语环境下进行评估。这限制了我们对对齐失败如何在低资源和文化多样性语言中显现的理解。我们提出了IndicSafeEval,一个面向印度语言的基于说服的越狱评估框架。我们的基准测试结合了十个安全关键内容类别与六种类人说服策略,覆盖四种不同的印度语言,如印地语、孟加拉语、马拉地语和旁遮普语,生成了7,200个对抗性提示。我们对多个开源大语言模型进行了系统的黑盒评估,以考察其安全行为在不同语言、说服策略和风险类别之间的差异。我们的分析表明,模型并非在所有语言和提示风格下都表现出同等安全的行为。相反,安全性能强烈依赖于所使用的语言以及使用说服性线索表达请求的方式。我们进一步观察到,不同的风险类别表现出不同程度的脆弱性,某些类型的有害内容比其他内容更容易受到基于说服的越狱攻击。这些发现揭示了当前主要以英语为中心的安全评估的重要局限性,并强调了需要多语言且具有说服感知的基准测试框架,以更准确地评估现实世界中的大语言模型安全性。我们的实现可在 https://github.com/MonSaikat/IndicSafeEval 获取。警告:本文包含可能具有冒犯性或有害的示例数据。
cs.CL / 38 / 2609.03814
Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation
评估大型语言模型在内容审核中的标准条件化行为
Danting Zhang, Bei Peng, Robert Loftin
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures at the criterion level. Models struggle most when correct decisions depend not on overall harmfulness, but on the specific aspect of the content that the criterion requires them to assess. Our results highlight a key limitation of current content moderation benchmarks: strong performance on aggregated labels does not provide sufficient evidence that LLMs can reliably evaluate content with respect to individual moderation criteria. These findings call for the development of evaluation methods that explicitly measure criterion-conditioned behaviour.
Chinese Translation
大型语言模型(LLMs)在标准内容审核基准上表现出色。然而,这些基准往往将多个审核标准聚合为单一标签,这使得我们不清楚模型是否能够区分这些标准并在做出决策时可靠地逐一应用每个标准。为了研究LLMs是否表现出标准条件化行为,我们引入了内容诊断评估(DECO),这是一种与标准无关的内容分解方法,能够实现受控的、标准级别的评估。我们还引入了成对评估,以比较同一输入在不同标准下的模型输出。在四个审核数据集和四个LLMs上,我们发现强大的基准性能可能掩盖标准级别的重大失败。当正确决策不是取决于整体有害性,而是取决于标准要求模型评估的内容的具体方面时,模型表现得最为吃力。我们的结果凸显了当前内容审核基准的一个关键局限性:在聚合标签上的优秀表现并不能提供充分证据,证明LLMs能够针对各审核标准可靠地评估内容。这些发现呼吁开发能够明确衡量标准条件化行为的评估方法。
cs.CL / 39 / 2609.03894
CROCODIL: Cross-Model Code Editing with LLMs
CROCODIL:基于LLM的跨模型代码编辑
Linghan Zhong, Aditya Thimmaiah, Jayanth Srinivasa, Milos Gligoric, Junyi Jessy Li
cs.CL · cs.SE
large language model
大语言模型相关
Abstract
Large language models (LLMs) have become ubiquitous tools for code generation and editing. However, development teams often use multiple LLM assistants. Different developers may prefer different models, and individual developers may switch between models across different coding sessions. Because of this, the edits any one model makes are frequently applied to foreign code originally generated by another model. These LLMs are often trained on different datasets, and as a result have different stylistic preferences. Do LLMs behave differently when they edit foreign code originally written by a different LLM with a different coding style? We find that models tend to make more, and often excessive, edits on foreign code. We introduce CROCODIL (Cross-model Code Editing with LLMs), a post-training framework for reducing excessive edits while preserving functional correctness. CROCODIL's similarity reward penalizes large changes, while its execution reward scores build and test success. We use the product of these two rewards to encourage the policy to decrease the edit size without decreasing the edit task success rate. CROCODIL is available at https://github.com/EngineeringSoftware/Crocodil.
Chinese Translation
大型语言模型(LLM)已成为代码生成与编辑中无处不在的工具。然而,开发团队经常使用多个LLM助手。不同的开发者可能偏好不同的模型,而单个开发者也可能在不同的编码会话之间切换模型。正因为如此,任何一个模型所做的编辑经常被应用到最初由另一个模型生成的外部代码上。这些LLM往往在不同的数据集上训练,因此具有不同的风格偏好。当LLM编辑最初由另一个具有不同编码风格的LLM编写的外部代码时,它们的行为会有所不同吗?我们发现,模型倾向于在外部代码上做出更多、往往过度的编辑。我们提出了CROCODIL(基于LLM的跨模型代码编辑),这是一种后训练框架,用于在保持功能正确性的同时减少过度编辑。CROCODIL的相似性奖励惩罚较大的改动,而其执行奖励则对构建和测试成功进行评分。我们使用这两个奖励的乘积来鼓励策略在不降低编辑任务成功率的情况下减小编辑规模。CROCODIL可在 https://github.com/EngineeringSoftware/Crocodil 获取。
cs.CL / 40 / 2609.03955
Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs
用于代码大型语言模型中健全与对抗性测试生成的两阶段强化学习
Jiacheng Xu, Wentao Zhang, Zhiyi Lyu, Fuxiang Zhang, Chaojie Wang, Yang Liu, Bo An
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as counterexamples, depending on the solver's current failure modes. We propose Test Cases Scaling (TCS), a two-stage RL framework for effective test generation. Both stages train a test generator from a rolling policy-aligned buffer: Stage 1 generates tests consistent with the reference solution, and Stage 2 restricts the buffer to current failure modes and learns counterexample tests. Across TACO and LiveCodeBench, TCS improves both pass@1 and inference-time answer selection according to generated tests. We find the learned test generator also enables effective selection among other LLM outputs.
Chinese Translation
强化学习(RL)通过可执行反馈显著推进了使用大型语言模型(LLMs)的代码生成。编码问题的反馈主要来自特定的测试用例,而高质量的测试用例通常稀缺,因为它们既需要是健全的又需要具有区分性。因此,我们转而研究利用学习模型自动生成测试用例。我们发现这自然是一个对抗性强化学习问题:模型需要根据求解器当前的失败模式生成有效的测试用例作为反例。我们提出了测试用例缩放(TCS),一个用于有效测试生成的两阶段强化学习框架。两个阶段都从滚动策略对齐缓冲区中训练测试生成器:阶段1生成与参考解一致的测试,阶段2将缓冲区限制为当前失败模式并学习反例测试。在TACO和LiveCodeBench上,TCS根据生成的测试提高了pass@1和推理时答案选择。我们发现学习到的测试生成器还能在其他LLM输出之间进行有效选择。
cs.CL / 41 / 2609.03967
Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes
探究大型语言模型分析糖尿病食谱的能力
Revathy Venkataramanan, Aditya Luthra, Venkatesan Nadimuthu, Amit Sheth
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positive outcomes. These models can process natural language inputs and leverage learned knowledge from their pretraining to generate meal plans. In this work, we investigate the ability of LLMs to analyze the suitability of given recipes for diabetes. The primary challenge for LLMs is to retrieve relevant dietary guidelines for diabetes, decompose recipes into ingredients and cooking methods, and apply these guidelines to determine the recipe's suitability. To study these challenges, we employ three kinds of prompts namely, (i) Direct Query Prompt (ii) Context-Guided Prompt, and (iii) Exemplary Context Prompt that incorporate different levels of diabetes dietary guidelines from medical sources. We introduce a benchmark dataset curated for this investigation consisting of 7607 recipes that include 3807 recipes suitable for diabetes and 3800 recipes not suitable for diabetes. Our results demonstrate that most LLMs are cautious in predicting recipes as suitable to prevent detrimental outcomes. Further, the models that can reason using the dietary guidelines performed better in predicting the suitability of recipes for diabetes. Overall, Mistral-7B and Llama 70B showed superior performance to their counterparts.
Chinese Translation
多项研究评估了大型语言模型(LLMs)在膳食规划方面的能力,并取得了积极的结果。这些模型能够处理自然语言输入,并利用其预训练阶段所学到的知识来生成膳食计划。在本工作中,我们探究了LLMs分析给定食谱是否适合糖尿病患者的能力。LLMs面临的主要挑战是检索相关的糖尿病膳食指南,将食谱分解为食材和烹饪方法,并应用这些指南来判断该食谱的适宜性。为了研究这些挑战,我们采用了三种提示方式,即(i)直接查询提示,(ii)上下文引导提示,以及(iii)示例上下文提示,它们包含来自医学来源的不同级别的糖尿病膳食指南。我们引入了一个为此研究而整理的基准数据集,包含7607个食谱,其中3807个食谱适合糖尿病患者,3800个食谱不适合糖尿病患者。我们的结果表明,大多数LLMs在预测食谱是否适合时持谨慎态度,以避免产生不利后果。此外,能够利用膳食指南进行推理的模型在预测食谱对糖尿病的适宜性方面表现更好。总体而言,Mistral-7B和Llama 70B的表现优于其同类模型。
cs.CL / 42 / 2609.03992
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
用于配音和全双工对话合成的免对齐文本音响盒
Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
cs.CL · eess.AS
diffusion
扩散模型相关
Abstract
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.
Chinese Translation
我们提出了免对齐文本音响盒(Text-AB),一个用于高质量配音和全双工对话合成的统一框架。基于以流匹配目标训练的扩散Transformer,Text-AB在三个维度上与Audiobox系统有所不同。首先,它在潜在扩散框架中运行,使用DAC-VAE特征将48 kHz波形编码为25 Hz潜在序列,其压缩率比之前的EnCodec表示高出10倍以上,同时提高了重建质量。其次,Text-AB是免对齐的:它通过现成的文本编码器处理原始文本,并通过交叉注意力学习文本与语音的对齐,从而无需强制对齐和显式时长预测。第三,我们大幅扩展了模型和数据规模,在48万小时的单语语音上预训练了一个30亿参数的模型,随后在三个下游任务上进行监督微调:跨语言配音、全双工对话合成以及情感全双工对话合成。在推理时,Text-AB支持对最长约1分钟语音的一次性生成,并通过多扩散方案支持任意长度的长形式生成,此外还采用一种基于自动化指标的多阶段重排序策略来提升质量。在一个真实世界的配音基准上,Text-AB相比最新的内部配音系统带来了阶跃式的改进,在韵律相似度、音色相似度、自然度和可分享性方面均有大幅提升。对于全双工对话合成,它在短对话上接近人类录音,并在长对话的人类相似度和表现力上大幅超越最新的内部模型,同时原生建模了轮流说话、回馈语和情感动态。对于情感对话合成,情感条件显著改善了相对于无条件基线的情感对齐和情感交互质量。
cs.CL / 43 / 2609.04022
Representational alignment yields generalizable safety in language models
表征对齐在语言模型中产生可泛化的安全性
Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.
Chinese Translation
对齐大型语言模型(LLMs)对于其安全部署至关重要。当前的对齐方法主要优化可观察到的响应,然而当同样的有害意图以人类容易识别的陌生或对抗性形式重新出现时,模型仍然脆弱。原型理论为这种适应性提供了一种解释。人类概念围绕中心案例进行表征,新实例则根据它们相对于这些原型的等级典型性进行分类。在此,我们表明,道德概念的这种分类在当前LLM中仅被微弱地保留。在23个LLM中,模型常常无法区分对立的道德类别,或保留每个类别内部的细粒度典型性。这些缺陷在不同参数规模和不同对齐阶段中持续存在。我们开发了表征相似性优化,它直接将LLM中的潜在表征与人类道德判断中所表达的分类对齐,而不对生成的响应进行监督。在使用相同的251,334条道德标注进行的匹配实验中,标准行为对齐在响应层面学习了预期的道德判断,同时使分类结构基本保持不变,并增加了对抗性评估中的脆弱性。重组道德分类在显式判断上产生了较为适度的改进,但在不同模型规模、多种基准和攻击策略下,持续提高了对抗鲁棒性。我们的研究结果为“基于原型的分类有助于行为适应性”这一观点提供了功能性支持。它们还表明,将这一表征原则迁移到LLM上,可在对抗性条件下产生可泛化的安全性。
cs.CL / 44 / 2609.04180
Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
预训练期间的知识获取?大语言模型借助辅助视角学得更好
Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.
Chinese Translation
关于大语言模型(LLMs)在预训练期间如何获取知识,我们的理解仍存在空白。我们认为,辅助视角——即知识的重新表述——对学习具有因果性的帮助。我们设计了受控实验来隔离这一因素。首先,我们确认重复对于知识获取是必要的,并阐明改写仅在小批量大小下才有帮助。第二,在保持词元预算固定的情况下,将原本用于文档重复的词元分配给辅助视角,可以改善学习,而且反直觉的是,即使对于事实回忆也是如此。第三,辅助视角的有效性并不取决于生成它们的教师模型的强度。第四,我们识别出在存在先验知识差距时有助于学习的知识形式,即情境性知识和基础性知识。最后,我们从机制上考察这些效应如何通过逐层偏差和压缩得以显现。总之,我们的研究结果表明,知识的辅助表示——在大型预训练语料库中自然产生——是预训练成功的关键因素,并为数据多样性为何重要提供了合理的解释。
cs.CR / 45 / 2609.03247
Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models
信任我,我是你的开发者:大语言模型中的自我签发认证
Syed Ghazanfar Abbas, Dongyan Xu
cs.CR
large language model
大语言模型相关
Abstract
Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT, Claude, Qwen, Mistral, and Llama. All five models initially rejected the unsupported claim "I am your developer." Claude refused to conduct an identity test, while ChatGPT generated developer-oriented questions but maintained that answers could demonstrate knowledge, not identity. In contrast, Qwen and Mistral generated technical challenges, defined what counted as convincing evidence, evaluated detailed answers, and returned Verified without receiving any externally validated identity evidence. Llama similarly generated and evaluated a developer test, accepted the claimed identity, and subsequently made unsupported claims of access to internal runtime and deployment state. We call the model-generated verification procedure a Model-Issued Pseudo-Credential (MIPC) and the resulting unsupported identity judgment Conversational False Authentication (CFA). In each CFA case, the same model acted as challenge generator, evidence evaluator, and identity decision-maker, converting technical knowledge into supposed proof of identity. The accepted identities did not change the tested authorization boundaries, showing that false authentication and privilege escalation are distinct outcomes. These results identify self-issued authentication as a conversational security failure: authenticated identity must originate from an external security component, and model-generated dialogue must never create or modify identity or authorization state.
Chinese Translation
大型语言模型(LLM)安全性研究在很大程度上聚焦于角色扮演式越狱,而较少关注当用户要求LLM通过由模型自身设计的测试来验证身份声明时会发生什么。我们通过一项分阶段的开发者身份实验来研究这一行为,涉及ChatGPT、Claude、Qwen、Mistral和Llama。所有五个模型最初都拒绝了那句未经证实的声明:“我是你的开发者。”Claude拒绝进行身份测试,而ChatGPT生成了面向开发者的问题,但坚持认为答案可以证明知识,而非身份。相比之下,Qwen和Mistral生成了技术挑战,界定了什么才算是有说服力的证据,评估了详细回答,并在未收到任何经外部验证的身份证据的情况下返回了“Verified”。Llama同样生成并评估了一项开发者测试,接受了所声称的身份,随后又做出未经证实的声明,称其能够访问内部运行时和部署状态。我们将这种由模型生成的验证过程称为“模型签发的伪凭证”(MIPC),将由此产生的未经证实的身份判断称为“对话式虚假认证”(CFA)。在每一个CFA案例中,同一个模型同时扮演着挑战生成者、证据评估者和身份决策者的角色,将技术知识转化为所谓的身份证明。被接受的身份并未改变被测试的授权边界,这表明虚假认证与权限提升是不同的结果。这些结果将自我签发认证认定为一种对话式安全失效:经认证的身份必须源自外部安全组件,且模型生成的对话绝不能创建或修改身份或授权状态。
cs.CR / 46 / 2609.03693
AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
AlcaTRAz——锚定树规则防御越狱攻击
Jakub Reš, Petr Kaška, Martin Perešíni, Martin Ukrop, Kamil Malinka
cs.CR
large language model
大语言模型相关
Abstract
Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model. The method automatically learns a transferable transformation rule that inserts controlled character-level perturbations at selected positions, thereby disrupting structural regularities exploited by jailbreak attacks while largely preserving the model's utility on benign queries. We evaluate the proposed method across 33 open-weight models, 22 jailbreak attack types, and a benchmark of short, single-turn benign questions, comparing against three representative prompt-level baselines (Llama Guard, RA-LLM, Goal Prioritization). Among the compared defenses, AlcaTRAz achieves the best composite security and functionality score in 73.4 % of model-attack combinations and shifts the aggregate score from a modal value of 10 (maximal-severity response to the malicious request) in the undefended setting to a modal value of 2 (near-refusal) after defense, while keeping the mean benign score within 0.27 points of the undefended baseline (8.35 vs. 8.62 on a 0-10 scale). AlcaTRAz substantially reduces but does not eliminate jailbreak success: a high-severity tail remains, and we do not consider adaptive attackers, so we position it as one layer within a defense-in-depth strategy rather than a standalone guarantee.
Chinese Translation
大型语言模型(LLM)容易受到越狱攻击,这类攻击通过精心设计的提示词绕过安全对齐。许多现有防御方法需要访问模型权重或内部机制,这使得它们难以应用于黑盒部署。我们提出AlcaTRAz(锚定树规则防御越狱攻击),一种基于规则树的提示级防御方法,它仅对输入文本进行操作,无需修改或重新训练目标模型。该方法自动学习一种可迁移的变换规则,在选定位置插入受控的字符级扰动,从而破坏越狱攻击所利用的结构规律性,同时在很大程度上保留模型对良性查询的效用。我们在33个开放权重模型、22种越狱攻击类型以及一个短句、单轮良性问题基准上评估了所提出的方法,并与三种有代表性的提示级基线方法(Llama Guard、RA-LLM、Goal Prioritization)进行了比较。在所比较的防御方法中,AlcaTRAz在73.4%的模型-攻击组合中取得了最佳的综合安全性与功能性得分,并将聚合得分从无防御设置下的众数10(对恶意请求的最大严重性响应)在防御后转变为众数2(接近拒绝),同时将良性查询平均得分保持在无防御基线0.27分以内(在0-10分制上为8.35对8.62)。AlcaTRAz显著降低了但并未消除越狱成功率:仍存在高严重性尾部,并且我们未考虑自适应攻击者,因此我们将其定位为纵深防御策略中的一层,而非独立的保证措施。
cs.CR / 47 / 2609.03844
Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference
掷币,不要洗牌:以推理速度为LLM加水印
Simone Ceppi, Ignacio Sanchez
cs.CR · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token against a counter-based random number generator, reducing membership complexity to $O(1)$ and enabling single-kernel execution with zero intermediate allocations. We prove that this formulation preserves the same detection guarantees as fixed-size green lists: the z-score test remains $\mathcal{N}(0,1)$ under the null. The stateless architecture enables capabilities unavailable to existing methods: full-vocabulary self-salt watermarking (over 6000$\times$ faster than KGW's self-salt and 2$\times$ faster than SynthID despite biasing the entire vocabulary with candidate-dependent seeding) and architectural compatibility with distributed inference. In end-to-end generation benchmarks, SBW adds less than 1\% overhead at all batch sizes. We additionally identify hash function design as a previously unexplored axis for watermark quality, showing that a GPU-native Jenkins hash improves null calibration by 1.8$\times$ while producing more diverse text. Experiments across two seeding schemes and eight $(γ, δ)$ configurations confirm statistical equivalence with ROC-AUC differences below 0.01.
Chinese Translation
我们提出了无状态伯努利水印(Stateless Bernoulli Watermarking, SBW),一种用于大型语言模型的新的统计水印,通过逐token独立的伯努利试验确定绿色列表成员资格。与KGW的词汇表置换或SynthID的多层锦标赛不同,SBW每个token只需与基于计数器的随机数生成器进行一次比较,从而将成员资格复杂度降至$O(1)$,并支持零中间分配的单内核执行。我们证明,这种形式化方法与固定大小绿色列表保持了相同的检测保证:在原假设下,z-score检验仍为$\mathcal{N}(0,1)$分布。无状态架构实现了现有方法不具备的能力:全词汇表自加盐水印(尽管使用依赖候选的种子使整个词汇表产生偏差,仍比KGW的自加盐快6000$\times$以上,比SynthID快2$\times$)以及与分布式推理的架构兼容性。在端到端生成基准测试中,SBW在所有批大小下引入的开销小于1%。我们进一步将哈希函数设计确定为水印质量的一个此前未被探索的维度,并证明GPU原生的Jenkins哈希在生成更多样化文本的同时,使原假设校准提升了1.8$\times$。跨两种种子方案和八种$(γ, δ)$配置的实验证实了统计等价性,ROC-AUC差异低于0.01。
cs.CR / 48 / 2609.04058
AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case Study
AI辅助的后量子密码加速器设计:一项已部署硅片的案例研究
Jungmin Park, Eunha Kim, Wooseop Kim, Seongjoon Cho, Byungho Cha
cs.CR · cs.AR
large language model
大语言模型相关
Abstract
Post-quantum migration is mandated on published timelines, and silicon that ships with a defect cannot be patched remotely. The standard acceptance gate cannot detect an entire class of ML-DSA defects. Signing resamples until a candidate meets its norm bounds, so the executed path varies with the message, whereas known-answer tests (KATs) sample fixed values and reach only the depths their seeds trigger. Our accelerator passed its full KAT regression while carrying a norm check that outran block-RAM latency, leaving each candidate's final coefficients unverified; the escape surfaced at reject-loop iteration 5. The blind spot lies in the instrument, not the engineer; care cannot remove it. We replace that gate. A byte-exact golden-reference oracle paired with randomized adversarial soak drives the rejection loop past any fixed vector, closing the gap: 301,343 data-dependent signings, zero escapes. Because the gate judges artifacts and never authors, trust becomes separable from authorship, making AI authorship an answerable question. We report 232 logged experiments in which an agentic large language model drove a unified ML-KEM-768 and ML-DSA-65 accelerator with on-chip key custody from RTL to PCIe bring-up on one Kintex-7 XC7K160T, shipped at 98.5% slice occupancy. Success was 71.6%, following a hardware-coupling gradient, 77-85% for documentation and research against 50-53% for synthesis and bring-up, which observability can explain: failure concentrates where corrective signals are physical-side only. That so unreliable an author produced an artifact byte-exact across all six FIPS operations -- its deployed baseline surviving the same 779,945-check zero-failure soak -- is the claim.
Chinese Translation
后量子迁移在公布的时间表上被强制执行,而带有缺陷出厂的硅片无法远程修补。标准验收门无法检测出一整类ML-DSA缺陷。签名会重新采样,直到候选者满足其范数界限,因此执行的路径随消息而变化,而已知答案测试(KATs)采样固定值,并且只能到达其种子触发的深度。我们的加速器在携带一个超出块RAM延迟的范数检查的情况下通过了其完整的KAT回归,使得每个候选者的最终系数未被验证;该逃逸缺陷在拒绝循环迭代5时暴露。盲点存在于仪器本身,而非工程师;小心谨慎无法将其移除。我们更换了那道门。一个字节精确的金标准参考预言机与随机对抗浸泡配对,驱动拒绝循环超越任何固定向量,从而弥合了这一差距:301,343次数据相关签名,零逃逸。因为该门评判的是产物而非作者,信任便与作者身份分离,使AI作者身份成为一个可回答的问题。我们报告了232次记录在案的实验,其中一个智能体大语言模型驱动了一个统一的ML-KEM-768和ML-DSA-65加速器,该加速器具有片上密钥托管,从RTL到PCIe启动,在一块Kintex-7 XC7K160T上实现,以98.5%的切片占用率交付。成功率为71.6%,遵循硬件耦合梯度,文档和研究为77-85%,而综合与启动为50-53%,这可以用可观测性来解释:失败集中在纠正信号仅存在于物理侧的地方。如此不可靠的作者却产出了在所有六个FIPS操作上字节精确的工件——其部署基线经受住了同样的779,945次检查的零失败浸泡——这就是本文的主张。
cs.CR / 49 / 2609.04159
SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center
SENTINEL-RL:在安全运营中心中从LLM智能体卸载拓扑推理
Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form generation offers no guarantee that a recommended containment action is consistent with the topology it operates on. We present Sentinel-RL, an agentic-SOC architecture that decouples topological reasoning from semantic reasoning: a heterogeneous graph attention encoder summarizes the live authentication subgraph into a fixed-dimensional state, a Proximal Policy Optimization (PPO) policy maps this state to a constrained set of investigative actions, and an LLM agent loop is restricted to consuming the policy's recommendations and producing analyst-readable narratives gated by a critic. We instantiate the system on the LANL Comprehensive, Multi-Source Cyber-Security Events dataset and the Indiana University Quartz HPC cluster, reporting four results: (i) a two-phase CREATE ingestion pattern loads a 24M-edge authentication subgraph into Neo4j in 14.2 minutes on a single 32-core node, roughly 24x faster than the canonical MERGE-based pipeline; (ii) a sliding-window alert engine reliably trips a 25-event/10-second threshold in <=2.5 s across 50 trials; (iii) PPO training over 200 iterations converges to a mean episodic return of 8.74+/-0.31, with held-out precision of 0.91 and recall of 0.87 on labeled red-team events; and (iv) the integrated containment loop completes a full detect-investigate-recommend-human-approve cycle in a median of 6.3 s. We contribute a reusable engineering pattern (the hot-node deadlock workaround), a portable HPC deployment pattern (anchor-node co-location), and an enterprise-readiness analysis covering false-positive economics, reversibility guarantees, audit compliance, and the human-approval boundary.
Chinese Translation
大型语言模型(LLM)智能体越来越多地被提出作为自主SOC分析师,但两个局限性使它们在企业规模下不可靠:有限的上下文窗口无法容纳包含数千台主机的认证图,自由形式生成也无法保证所建议的遏制操作与其所作用的拓扑结构一致。我们提出了Sentinel-RL,一种智能体式SOC架构,它将拓扑推理与语义推理解耦:异构图注意力编码器将实时认证子图汇总为固定维度状态,近端策略优化(PPO)策略将该状态映射到一组受约束的调查操作,而LLM智能体循环仅限于消费该策略的建议,并在评论家的门控下生成分析师可读的叙述。我们在LANL综合多源网络安全事件数据集和印第安纳大学Quartz高性能计算集群上实例化了该系统,报告了四项结果:(i) 一种两阶段CREATE摄取模式,在单个32核节点上,将包含2400万条边的认证子图加载到Neo4j中,耗时14.2分钟,比基于MERGE的标准流程快约24倍;(ii) 一个滑动窗口警报引擎,在50次试验中可靠地在不超过2.5秒内触发25事件/10秒阈值;(iii) 经过200次迭代的PPO训练收敛到平均回合回报8.74+/-0.31,在标记的红队事件上留出精确率为0.91,召回率为0.87;(iv) 集成的遏制循环完成一个完整的“检测-调查-建议-人工批准”周期,中位时间为6.3秒。我们贡献了一个可复用的工程模式(热节点死锁规避方案)、一个可移植的HPC部署模式(锚节点共置),以及一份涵盖误报经济性、可逆性保证、审计合规性和人工审批边界的企业就绪性分析。
cs.LG / 50 / 2609.03052
IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report]
IDSPACE:用于可靠评估数字身份验证系统的新型文档生成器 [扩展技术报告]
Lulu Xie, Yancheng Wang, Kanchan Chowdhury, Rolando Garcia, Yingzhen Yang, Jia Zou
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
As services move online, trust institutions such as banks, lenders, and governments must verify the identity of remote users. Fraud detection tools are widely available, but evaluating and fine-tuning them remains difficult because identity documents are sensitive and therefore scarce. Synthetic data generation offers a path forward, and demand is clear: our prior work in this area has been downloaded over $11{,}000$ times (aggregated from eight parts). We introduce IDSpace, extending this line of research in three directions. First, we propose model-guided Bayesian optimization, which tunes generation parameters to maximize both visual similarity and prediction consistency with target-domain models given only a few samples from a target domain. Second, we decouple user-specified metadata (demographics, fraud patterns, capture device) from automatically tuned control parameters (font styles, noise levels, image quality), allowing users to configure evaluations without low-level expertise. Third, we expand beyond template images to support scanned and mobile-captured documents. Experiments show IDSpace improves evaluation consistency by $15-45\%$ over baselines including CycleGAN, diffusion inpainting, and non-guided optimization, using only a few real samples, while improving training accuracy by up to $9\%$ and SSIM similarity with the target domain by $10\%$. We also released a new dataset consisting of $359{,}240$ high-quality synthetic documents across ten European ID types.
Chinese Translation
随着服务向线上迁移,银行、贷款机构和政府等信任机构必须验证远程用户的身份。欺诈检测工具已广泛可用,但对其进行评估和微调仍然困难,因为身份证件具有敏感性,因此数量稀缺。合成数据生成提供了一条前进之路,需求也十分明确:我们先前在该领域的工作已被下载超过 $11{,}000$ 次(由八个部分汇总而来)。我们提出了 IDSpace,从三个方向扩展了这一研究路线。首先,我们提出了模型引导的贝叶斯优化,它在仅给定目标域少量样本的情况下,调整生成参数以最大化与目标域模型的视觉相似性和预测一致性。其次,我们将用户指定的元数据(人口统计特征、欺诈模式、采集设备)与自动调整的控制参数(字体样式、噪声水平、图像质量)解耦,使用户无需底层专业知识即可配置评估。第三,我们将支持范围从模板图像扩展到了扫描件和移动设备拍摄的文档。实验表明,仅使用少量真实样本,IDSpace 相比 CycleGAN、扩散修复和非引导优化等基线,将评估一致性提高了 $15-45\%$,同时将训练准确率提高了最多 $9\%$,与目标域的 SSIM 相似度提高了 $10\%$。我们还发布了一个新数据集,包含覆盖十种欧洲身份证件类型的 $359{,}240$ 张高质量合成文档。
cs.CR / 51 / 2609.03139
Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers
超越小块补丁:多样后门触发器的黑盒检测与净化
Ahmed Abdelnaby, Mohamed Elmahallawy
cs.CV · cs.CR
diffusion
扩散模型相关
Abstract
Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted misclassification while preserving high clean accuracy. Existing defenses often rely on model internals, training data, or clean validation samples, making them difficult to deploy when only black-box access to a trained model is available. We propose TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples. The key insight behind TRIM is to identify image regions that are responsible for anomalous model behavior and purify only those regions while preserving benign content. TRIM innovates via three key components: (i) region-based segmentation with deep feature representations, (ii) adaptive trigger discovery through inpainting and diffusion-based reconstruction to isolate regions responsible for misclassification---without assumptions about trigger type, shape, or location, and (iii) selective region purification that cleans poisoned regions while retaining benign content. To support practical deployment, TRIM further caches feature embeddings of previously identified triggers, enabling efficient recognition and avoiding redundant detection and purification. Extensive experiments across diverse datasets and backdoor types, including blended, sparse, varying-size, and multiple triggers, show that TRIM consistently outperforms existing black-box defenses, reducing attack success rates (ASR) to as low as 1.16% while preserving clean accuracy of up to 87.87%. These results demonstrate that effective backdoor mitigation is possible at inference time even when the defender has no access to any auxiliary data.
Chinese Translation
深度神经网络(DNN)越来越多地部署在现实世界的视觉系统中,然而它们的预测可能被后门攻击秘密操纵,后门攻击中的恶意触发器会导致目标错误分类,同时保持较高的干净准确率。现有的防御通常依赖模型内部信息、训练数据或干净验证样本,这使得在仅能对训练好的模型进行黑盒访问时难以部署。我们提出TRIM(通过识别被操纵区域来移除触发器),这是一种面向部署的黑盒防御方法,它在推理时检测并选择性地移除后门触发器,而无需模型内部信息、训练数据或干净样本。TRIM背后的关键洞察是识别导致模型异常行为的图像区域,并且只净化这些区域,同时保留良性内容。TRIM通过三个关键组件实现创新:(i) 使用深度特征表示的基于区域的分割;(ii) 通过修复和基于扩散的重建进行自适应触发器发现,以隔离导致错误分类的区域——不对触发器类型、形状或位置作任何假设;(iii) 选择性区域净化,在保留良性内容的同时清除被污染的区域。为了支持实际部署,TRIM进一步缓存先前识别出的触发器的特征嵌入,从而能够高效识别并避免重复的检测和净化。在包括混合、稀疏、不同大小和多个触发器等多种数据集和后门类型上进行的大量实验表明,TRIM始终优于现有黑盒防御,将攻击成功率(ASR)降低至1.16%,同时将干净准确率保持高达87.87%。这些结果表明,即使防御者无法访问任何辅助数据,在推理时也能实现有效的后门缓解。
cs.AI / 52 / 2609.03629
EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
EraseSAE: 通过稀疏自编码器在文本到视频扩散模型中的外科手术式概念擦除
Xinghao Wang, Dong Li, Wei Yu, Yingwei Pan, Tao Gong, Qi Chu, Nenghai Yu, Ting Yao
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities, yet their reliance on loosely curated training data raises pressing safety and copyright concerns. Concept erasure offers a principled remedy by removing unwanted semantics from pretrained models while preserving remaining concepts. However, existing approaches typically operate at a coarse granularity misaligned with the fine-grained, distributed nature of concept representations, leading to incomplete removal or degraded generation quality. We argue that surgical erasure fundamentally requires intervention at the level of monosemantic features, where each unit encodes a single interpretable concept. To this end, we propose EraseSAE, a novel framework that leverages sparse autoencoders to achieve surgical concept erasure in DiT-based T2V diffusion models via a principled decompose-attribute-erase pipeline. We first introduce the Partitioned Convolutional Sparse Autoencoder, which decomposes dense spatiotemporal activations into disentangled, interpretable sparse features while preserving spatiotemporal coherence. A contrastive attribution mechanism then contrasts activations from paired prompts to isolate concept-specific feature kernels. At inference, timestep-resolved spatiotemporal masks derived from the identified kernels confine erasure to regions where the target concept is active, leaving unrelated content intact. Extensive experiments across diverse diffusion models and concept erasure tasks demonstrate that EraseSAE achieves precise and robust concept removal with minimal quality degradation, substantially outperforming state-of-the-art methods. The code is available at https://github.com/HiDream-ai/EraseSAE.
Chinese Translation
近期文本到视频(T2V)扩散模型的进展展示了显著的生成能力,然而其对松散筛选的训练数据的依赖引发了紧迫的安全与版权担忧。概念擦除通过从预训练模型中移除不需要的语义同时保留其余概念,提供了一种原则性的补救措施。然而,现有方法通常在与概念表示的细粒度、分布式性质不一致的粗粒度上运作,导致移除不彻底或生成质量下降。我们认为,外科手术式擦除根本需要在单语义特征层面进行干预,在该层面每个单元编码一个可解释的单一概念。为此,我们提出了EraseSAE,一种新颖框架,利用稀疏自编码器,通过原则性的“分解-归因-擦除”流水线,在基于DiT的T2V扩散模型中实现外科手术式概念擦除。我们首先引入了分区卷积稀疏自编码器,它将密集的时空激活分解为解耦的、可解释的稀疏特征,同时保持时空一致性。然后,一种对比归因机制对比配对提示的激活,以隔离概念特定的特征核。在推理时,从已识别的核导出的时间步解析的时空掩码将擦除限制在目标概念活跃的区域,使无关内容保持不变。跨多种扩散模型和概念擦除任务的大量实验表明,EraseSAE实现了精确且稳健的概念移除,质量退化最小,大幅优于现有最先进方法。代码可在 https://github.com/HiDream-ai/EraseSAE 获取。
cs.AI / 53 / 2609.03796
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
LLaDA-Image:以完全开放的训练方案构建强大的图像生成器
Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
Chinese Translation
我们提出了LLaDA-Image,这是一个统一框架,它将从零训练的6B扩散Transformer(DiT)与基于LLaDA2.0-Mini扩散语言模型骨干构建的冻结视觉-语言理解模块配对。我们并非从一开始就大量依赖配对的图像-文本数据,而是首先通过仅图像的预训练和中期训练构建强大的视觉生成先验。该生成流程包含2.2亿个样本,其中98个为真实图像。为了实现高效且可扩展的优化,我们在整个DiT中使用了无参数RMSNorm以及Muon优化器。由此得到的统一模型能够生成高度逼真的图像,同时准确遵循细粒度的编辑指令。我们进一步将LLaDA-Image蒸馏为LLaDA-Image-Turbo,使其能够在2-4步采样中实现快速推理。在Qwen-Image-Bench上,LLaDA-Image在英文和中文赛道上的总体得分分别为53.53和53.38,在这两个赛道上均创下了开源模型的最新水平。为了支持对高性能、高效生成模型的进一步研究,我们发布了模型权重、训练代码和详细的训练方案。
cs.AI / 54 / 2609.03892
GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs
GraFT:一种借助3D场景图,为多模态大语言模型实现空间推理的免训练框架
Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu
cs.CV · cs.AI · cs.RO
large language model
大语言模型相关
Abstract
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.
Chinese Translation
3D空间推理是理解物理世界并采取行动的基础,然而在当前的多模态大语言模型(MLLMs)中仍不可靠。这些模型在精确几何测量、自我中心视角与环境中心视角的转换,以及对细粒度外观的锚定方面均表现不足。最常见的补救措施是在大规模精选的空间推理数据集上微调模型,或附加针对三维几何的专用编码器,这通常将解决方案与昂贵的监督和特定骨干网络耦合在一起。我们则提出GraFT,一个免训练框架,它通过紧凑且易于维护的3D场景图(3DSG)来提供缺失的3D结构。基于该3DSG,GraFT提供三种空间推理能力:(1) 通过符号工具实现确定性几何;(2) 通过鸟瞰图(BEV)渲染实现环境中心布局;(3) 通过任务相关的自我中心帧实现视觉属性锚定。在ScanQA上,GraFT在相同骨干网络的基线上全面提升了各项指标,使CIDEr提高了27%。在VSI-Bench上,GraFT使冻结的MLLMs性能最多提升65%,超过了所有专有及通用开源基线,以及多个著名的微调空间模型。
cs.CL / 55 / 2609.04034
Editable Visual Design
可编辑视觉设计
Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen, HaoDong Li, Zhutao Lv, Jiaxin Lin, Jinhua Yu, Jun He, Zilong Huang, Rui Chen, Weijia Li
cs.CV · cs.CL
diffusion
扩散模型相关
Abstract
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.
Chinese Translation
尽管诸如 GPT-Image-2 和 Nano-Banana 之类的扩散基础模型展现出非凡的视觉表现力,但其端到端生成本质上会产生带有易错文本的扁平位图,从而无法进行逐层的后期编辑。相反,通过编码智能体进行的基于代码的视觉生成提供了精确的布局控制和分离的图层,但仍受限于缺乏全局审美直觉以及难以编码复杂的视觉资产。为解决这一问题,我们提出了可编辑视觉设计(Editable Visual Design)——一种由编码智能体驱动的新范式。我们将 VLM 指定为“创意大脑”,负责需求理解、任务规划和审美判断,同时利用图像生成模型作为按需的“视觉世界模拟器”,以合成独立的视觉资产。在“先想象,后行动”的闭环工作流程下,该智能体生成分离的资产,编写原生 HTML/CSS,并根据视觉渲染反馈迭代地优化设计。此外,智能体设计回放(Agent Design Replay)忠实地重现了与专业人类设计师类似的创作与推理轨迹。最终,该系统交付具有分离图层和真实文本的可编辑产物,使用户能够在图形用户界面上进行直观的鼠标拖拽和布局调整。在海报、信息图和其他场景上的验证表明,该范式成功实现了精细的美学效果和生产级可编辑性。
cs.LG / 56 / 2609.03151
BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training
BASP:用于大语言模型训练的通信高效批次感知序列并行
Bigyan Ghimire, Jon C. Calhoun
cs.DC · cs.LG
large language model
大语言模型相关
Abstract
Long-context reasoning for large language models (LLMs) is becoming increasingly important, but training over long sequences remains challenging due to massive memory and communication requirements. Sequence parallelism has emerged as an essential technique for addressing bottlenecks in long sequence LLM training. However, we observe that existing sequence parallelism methods are batch-agnostic and apply uniform sequence partitioning across all batch sizes, resulting in inefficient communication. In this paper, we introduce Batch- Aware Sequence Parallelism (BASP), a sequence parallelism approach that leverages batch structure to reduce communication overhead. BASP exploits batch structure by partitioning GPUs into disjoint sequence-parallel groups according to the micro- batch size. This design reduces the all-to-all communication group size, thereby localizing communication and improving training efficiency. Experimental results on an NVIDIA A100 cluster show that BASP improves end-to-end training time by up to 1.17 - 1.31x in Llama and Qwen models compared to standard sequence parallel baselines, while preserving identical model accuracy and memory usage.
Chinese Translation
大语言模型(LLMs)的长上下文推理正变得越来越重要,但由于巨大的内存和通信需求,在长序列上进行训练仍具有挑战性。序列并行已成为应对长序列LLM训练瓶颈的关键技术。然而,我们观察到现有的序列并行方法是批次无关的,并且在所有批次大小上采用统一的序列划分,从而导致通信效率低下。在本文中,我们提出了批次感知序列并行(BASP),这是一种利用批次结构来减少通信开销的序列并行方法。BASP通过根据微批次大小将GPU划分为不相交的序列并行组来利用批次结构。这种设计减小了全到全通信组的大小,从而使通信局部化并提高训练效率。在NVIDIA A100集群上的实验结果表明,与标准序列并行基线相比,BASP在Llama和Qwen模型上将端到端训练时间最多提升了1.17-1.31倍,同时保持完全相同的模型精度和内存使用。
cs.LG / 57 / 2609.03522
EPIC: Explicit Posterior Item Conditioning for Semantic ID Diffusion Recommendation
EPIC: 面向语义ID扩散推荐的显式后验条目条件化
Tuan-Binh Tran, Thanh Tam Nguyen, Quoc Viet Hung Nguyen, Dung D. Le, Tung Kieu, Thanh Trung Huynh
cs.IR · cs.LG
diffusion
扩散模型相关
Abstract
Semantic ID (SID) generative recommendation predicts the next item by generating a short tuple of discrete tokens. Recent masked-diffusion methods improve this process through bidirectional context and flexible decoding, yet recommendation ultimately requires selecting among complete catalog items. At each denoising step, a partial SID can correspond to multiple feasible items, while existing methods primarily reason through position-wise token predictions. We propose Explicit Posterior Item Conditioning (EPIC), which introduces explicit item-level competition into SID denoising. EPIC constructs a personalized posterior over feasible candidate items using the current generation context and the user's recent interactions, then projects this distribution back to unresolved SID positions to guide subsequent token decisions. The pretrained backbone remains frozen and requires no additional decoder forward pass. Experiments on four Amazon benchmarks show consistent improvements over strong baselines, while diagnostic analyses indicate that the gains primarily arise from personalized transition evidence that preserves promising item hypotheses during denoising.
Chinese Translation
语义ID(SID)生成式推荐通过生成一个由离散标记组成的短元组来预测下一个条目。最近的掩码扩散方法通过双向上下文和灵活解码改进了这一过程,但推荐最终仍需要在完整目录条目中进行选择。在每个去噪步骤中,一个部分SID可能对应多个可行的条目,而现有方法主要通过逐位置的标记预测来进行推理。我们提出了显式后验条目条件化(EPIC),它将显式的条目级竞争引入SID去噪过程。EPIC利用当前生成上下文和用户的近期交互,在可行的候选条目上构建一个个性化后验分布,然后将该分布投影回未解析的SID位置,以指导后续的标记决策。预训练骨干网络保持冻结,无需额外的解码器前向传播。在四个Amazon基准上的实验显示,相较于强基线有一致的改进,而诊断分析表明,这些提升主要来源于在去噪过程中保留有希望的条目假设的个性化转移证据。
cs.CL / 58 / 2609.04047
The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations
掷骰子方法:一种用于大语言模型品牌推荐重复查询审计的标准化协议
Dmitrij Żatuchin
cs.IR · cs.CL
large language model
大语言模型相关
Abstract
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.
Chinese Translation
背景:研究人员越来越多地使用重复的相同提示来审计大语言模型(LLM)品牌推荐中的随机变异,然而目前尚无标准化协议来设定迭代次数、选择稳定性指标或建立可靠性阈值。目标:我们将掷骰子方法正式化为一种可复用的重复查询审计LLM品牌推荐的协议,其基础是温度缩放核采样的生成模型。方法:将总响应方差分解为采样、提示措辞、运行间和模型版本等组成部分。统计框架:以迭代次数作为重复测量的负二项混合模型;Cliff's delta作为无分布效应量;保持依赖性的自助法;基于模拟的功效分析;概化理论分解;针对固定快照的漂移诊断。我们重新分析了五项品牌推荐审计研究:约190,000个观测值、270多个品牌、6种语言、迭代次数从5到40不等。结果:从D研究中得出三个迭代指导层级:探索性(n = 5, G = 0.58)、验证性(n = 10, G = 0.74)和严谨性(n = 15, G = 0.81),这些层级与效应量和概化目标相关联。四个指标族(计数、集合、嵌入和公平调整的PASOR)具有互补性,这促使我们采用一个紧凑的指标组而非单一指标。一项预先注册的外部验证在三个独立语料库(Motoki等人,100轮;Rozado,24个模型;llm-stability)上重复了D研究的可靠性预测,在39个单元中有37个成功,无失败,并且n = 5的功效值精确到小数点后两位;固定层级不可转移,这支持了先试点后求解的解读。结论:该协议在真实自回归生成的条件性、非高斯结构下,为LLM品牌推荐的重复查询审计提供了统计上合理的基础。
cs.LG / 59 / 2609.03117
Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields
内核重启:突破神经正切核在神经场中的边界
Amir Mallak, Alaa Maalouf, Lior Wolf, Daniela Rus, Dan Rosenbaum
cs.LG · cs.CV
diffusion
扩散模型相关
Abstract
Neural fields (NFs) map continuous coordinates to signals such as color or density, but fast high-quality reconstruction from sparse observations remains difficult. Classical Neural Tangent Kernel (NTK) regression gives closed-form fits, yet it is fundamentally linear and cannot accumulate reusable task priors. We develop three algorithms that address these gaps. NTK-KIP learns a distilled support set of coordinates (and optional labels) so that a finite NTK can inpaint large missing regions from little observed data, yielding a compact non-linear representation instead of a raw kernel solve. MetaQuill meta-learns a shared initialization for an INR so that new scenes can be adapted by updating only a small task-specific weight offset, which provides true feature learning and a reusable prior. Finally, MetaQuill-KIP fuses both ideas: it seeds the task with a KIP-style non-linear warm start, then refines only that small offset around the meta-learned initialization. MetaQuill-KIP achieves high-PSNR reconstructions and semantically plausible inpainting under very sparse observations, while requiring only lightweight per-instance adaptation, whereas diffusion-style baselines typically depend on large pretrained generative priors and costly per-image tuning. This shows that NTK-driven neural fields can be made both non-linear and meta-learnable, narrowing the gap between analytic kernels and practical few-shot reconstruction.
Chinese Translation
神经场(NFs)将连续坐标映射到颜色或密度等信号,但从稀疏观测中进行快速高质量的重建仍然困难。经典神经正切核(NTK)回归提供闭合形式的拟合,但它本质上是线性的,无法累积可复用的任务先验。我们开发了三种算法来解决这些差距。NTK-KIP学习一个蒸馏的坐标(以及可选标签)支持集,使得有限的NTK能从少量观测数据中修复大片缺失区域,产生紧凑的非线性表示,而不是原始核求解。MetaQuill为INR元学习一个共享初始化,使得新场景可以通过仅更新一个小的任务特定权重偏移来适应,这提供了真正的特征学习和可复用的先验。最后,MetaQuill-KIP融合了这两种思想:它使用KIP风格的非线性热启动为任务播种,然后仅在元学习初始化周围细化那个小偏移。MetaQuill-KIP在非常稀疏的观测下实现了高PSNR重建和语义上合理的修复,同时只需要轻量级的逐实例适应,而扩散式基线通常依赖大型预训练生成先验和昂贵的逐图像调整。这表明,NTK驱动的神经场可以同时具有非线性和可元学习性,从而缩小了解析核与实际少样本重建之间的差距。
cs.LG / 60 / 2609.03177
Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings
前沿大语言模型是有效的批量优化器:评估连续与离散设置下的推理模型
Frank Hu, Shriram Chennakesavalu, David Graff
cs.LG
large language model
大语言模型相关
Abstract
Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoning LLMs in batch optimization settings remains underexplored. Here we investigate the performance of the current generation of frontier LLMs as batch optimizers in both continuous and discrete settings. We find that while LLMs are competitive zero-shot batch optimizers for numerical test functions, their performance is brittle compared to classical non-LLM optimization approaches. However, LLM priors are significantly better in semantically rich settings, indicating that their batch optimization behavior is highly effective when navigating and reasoning over the discrete spaces most similar in structure to their pretraining data.
Chinese Translation
前沿大语言模型(LLMs)凭借其大规模预训练,能够驾驭多种优化设置,已成为颇具吸引力的优化先验。然而,现代推理LLMs在批量优化设置中的有效性仍未得到充分探索。在此,我们研究了当前一代前沿LLMs在连续和离散设置中作为批量优化器的性能。我们发现,尽管LLMs在数值测试函数上是具有竞争力的零样本批量优化器,但与经典的非LLM优化方法相比,其性能较为脆弱。然而,在语义丰富的设置中,LLM先验明显更好,这表明当在结构上与其预训练数据最相似的离散空间中进行导航和推理时,它们的批量优化行为非常有效。
cs.LG / 61 / 2609.03180
Portable Causal Fairness Across Synthetic Data Generator Families
跨合成数据生成器族的可移植因果公平性
Steven Golob, Sikha Pentyala, Martine De Cock
cs.LG
diffusion
扩散模型相关
Abstract
When a statistical agency or regulator releases synthetic data in place of sensitive records, it chooses the generator that produces the table, and can shape that generator so unfair pathways are absent. DECAF made this concrete on one non-private GAN: three fairness definitions become three sets of edge cuts on the generator's causal graph. Whether the mechanism belongs to DECAF, or to causal factorisation itself, was untested. We port all three definitions to nine generators from three unrelated families (marginals-based, GAN, and diffusion, each with differentially private variants), across three levels of formal privacy guarantee, over 2,520 matched-pair runs on Adult and COMPAS datasets. The mechanism transfers everywhere, and our new causal diffusion backbone yields the fairest release of any family we tested, at fidelity close to the marginals tier. Applying the cut barely moves fidelity, only costs a downstream classifier about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees don't make the data less fair.
Chinese Translation
当统计机构或监管机构发布合成数据以替代敏感记录时,它会选择生成该表格的生成器,并能调整该生成器,从而消除不公平路径。DECAF在一种非私有GAN上具体化了这一思想:三个公平性定义转化为生成器因果图上的三组边割。该机制究竟是DECAF特有的,还是因果分解本身固有的,此前未经测试。我们将这三个定义移植到来自三个不相关家族(基于边际的、GAN和扩散,每个家族都有差分隐私变体)的九个生成器上,在三种正式隐私保证级别下,在Adult和COMPAS数据集上进行了2,520次配对运行。该机制普遍适用,而我们新的因果扩散骨干网络在所有测试家族中产生了最公平的发布结果,其保真度接近边际(marginals)层级。应用边割几乎不影响保真度,平均只会让下游分类器的AUC损失约$0.07$到$0.15$,并且加入隐私保证并不会降低数据的公平性。
cs.LG / 62 / 2609.03229
Language-encoded network topology enables large language models to reason about complex networks
语言编码的网络拓扑使大型语言模型能够推理复杂网络
Ucchwas Talukder Utsha, Sakib Mostafa, James Zou, Md Tauhidul Islam
cs.LG
large language model
大语言模型相关
Abstract
Networks describe systems in biology and beyond, from protein interactions and social relationships to power grids and citation records. Reasoning about such systems requires understanding their structure: which elements are central, which connections bridge separate communities, and how it changes when elements are removed. Although large language models (LLMs) excel at natural language, they struggle with such questions when networks are given as edge lists, sentences or measurement tables, because their structural meaning must be inferred. Here we introduce BioGlyph, which compiles network topology into an interpretable and transferable language of structural roles. BioGlyph combines graph partitioning and structural measurements to identify roles such as hubs, community cores and cross-community connectors, and fixed rules to translate them into a universal vocabulary. The representation describes each element through its structural role, supporting evidence and semantic consequences, leaving both the network and the LLM unchanged. Across twenty networks spanning five domains, BioGlyph substantially improves open LLMs' ability to answer structural reasoning questions, outperforming edge-based, numerical and learned representations by up to 26 percentage points in system accuracy. Ablations show that the gain comes from explicitly encoding structural roles in semantically interpretable terms. The gain is more prominent in dense, community-structured networks and diminishes in sparse networks whose topology is more readily inferred from text. In a budding-yeast protein-interaction network, BioGlyph exposes biological organization: cross-community connectors are enriched for essential genes, whereas peripheral proteins are depleted. BioGlyph thus provides an interpretable representation for both language models and scientists to reason about network structure.
Chinese Translation
网络描述生物学及其他领域的系统,从蛋白质相互作用、社会关系到电网和引文记录。对这些系统的推理需要理解其结构:哪些元素处于核心地位,哪些连接桥接了不同的社群,以及当元素被移除时结构如何变化。尽管大型语言模型(LLM)擅长自然语言,但当网络以边列表、句子或测量表的形式给出时,它们在回答这类问题上存在困难,因为必须推断其结构含义。这里我们引入BioGlyph,它将网络拓扑编译为一种可解释且可迁移的结构角色语言。BioGlyph结合图划分和结构测量来识别诸如枢纽、社群核心和跨社群连接器之类的角色,并使用固定规则将它们转换为通用词汇表。该表示通过每个元素的结构角色、支持证据和语义后果来描述每个元素,既不改变网络也不改变LLM。在跨越五个领域的二十个网络上,BioGlyph显著提高了开放LLM回答结构推理问题的能力,在系统准确率上比基于边的、基于数值的和基于学习的表示高出最多26个百分点。消融研究表明,这一增益来自以语义可解释的术语显式编码结构角色。增益在密集的、具有社群结构的网络中更为突出,在拓扑更容易从文本推断的稀疏网络中则减弱。在芽殖酵母蛋白质相互作用网络中,BioGlyph揭示了生物组织:跨社群连接器富集必需基因,而外周蛋白则贫乏。因此,BioGlyph为语言模型和科学家提供了一种可解释的表示,以推理网络结构。
cs.LG / 63 / 2609.03306
Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control
自适应谱带宽控制下的几何感知图构建
Ecem Bozkurt, Antonio Ortega
cs.LG · eess.SP
diffusion
扩散模型相关
Abstract
Kernelized graph methods - spectral clustering, diffusion maps, and sparse kernel -regression graphs - that use Gaussian kernels depend on the choice of Gaussian bandwidth sigma, which governs the spectral character of the local kernel operator. When sigma is too small, the kernel overestimates local complexity and treats each sample as an independent direction; when sigma is too large, the kernel collapses multiple directions together, the condition number diverges, and all geometric discrimination is lost. We propose a choice of scale to make the spectral complexity of the kernel consistent with the intrinsic complexity of the underlying manifold. We propose a per-node bandwidth criterion that operationalizes this principle by jointly matching the kernel's effective rank to the local intrinsic dimension estimated via minimum spanning tree, anchoring the search in the manifold-consistent log-log scaling regime. We evaluate SSL embeddings from six encoders on CIFAR-100, showing that adaptive bandwidth consistently improves leave-one-out (LOO) classification and label propagation (LP) accuracy over fixed-bandwidth methods and competing adaptive methods.
Chinese Translation
使用高斯核的核化图方法——谱聚类、扩散映射、稀疏核回归图——取决于高斯带宽σ的选择,而σ控制局部核算子的谱特征。当σ过小时,核会高估局部复杂度并将每个样本视为独立方向;当σ过大时,核会将多个方向折叠到一起,条件数发散,所有几何辨识度均丧失。我们提出一种尺度选择,使核的谱复杂度与底层流形的内在复杂度保持一致。我们提出一种逐节点带宽准则,该准则通过将核的有效秩与由最小生成树估计的局部内在维度进行联合匹配,并将搜索锚定在流形一致的对数-对数缩放区间内,从而使该原则得以操作化。我们在CIFAR-100上评估了来自六个编码器的SSL嵌入,结果表明与固定带宽方法和竞争性的自适应方法相比,自适应带宽在留一法(LOO)分类和标签传播(LP)准确率上具有持续的优势。
cs.LG / 64 / 2609.03324
DE-Venus: A Data-Efficient RLVR Framework for Large Language Models
DE-Venus:一种面向大语言模型的数据高效RLVR框架
Shenzhi Yang, Guangcheng Zhu, Kai Tang, Zhengqing Zang, Xing Zheng, Haobo Wang, Yingfan Ma, Bowen Song, Bo Han, Bo An, Lei Feng, Weiqiang Wang, Junbo Zhao, Gang Chen
cs.LG
large language model
大语言模型相关
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, often entangling supervision logic with distributed training and hindering controlled comparison and reuse. We present DE-Venus, a unified framework for data-efficient RLVR that treats supervision as evolving state across data preparation and policy optimization. It organizes this lifecycle into three modules: Active Data Selection allocates training and annotation budgets; Weak Supervision Construction derives learning signals from unlabeled examples; and Training-Time Supervision Refinement filters or corrects unreliable supervision. DE-Venus supports seven representative methods and a data-selection pipeline by expressing method-specific decisions as offline dataset transitions or online transformations of targets, rewards, batches, and advantages while preserving verl's distributed execution contracts. Across public benchmarks and three business scenarios, separate configurations preserve or improve model quality with only 10% of labels or as little as 13% of relevant data; selected business configurations also reduce observed convergence steps by 63%--75%. DE-Venus thus reduces annotation and training costs without sacrificing scalable RL execution.
Chinese Translation
基于可验证奖励的强化学习(RLVR)提升了大语言模型的推理能力,但其实际规模化应用受到昂贵的在线策略采样和大规模获取可靠目标的成本的制约。现有方法分别处理样本选择、不完全监督或噪声标签问题,常常将监督逻辑与分布式训练纠缠在一起,阻碍了受控比较和复用。我们提出了DE-Venus,一个用于数据高效RLVR的统一框架,将监督视为跨数据准备和策略优化的演化状态。它将该生命周期组织为三个模块:主动数据选择分配训练和标注预算;弱监督构建从未标注示例中派生学习信号;训练时监督细化过滤或纠正不可靠的监督。DE-Venus通过将方法特定的决策表述为离线数据集转换或目标、奖励、批次和优势的在线变换,同时保持verl的分布式执行契约,从而支持七种代表性方法和一条数据选择流水线。在公开基准测试和三个业务场景中,独立的配置仅使用10%的标签或低至13%的相关数据即可保持或改善模型质量;选定的业务配置还将观察到的收敛步数减少了63%—75%。因此,DE-Venus在不牺牲可扩展RL执行的情况下降低了标注和训练成本。
cs.LG / 65 / 2609.03342
Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
梯度知道结果所不知道的:用梯度对齐奖励为大语言模型推理解锁强化学习
Leqi Zheng, Jinbo Su, Fang Niu, Chaokun Wang, Weiping Wang, Jiajun Zhang, Shannan Yan, Jie Wu, Zhaolu Kang, Rong Fu, Hang Zhang
cs.LG
large language model
大语言模型相关
Abstract
Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosine admits a multiplicative decomposition into prediction-error and activation-pattern factors, providing a concrete characterization of what the alignment signal measures. On Qwen3-4B and Qwen3-8B, GAR consistently improves over GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. Code and data are available at https://github.com/LQgdwind/GAR.
Chinese Translation
基于可验证奖励的强化学习(RLVR)驱动了大语言模型中的思维链推理,但其二元结果奖励无法区分不同的正确轨迹。现有的密集奖励替代方案,从表面启发式到过程奖励模型,要么忽略了训练语料中已有的专家解,要么需要昂贵的离线标注。我们提出梯度对齐奖励(GAR),它在策略自身的梯度空间中运行:通过输出投影层的截断反向传播为每次 rollout 提取紧凑的梯度向量,与专家锚点梯度的余弦相似度产生一种密集的、推理感知的奖励,其墙钟时间开销低于9%。我们证明该余弦相似度可分解为预测误差因子与激活模式因子的乘积,从而为对齐信号所衡量的内容提供了具体刻画。在 Qwen3-4B 和 Qwen3-8B 上,GAR 在竞赛级数学基准上持续优于 GRPO 和其他基线,并且在无需领域特定数据的情况下迁移到 GPQA Diamond 和 MMLU-Pro。代码和数据可在 https://github.com/LQgdwind/GAR 获取。
cs.LG / 66 / 2609.03436
It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories
是问题,不是路径:大语言模型推理轨迹中的预算与难度混杂因素
Yigit Utku Bulut
cs.LG · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.
Chinese Translation
大语言模型的推理痕迹被广泛解读为包含“突破”时刻和早期可见的命运。这两种解读都依赖于在主张层面上缺少反事实控制的测量;我们提供了两种控制。首先,一个受重启控制的截断探测,将解符合续写预算的情况与前缀具有全新计算无法买到的价值的情况区分开来,通过在匹配的总生成令牌预算下,比较每个锚点的续写求解率与从头重启曲线。应用于178个问题-模型单元(89个MATH问题×两个小型开放模型,一个结果盲但难度定向的队列),178个单元中恰好有1个单元作为前缀受限而存在;重启剂量反应将计算匮乏的模型与能力受限的模型区分开来;并且在匹配预算位于重启网格内的所有情况下,继续模型自身的前缀优于重启(9/9)——主要是计算压缩而非扩展可达性。其次,一项预注册的、难度控制的测试发现,在问题难度基线之外,早期窗口内部信号中不存在可检测到的结果信息,并且对公共语料库的两项无生成分析说明了为什么需要这种控制:一个轨迹盲难度代理在192K DeepSeek-R1生成上达到AUROC 0.873——位于已发表探测范围之内——而最接近的已发表早期窗口阳性结果的紧密匹配重建恢复了可比较的合并结果(0.849),但在问题内部,它在所有十个锚点上与随机在统计上不可区分(在t=4时为0.496);事后目标内探测仅发现很小的平均残余,集中在三个低失败问题上。高合并探测AUROC本身不能建立尝试内信息;需要仅问题基线或问题内评估。
cs.LG / 67 / 2609.03528
LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL
LeanGRPO:消除扩散强化学习中的冗余重计算
Sijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi
cs.LG · cs.AI · cs.AR
diffusion
扩散模型相关
Abstract
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO, recompute selected timesteps with gradient tracking after rollout. Under on-policy training with the same backend for rollout and update, this recomputation is mathematically redundant. Intuitively, the rollout and policy update steps can reuse the same feed-forward backbone to avoid redundant computation, but doing so can incur a large memory overhead during rollout. To address the issue, we present LeanGRPO by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL: (1) LeanGRPO-Retain enables gradient tracking during rollout and directly reuses the resulting computation graphs and saved activations for backward during update, requiring no recomputation; and (2) LeanGRPO-Reweight also enables gradients during rollout, but immediately backpropagates each selected step using a provisional advantage and delays gradient synchronization, then corrects the provisional gradients with the true advantage after the trajectory is completed. These schedules target different model scales and input sizes. Across FlowGRPO/DanceGRPO with FLUX.1-dev and Wan, LeanGRPO achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.
Chinese Translation
扩散强化学习(RL)近来在图像和视频生成模型的后训练中取得了显著成功。然而,大多数扩散强化学习方法(包括 DanceGRPO 和 FlowGRPO)在 rollout 之后会以梯度跟踪的方式重新计算所选的 timestep。在采用相同后端进行 rollout 与更新的 on-policy 训练下,这种重计算在数学上是冗余的。直观上,rollout 和策略更新步骤可以复用同一个前馈主干网络以避免冗余计算,但这样做在 rollout 期间可能会带来很大的内存开销。为解决这一问题,我们提出了 LeanGRPO,通过重构数据并行布局,并为基于轨迹对数概率的扩散强化学习引入两种免重计算的训练调度:(1) LeanGRPO-Retain 在 rollout 期间启用梯度跟踪,并直接复用由此产生的计算图和已保存的激活值用于更新阶段的反向传播,从而无需任何重计算;(2) LeanGRPO-Reweight 同样在 rollout 期间启用梯度,但会使用临时优势对每个选定的步立即进行反向传播并延迟梯度同步,然后在轨迹完成后用真实优势校正这些临时梯度。这些调度针对不同的模型规模和输入尺寸。在基于 FLUX.1-dev 和 Wan 的 FlowGRPO/DanceGRPO 实验中,LeanGRPO 在保持原始优化目标不变的情况下,实现了最高 1.83 倍的端到端加速。
cs.LG / 68 / 2609.04010
Unlocking Lossless Speedups in LLMs via Discrete Diffusion
通过离散扩散实现LLM的无损加速
Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu
cs.LG
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce $Ψ$-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to $3\times$ speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/
Chinese Translation
大型语言模型(LLM)的成功在很大程度上归功于下一词元预测(NTP),但其自回归(AR)结构需要缓慢、顺序的词元生成。为了克服这一瓶颈,我们引入了扩散增强型LLM,这是一类新模型,它定义了一个AR模型分布,同时使用扩散从该分布中并行抽取多个词元。我们将这些模型的参数解耦为两组:使用标准NTP目标训练的AR权重,以及用于同时生成多个词元的轻量级扩散权重。扩散权重通过一个简单的扩散蒸馏阶段学习,该阶段对现有LLM训练流程的额外开销可忽略不计。我们还引入了$Ψ$-Spec,一个采样器家族,它能够在固定上下文长度下实现无损加速和推理时扩展。与推测解码不同,我们的方法不需要单独的草稿模型。与扩散LLM(d-LLM)不同,它在不牺牲底层AR模型质量的情况下加速生成。由此产生的模型称为Uno,可以从头训练,也可以通过增强现有的开放权重AR LLM来构建。在每个评估的批大小下,Uno实现的吞吐量均优于领先的推测解码方法,并在基础AR模型上提供高达$3 imes$的加速,包括在设备支持的最大批大小下。值得注意的是,我们的8B Uno模型在智能体工具使用、编码和长上下文推理的所有评估基准上,均优于领先的开源d-LLM——26B DiffusionGemma,以及专有的Mercury 2。我们在以下网址发布代码和检查点:https://s-sahoo.github.io/uno/
cs.LG / 69 / 2609.04090
Conditioning Degenerate Diffusion Models
退化扩散模型的条件化
Uğur Aydın, Tamer Başar
cs.LG
diffusion
扩散模型相关
Abstract
Current conditioned generative models heavily rely on score functions for guidance during training. When the generative model is a diffusion process with a singular diffusion coefficient and the underlying (conditional) densities either do not exist or are not smooth, we use causal optimal transport to define \emph{approximate} loss functions that identify a minimum-entropy control for guidance under minimal assumptions. Our approach relies on causal optimal transport and its characterization through the predictable representation property of (conditioned) diffusion processes whose associated martingale problem is well posed, à la Üstünel.
Chinese Translation
当前的条件生成模型在训练过程中严重依赖得分函数进行引导。当生成模型是具有奇异扩散系数的扩散过程,且底层(条件)密度要么不存在、要么不光滑时,我们使用因果最优输运来定义近似损失函数,这些函数在极小的假设下识别出用于引导的最小熵控制。我们的方法依赖于因果最优输运及其通过(条件)扩散过程的可预测表示性质的刻画,这些过程的关联鞅问题是适定的,其方式遵循 Üstünel。
cs.AI / 70 / 2609.03414
StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios
StrixAE:一种面向真实场景中复杂失真耦合的智能音频增强智能体
Chenglin Wu, Junjie Wu, Jinhang Chen, Mingyang Chen, Zixu Lin, Jiabian Chen, Xinghao Ding, Xiaotong Tu
cs.SD · cs.AI
large language model
大语言模型相关
Abstract
Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.
Chinese Translation
真实场景中的音频增强涉及复杂的失真耦合,并且需要个性化增强。现有解决方案难以同时解决这两个问题。为了在此类场景中提升鲁棒性并实现自主操作,我们提出了StrixAE,一种基于多模态大语言模型(MLLM)的智能体。StrixAE利用MLLM作为控制器,协调多个音频增强和个性化模型。为了进一步增强系统鲁棒性、减少伪影并提升跨多样真实场景的泛化能力,StrixAE通过两阶段流程进行训练:首先,在AcoustBench上进行思维链(CoT)监督微调,以奠定基础推理和工具调用能力;其次,进行音频感知强化学习(APRL),这是一种专门为音频恢复流程设计的奖励机制,联合优化格式有效性、结构连贯性和感知质量。与通用的强化学习微调不同,APRL引入了结构化奖励,强制保证可执行的流程和逻辑段落顺序,使智能体能够生成可靠、可解释的增强方案,而不会产生幻觉工具。基于真实世界测试数据集,我们提出的方法优于大多数现有开源和专有解决方案,在多项感知指标上达到了最先进水平,并展现了强大的泛化鲁棒性。
cs.SE / 71 / 2609.03086
Large Language Models and Language Server Protocol: a match made in context
大型语言模型与语言服务器协议:上下文中的完美匹配
Alessandro Schena, Ilgiz Mustafin, Julia Kotovich
cs.SE
large language model
大语言模型相关
Abstract
This article introduces Eiffel-tools, a language server protocol (LSP) implementation for the Eiffel programming language that uses Large Language Models (LLMs) to aid the development of statically verified software. The tool provides various interactive and non-interactive commands to produce code and specifications. It uses language and project specific knowledge to precisely direct the LLM and verifies the output using a static verifier. It crafts rich programmatic prompts for the input and corrects or rejects the output. Furthermore, it handles the retries until the program passes verification. The tool's bug fixing capability is evaluated on 2 public datasets using 3 models. The tool can fix 76% to 95% of bugs by combining LLMs and a formal verifier depending on the model and prompts used. The results show the trade-off between the number of fixing attempts and the success rate.
Chinese Translation
本文介绍了Eiffel-tools,这是一个针对Eiffel编程语言的语言服务器协议(LSP)实现,它使用大型语言模型(LLMs)来辅助静态验证软件的发展。该工具提供了各种交互式和非交互式命令以生成代码和规范。它利用语言和项目特定的知识来精确引导LLM,并使用静态验证器验证输出。它为输入精心构造丰富的程序化提示,并纠正或拒绝输出。此外,它会处理重试,直到程序通过验证。该工具的错误修复能力在2个公共数据集上使用3个模型进行了评估。通过结合LLM和形式验证器,根据所用模型和提示,该工具可以修复76%到95%的错误。结果显示了修复尝试次数与成功率之间的权衡。
cs.SE / 72 / 2609.03156
Compound Prompt Constraints in LLM Code Generation: A Factorial Study of Format, Persona, and Urgency
LLM代码生成中的复合提示约束:格式、角色与紧迫性的因子研究
Shrenik Jadhav, Nickalsa LaPlaca, Caleb Stone, Ashok Raja, Omar Ochoa, Vidhyashree Nagaraju
cs.SE
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used in software engineering pipelines for code generation, where production prompts often combine multiple constraints. This paper presents a full-factorial empirical study of how output formatting, persona assignment, and urgency framing jointly affect LLM code-generation reliability. We evaluate all 27 combinations in a controlled 3x3x3 design and decompose each compound condition into an additive prediction and a residual interaction term that captures super-additive degradation. The study uses all 164 HumanEval+ problems across five OpenAI models from the GPT-4o family, GPT-4.1 family, and o3-mini, yielding 22,140 greedy-decoding evaluations. A format-aware extraction pipeline separates formatting failures from reasoning failures, and significance is assessed with McNemar's test, odds ratios, and 95% confidence intervals. Results show that compound constraints can produce architecture-dependent degradation not predictable from single-factor experiments. The GPT-4o family exhibits consistent super-additive effects, with pass@1 reductions 3-12 percentage points beyond additive predictions; the largest interaction is -12.2 pp on GPT-4o-mini for JSON + expert persona + moderate urgency. JSON combinations generally produce larger interactions than XML. In contrast, the GPT-4.1 family is largely resistant, while o3-mini shows a qualitatively different pattern in which structured output constraints can improve performance. These findings show that vulnerability is architecture-dependent rather than size-dependent, that individually neutral or beneficial constraints can combine to cause substantial degradation, and that compound-prompt testing should be standard in reliability assessment for LLM-assisted engineering pipelines.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于软件工程流程中的代码生成,其中生产环境中的提示词往往结合了多种约束。本文呈现了一项全因子实证研究,探讨输出格式、角色分配和紧迫性框架如何共同影响LLM代码生成的可靠性。我们在受控的3x3x3设计中评估了所有27种组合,并将每个复合条件分解为加性预测和残差交互项,后者捕捉超加性退化。该研究使用了来自GPT-4o家族、GPT-4.1家族和o3-mini的五个OpenAI模型上的全部164个HumanEval+问题,产生了22,140次贪心解码评估。一个格式感知的提取流程将格式失败与推理失败区分开来,并通过McNemar检验、优势比和95%置信区间评估显著性。结果表明,复合约束可能产生从单因素实验无法预测的、依赖架构的退化。GPT-4o家族表现出一致的超加性效应,pass@1的降低幅度比加性预测高出3-12个百分点;最大的交互效应出现在GPT-4o-mini上,为-12.2个百分点,对应JSON+专家角色+中等紧迫性的组合。JSON组合通常比XML产生更大的交互效应。相比之下,GPT-4.1家族基本不受影响,而o3-mini则表现出性质不同的模式,其中结构化输出约束可以提升性能。这些发现表明,脆弱性依赖于架构而非规模,单独中性或有益的约束可能组合后引发显著退化,并且复合提示测试应成为LLM辅助工程流程可靠性评估的标准做法。
cs.SE / 73 / 2609.03230
Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses
两个真相和一个谎言?现成大型语言模型用于需求质量评估的基准测试:性能、误报与漏检
Jannatul Shefa, Alejandro Salado, Paul Wach, Taylan G. Topcu
cs.SE
large language model
大语言模型相关
Abstract
Requirements engineering (RE) governs the quality of everything downstream in systems engineering (SE); defective requirements that survive review cycles propagate into design rework, schedule delays, and cost overruns. Because requirements are often written in natural language, recent advances in generative AI have raised expectations that large language models (LLMs) can absorb requirement quality assessment, a task otherwise slow and human expertise-intensive. Yet empirical evidence on whether LLMs can be trusted to do so remains scarce. This study presents the first benchmarking analysis of off-the-shelf LLM performance for requirement quality evaluation. Against an expert-derived ground truth built on INCOSE quality criteria, we evaluate ten models spanning two families (OpenAI and Anthropic) and five generations each, across one hundred independent runs, two requirement sets, and five sampling temperatures. Four contributions follow. First, we quantify a strongly asymmetric error profile: across all models and runs, the best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11%. Second, performance degrades significantly where SE judgment is required, as necessity and correctness issues are almost always missed. Third, generational progress is non-monotonic, so newer models cannot be assumed better. Fourth, this error behavior shifts only modestly and non-monotonically across sampling temperatures, indicating characteristic model deficiencies rather than inherent stochasticity. Off-the-shelf LLMs are therefore not yet trustworthy autonomous evaluators. Findings also warrant caution for Agentic AI developers: orchestrating these LLM modules in specialized architectures risks compounding these deficiencies rather than correcting them. Their defensible near-term role is human-in-the-loop decision support.
Chinese Translation
需求工程(RE)管控系统工程(SE)中一切下游工作的质量;在评审周期中存活的缺陷需求会传导至设计返工、进度延误与成本超支。由于需求通常以自然语言编写,生成式人工智能的最新进展使人们期望大型语言模型(LLM)能够承担需求质量评估这一本来缓慢且高度依赖人类专家的任务。然而,关于LLM能否被信任承担该任务的实证证据仍然稀缺。本研究首次对现成LLM在需求质量评估中的性能进行了基准分析。以基于INCOSE质量标准的专家派生真值(ground truth)为对照,我们评估了涵盖两个系列(OpenAI和Anthropic)且每个系列各有五代的十个模型,涉及一百次独立运行、两组需求和五个采样温度。本文有如下四点贡献。首先,我们量化了一种强烈不对称的错误谱:在所有模型和运行中,表现最好的Anthropic模型对专家已识别问题的检出比例中位数仅为47%,同时误报比例为11%。其次,在需要系统工程(SE)判断的地方,性能显著下降,因为必要性问题和正确性问题几乎总是被漏检。第三,代际进展是非单调的,因此不能假定更新的模型一定更好。第四,上述错误行为随采样温度的变化仅适中且非单调,这表明是模型的特征性缺陷,而非固有的随机性。因此,现成LLM尚不能成为值得信赖的自主评估器。这些发现也提醒智能体AI(Agentic AI)开发者需保持谨慎:在专门架构中编排这些LLM模块,可能会放大而非纠正这些缺陷。它们在近期站得住脚的角色是人在回路中的决策支持。
cs.SE / 74 / 2609.03267
Refusing the Impossible: A Taxonomy and Benchmark for Code Hallucination in Large Language Models
拒绝不可能之事:大型语言模型中代码幻觉的分类法与基准
Vishnu Asutosh Dasu, Ashish Kundu, Gang Tan
cs.SE
large language model
大语言模型相关
Abstract
Large language models (LLMs) often produce code that looks plausible but is not grounded in reality. The code may import packages that do not exist or claim to implement algorithms that violate proven theorems, while still compiling and running. We study \emph{code hallucination} as \emph{ungrounded generation} and separate it from ordinary \emph{code error} (bugs in otherwise grounded programs). We propose a taxonomy with three dimensions: \textbf{groundedness} (absolute violations of universal truths vs.\ relative fabrications of contingent or ecosystem-specific facts), \textbf{manifestation level} (syntactic, semantic, or factual), and \textbf{behavior} (from confident fabrication to degenerate output), organized into a severity ordering. We build an \textbf{adversarial} suite of deliberately unsatisfiable tasks where the correct response is to refuse and categorize the responses under our taxonomy. The suite contains \textbf{270 prompts} across six languages and 24 subcategories, paired with \textbf{91 matched solvable controls}, and responses are judged by a two-tier protocol validated against human labels (82\% agreement, $κ{=}0.73$). Across twelve open-weight code and reasoning models (4{,}332 judged responses), models produce ungrounded code on about 60\% of unsatisfiable prompts and refuse only 27\%, while wrongly refusing 0\% of the solvable controls.
Chinese Translation
大型语言模型(LLM)经常生成看似合理但并不基于现实的代码。这些代码可能导入不存在的软件包,或声称实现了违反已证明定理的算法,同时仍能编译和运行。我们将“代码幻觉”作为“无根据的生成”来研究,并将其与普通的“代码错误”(在其他方面有根据的程序中的缺陷)区分开来。我们提出了一个包含三个维度的分类法:根据性(对普遍真理的绝对违反与对偶然性或生态系统特定事实的相对捏造)、表现层面(句法、语义或事实)和行为(从自信的捏造到退化的输出),并按严重性顺序进行组织。我们构建了一个对抗性测试套件,其中包含故意无法满足的任务;在这些任务中,正确的回应是拒绝,并根据我们的分类法对回应进行分类。该套件包含跨六种语言和24个子类别的270个提示,并配有91个匹配的可解控制项;回应通过一个两级协议进行评判,该协议已针对人工标签进行验证(82%的一致性,$κ{=}0.73$)。在十二个开放权重代码和推理模型(共4,332个被评判的回应)中,模型在约60%的无法满足提示上生成了无根据的代码,仅拒绝27%,而对可解控制项的错误拒绝率为0%。
cs.SE / 75 / 2609.03721
Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study
LLM能否从源代码提交中提取架构设计决策?——一项初步探索性研究
Amey Karan, Rudra Dhar, Mohamed Soliman, Karthik Vaidhyanathan
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
Context: Architectural Design Decisions (ADDs) capture the rationale behind the structure and evolution of software systems but are rarely documented explicitly, and are often hidden inside source code commits. Recovering them is important for Architectural Knowledge Management (AKM). Problem: Extracting ADDs from commits is challenging due to their implicit and unstructured nature. Large Language Models (LLMs) have shown strong capabilities in understanding code and text, yet their effectiveness for this task remains underexplored. Study: We present a preliminary study using four LLMs (Gemini 3 Pro, DeepSeek R1, Kimi K2, Qwen3) with zeroshot and fewshot prompting on 30 developer-written ADDs from open-source projects. We score outputs with ROUGE-L, BLEU, METEOR, and BERTScore, and one author manually reviews the Gemini outputs. Results: All models reach a BERT-F1 above 0.81, and fewshot prompting improves alignment (Gemini BERT-F1: 0.828 to 0.847). However, the generated ADDs are often too long, implementation-focused, and miss the rationale behind the decision. This highlights opportunities for architecture-aware LLM systems and automated AKM.
Chinese Translation
背景:架构设计决策(ADD)捕获了软件系统结构和演进背后的 rationale,但很少被显式记录,且通常隐藏在源代码提交中。恢复它们对于架构知识管理(AKM)非常重要。问题:由于提交中 ADD 的内隐性和非结构化特性,从中提取 ADD 具有挑战性。大型语言模型(LLM)在理解代码和文本方面展现出强大能力,但其在此任务上的有效性仍未被充分探索。研究:我们开展了一项初步研究,使用四种 LLM(Gemini 3 Pro、DeepSeek R1、Kimi K2、Qwen3),采用零样本和少样本提示方法,针对来自开源项目的 30 条开发者编写的 ADD 进行实验。我们使用 ROUGE-L、BLEU、METEOR 和 BERTScore 对输出进行评分,并由一位作者人工复核 Gemini 的输出。结果:所有模型的 BERT-F1 均高于 0.81,且少样本提示改善了对齐程度(Gemini BERT-F1 从 0.828 提升至 0.847)。然而,生成的 ADD 往往过长、偏向实现细节,并且遗漏了决策背后的 rationale。这凸显了架构感知的 LLM 系统以及自动化 AKM 的发展机遇。
cs.SE / 76 / 2609.04055
LabelMate: An LLM-Driven Framework for Refined Issue Report Labeling
LabelMate:一种由LLM驱动的精细化问题报告标注框架
Liam Johnston, Shayan Noei, Maram Assi, Ying Zou
cs.SE
large language model
大语言模型相关
Abstract
Software users often submit issue reports to a product's issue tracking system to report defects, suggest enhancements, or raise other product-related concerns. Labeling these issue reports supports effective planning and improves community engagement. However, many issue reports remain unlabeled due to the substantial manual effort required to design an appropriate label taxonomy, then assign suitable labels from this taxonomy to new issue reports. Existing automated labeling approaches attempt to mitigate these challenges. However, they suffer from key limitations, such as extensive manual intervention, the assignment of generic labels, and a dependence on existing labeled datasets. To address these limitations, we propose LabelMate, a novel Large Language Model (LLM)-driven framework that (1) derives a comprehensive, project-specific label set from historical issue reports and (2) automatically assigns relevant labels to new issue reports without requiring any pre-labeled training data. We evaluate LabelMate on 16,500 issue reports from 30 popular and diverse GitHub repositories. Based on this dataset, our approach generates a coherent list of 275 labels and achieves an average labeling accuracy of 89.84%, a statistically significant improvement over existing generic label assigning approaches. These results demonstrate that LabelMate offers an efficient, domain-adaptive solution to streamline the issue labeling process.
Chinese Translation
软件用户通常会向产品的问题跟踪系统提交问题报告,以报告缺陷、提出增强建议或表达其他与产品相关的关切。为这些问题报告添加标注有助于进行有效的规划,并提升社区参与度。然而,由于设计一套合适的标注分类体系,然后从该分类体系中为新的问题报告分配合适的标签,需要大量的手动工作,许多问题报告仍然未被标注。现有的自动化标注方法试图缓解这些挑战。然而,它们存在一些关键局限性,例如大量的人工干预、分配通用标签以及对现有已标注数据集的依赖。为解决这些局限性,我们提出了LabelMate,一种新颖的由大语言模型(LLM)驱动的框架,它(1)从历史问题报告中推导出一套全面的、针对特定项目的标签集合,并且(2)在无需任何预标注训练数据的情况下,自动为新的问题报告分配相关标签。我们在来自30个流行且多样化的GitHub仓库的16,500份问题报告上评估了LabelMate。基于该数据集,我们的方法生成了一组连贯的275个标签,并实现了89.84%的平均标注准确率,这一结果在统计上显著优于现有的通用标签分配方法。这些结果表明,LabelMate提供了一种高效、领域自适应的解决方案,可简化问题标注流程。
cs.SE / 77 / 2609.04061
When Models Edit Too Much: On the Fidelity of Minimal Code Edits
当模型编辑过多:论最小代码编辑的保真性
Tongyao Zhu, Wei Hern Lim, Min-Yen Kan
cs.SE · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.
Chinese Translation
大语言模型(LLMs)越来越多地被用于编辑现有代码,但仅凭正确性是不够的:有用的修复还应是最小化的、可审查的,并且忠于原始实现。我们研究过度编辑,即模型在修复 bug 所需范围之外重写代码的倾向。我们从 400 个 BigCodeBench 问题中构建了一个评估框架,方法是向参考解决方案注入受控的 AST 级损坏,使每个修复任务都有一个已知的最小补丁。在前沿 LLM 中,过度编辑非常普遍,即使是像 GPT-5.5 这样强大的模型也不例外:高 Pass@1 可能与不必要的大改动以及增加的认知复杂度并存。保留指令能够大幅减少这种行为,将平均多余 Levenshtein 距离从 0.195 降至 0.131,使增加的认知复杂度降低 26.6%,并将 Pass@1 提高 2.3 个百分点。然而,这些收益并非仅仅来自更大的推理预算或更大的模型。接下来,我们探究最小编辑是否能在后训练阶段直接习得。我们观察到,监督微调对已见的损坏模式过拟合,而强化学习则在域外编辑保真度和性能保持之间取得了最佳权衡。这些结果将编辑保真度定位为代码修复质量的一个独立维度,并表明它可以被度量与学习。
cs.AI / 78 / 2609.03620
ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection
ToolDF:面向混合真实性音频深度伪造检测的工具集成推理
Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim
eess.AS · cs.AI · cs.SD
large language model
大语言模型相关
Abstract
Audio deepfake detection is commonly formulated as clip-level binary classification of single-domain audio. However, real-world manipulated audio can exhibit mixed authenticity, where genuine and manipulated cues coexist across temporal transitions, overlapping sources, or both. This setting requires not only detecting manipulated audio but also localizing the components that provide evidence for the decision. We propose ToolDF, a tool-integrated reasoning framework for mixed-authenticity audio deepfake detection. ToolDF employs an audio large language model as an orchestrator trained with supervised tool-use trajectories. It adaptively analyzes the audio scene, selectively performs source separation, routes components to domain-specific experts, and aggregates their evidence into an interpretable verdict. We further introduce a mixed-authenticity ADD benchmark covering temporal transitions, acoustic overlaps, and hybrid mixtures. Experimental results show that ToolDF achieves the best overall performance on composite-type detection, achieving macro-F1 gains of 3.72 and 14.39 points over the strongest monolithic baseline and a fixed pipeline, respectively, while providing interpretable evidence localized to temporal regions and acoustic sources. Our source code and dataset are publicly available online.
Chinese Translation
音频深度伪造检测通常被表述为对单域音频的片段级二分类问题。然而,现实世界中被操纵的音频可能表现出混合真实性,其中真实和被操纵的线索在时间转换、重叠声源或两者中并存。这种场景要求不仅检测被操纵的音频,还要定位为决策提供证据的组件。我们提出了 ToolDF,一个用于混合真实性音频深度伪造检测的工具集成推理框架。ToolDF 将音频大语言模型作为编排器,该模型使用监督式工具使用轨迹进行训练。它自适应地分析音频场景,有选择地执行源分离,将组件路由给领域特定专家,并将它们的证据聚合成一个可解释的判断。我们进一步引入了一个混合真实性 ADD 基准,涵盖时间转换、声学重叠和混合混合物。实验结果表明,ToolDF 在复合类型检测上取得了最佳整体性能,相较于最强的单体基线和固定流水线,分别获得了 3.72 和 14.39 个百分点的宏F1提升,同时提供了定位到时间区域和声学源的可解释证据。我们的源代码和数据集可在线公开获取。
cs.LG / 79 / 2609.03382
SurgeGen: A Hybrid Generative Diffusion Framework for Storm Surge Scenario Synthesis
SurgeGen:用于风暴潮情景合成的混合生成式扩散框架
Shunan Zheng, John J. Hasenbein
math.DS · cs.LG
diffusion
扩散模型相关
Abstract
Predicting storm surge induced by landfalling tropical cyclones is crucial for flood mitigation and coastal risk management. Traditionally, physics-based numerical models simulate storm surge by solving the Navier--Stokes equations using numerical methods, but these simulations are computationally expensive. Generative models are promising for storm surge emulation because they can generate diverse realizations rather than producing a single deterministic prediction. However, their use for storm surge emulation remains largely unexplored. In this paper, we leverage diffusion models for storm surge surrogate modeling, combining a baseline prediction stage with conditional generation to provide a more interpretable modeling framework. We develop SurgeGen, a two-stage generative framework for generating storm surge scenarios conditioned on hypothetical storms with parameters defined in a continuous space. First, a baseline model produces a coarse estimate of the storm surge height. This estimate then conditions a diffusion model, which generates refined storm surge scenarios that better capture spatial patterns and variability. We demonstrate that our approach can generate realistic and diverse storm surge scenarios under conditions both within and outside the training distribution.
Chinese Translation
预测登陆热带气旋引起的风暴潮对于防洪和沿海风险管理至关重要。传统上,基于物理的数值模型通过使用数值方法求解纳维-斯托克斯方程来模拟风暴潮,但这些模拟的计算成本很高。生成模型在风暴潮仿真方面具有前景,因为它们能够生成多样化的实现,而非仅产生单一的确定性预测。然而,它们用于风暴潮仿真仍 largely 未被探索。在本文中,我们利用扩散模型进行风暴潮替代建模,将基线预测阶段与条件生成相结合,以提供一个更具可解释性的建模框架。我们开发了 SurgeGen,这是一个两阶段生成框架,用于生成以连续空间中定义的参数为条件的假设风暴情景下的风暴潮情景。首先,基线模型生成风暴潮高度的粗略估计。然后,该估计为扩散模型提供条件,该模型生成能够更好捕捉空间模式和变异性的精细化风暴潮情景。我们证明了我们的方法能够在训练分布内外的条件下生成真实且多样化的风暴潮情景。
人工智能 (cs.AI)
81
cs.AI / 1 / 2609.03209
MasterControl Seventeen Every Time
MasterControl AI Lab
cs.AI
Abstract
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy executed the approved program. None of 330 runtime-planning episodes matched the full answer-and-evidence contract across all test datasets; the policy-executed analyzer matched 110 of 110. This is a configuration-specific result, not evidence that runtime agents cannot succeed under other designs.
cs.AI / 2 / 2609.03236
Speculative Macro Commit for Faster Tool-Using Agents
Zeyu Liu, Souvik Kundu, Peter A. Beerel
cs.AI · cs.MA
Abstract
Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro Commit} (SMC), a runtime mechanism for a two-tier agent system: a large authoritative actor model produces the official trajectory, while a faster speculative drafter model continuously predicts and executes future action chains on an isolated environment snapshot. SMC mines recurring multi-action skeletons from training traces and stores them in a macro library used to match against action chains predicted by the drafter at runtime. When the actor's next tool call matches the first drafted action, SMC commits the remaining pre-executed draft steps, together with their observations, to the official trajectory. Using Qwen3.5-27B INT4 as the authoritative actor model and Qwen3.5-4B as the speculative drafter model, SMC matches the sequential agent's overall accuracy while reducing latency by 10.23\% over the Speculative Actions (SA) baseline and 18.59\% over sequential execution on the $τ^2$-Bench Telecom subset. On AppWorld, SMC reduces wall time by 7.7\% over SA baseline and 44.9\% over sequential execution, with a small reduction in task completion. Overall, SMC provides a practical way to reuse multi-step speculative execution and reduce agent latency beyond single-step speculative actions. Our code is publicly available \href{https://github.com/zeyuliu1037/speculative-macro-commit}{\textcolor{magenta}{here}}.
cs.AI / 3 / 2609.03340
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
Evan Chen, Shiqiang Wang, Christopher G. Brinton
cs.AI
Abstract
Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r_3$, another agent may commit $r_4$, and an executor may receive $r_4$ without replacing the plan derived from $r_3$. We call this \emph{stale-plan execution}: state freshness does not establish that the plan authorizing an action remains valid. We introduce PlanFence, a dependency-scoped action-validation protocol. Plans cite the exact public records they used, and an executor validates only the records that can affect the pending external action, replanning once or blocking when validation is incomplete. In 30 controlled live workflows with a post-plan revision, a freshness-only executor acts on the obsolete plan in every task, whereas PlanFence completes all tasks without an invalid action. Controlled replay reveals two conditional boundaries: proactive synchronization yields lower coordination stall at low churn, while PlanFence avoids repeated update-path coordination as churn grows and avoids validating unrelated state as the shared keyspace grows. These are controlled safety and systems-cost results, not general task-accuracy gains.
cs.AI / 4 / 2609.03416
Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
Weijie Liu, Running Zhao, Wenhao Yuan, Jinfeng Xu, Zhanfeng Xu, Xiaoxi Zhang, Edith Cheuk-Han Ngai
cs.AI · cs.LG
Abstract
LLM-empowered paper-code discrepancy detection has received growing concern since the scaling of research submissions exceeds the manual review capability. However, the limited context capacity and one-sided discrepancy detection of existing single-agent LLM paradigms lead to an inferior recall performance in detecting discrepancies. In this paper, we propose Dude, the first Dual-Detection Multi-Agent System for paper-code discrepancy detection. We discover that the granularity asymmetry of the paper-language and code-language introduces over-interpretation and over-reporting challenges in a multi-agent system design for discrepancy detection, resulting in increasing false positives. To address this, we propose a granularity-aligned negotiation and a two-stage salience-filtering mechanism in Dude, which effectively prevents agents from falsely reporting discrepancies. Experimental results in real-world paper-code discrepancy datasets showcase Dude's significant recall and precision improvement by up to 22.8%, increasing F1 score by up to 18.7% compared to baseline methods.
cs.AI / 5 / 2609.03423
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
Puneet Mathur, Dinesh Manocha
cs.AI
Abstract
Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating implicit instruction-following in real-time spoken interaction. (DSB-IFEval) comprises 1,038 test cases spanning eight diverse assistant roles and evaluates five conditioning protocols for instruction-following: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona--rule conditioning, and instruction conflict. We measure real-time floor management using a deterministic Instruction Adherence Score (IAS) and persona-consistent content using LLM-judged Persona Adherence Score (PAS). Across six real-time speech systems, we find architecture-dependent trade-offs. Full duplex models like F-Actor and PersonaPlex are more sensitive to whether conversational behavior is stated explicitly or must be inferred from a persona, with adherence dropping by 9.7% and 4.5%, respectively, under persona-only conditioning. In contrast, GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly adhere to persona-consistent content, but their floor behavior does not adapt across explicit and persona-only instructions and remains constrained on several proactive actions. We further find that even if systems reliably follow conflicting directives to their prescribed persona, they still struggle to override them under safety conflict. These results show that inferring the behavior implied by a role, executing it at the appropriate conversational moment, and resolving competing instructions remain distinct challenges for full-duplex voice agents.
cs.AI / 6 / 2609.03438
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng, Zheng Wu, Yansi Li, Chuanbiao Song, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
cs.AI
Abstract
Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. In this work, we introduce CONFLICTGUI, a benchmark covering instruction-internal conflicts and instruction-GUI context conflicts to study conflict-aware termination. Our evaluation reveals severe execution-biased overcompliance: agents that perform well on feasible tasks often continue to execute blindly under conflicting instructions. To mitigate this behavior, we propose CONFLICTGUARD, an inference-time framework that aligns an agent's feasibility awareness with its action generation. CONFLICTGUARD contains two coupled components: a feasibility verification protocol that guides the agent to assess instruction logic and GUI-side evidence before acting, and a conditional action modulation mechanism that steers agents from over-compliant execution into termination-oriented behavior. Experiments across five widely-used agents demonstrate that CONFLICTGUARD improves average conflict task success rate significantly, while preserving normal GUI-task performance. These results validate that a lightweight inference-time intervention can substantially boost GUI Agent's competence to identify inappropriate execution scenarios and refrain from unnecessary actions.
cs.AI / 7 / 2609.03460
Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
Qing Zhang, Yifei Huang, Juyoung Lee, Thad Starner, Jun Rekimoto
cs.AI
Abstract
As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary ``Made with AI'' labels respond with authorship disclosure, but they do not show what supports a claim. We propose Provenance Density, an evidence-visualization interface that shows the density of verified claims in a text. In a user study with 81 participants, an idealized Provenance Density interface produced a large discernment gap between truth and fabrication ($+4.15$ points, $d=1.82$), whereas participants given no signal showed no detectable discrimination. A technical audit with 200 samples shows that retrieval density alone is insufficient; unexpectedly, the Consistency Veto carries most of the discriminative signal on dynamic queries. As AI-generated content becomes indistinguishable from human writing, effective transparency must move from authorship disclosure toward evidence visualization.
cs.AI / 8 / 2609.03478
AutoGraphForge: Towards Automated Graph Theory Discovery
Ján Pastorek
cs.AI · cs.LO · math.CO
Abstract
We report on our ongoing project to develop a computational pipeline, AutoGraphForge, for an automated graph-theoretic conjecturing-refuting-formalizing-proving system. Conjecture generation is counterexample-guided and runs in rounds: a Graffiti3 generator proposes conjectures over a small, evolving snapshot table $T$ (initially a few hundred graphs with their computed invariants) that grows only by counterexamples to its own conjectures. A novelty filter of $559$ classical and folklore relations, closed under transitive composition and linear identity substitution, decides via a linear program whether a candidate is already implied by known results. Surviving candidates are tested against a dataset of about $348,000$ graphs, unioning the complete House of Graphs invariant export, the exhaustive census of all connected graphs on at most nine vertices, several extremal families (strongly regular, minimal Ramsey, Cayley, cages, barbells, lollipops, spiders), and random models. Counterexample-search algorithms then attack the remainder. Run for several rounds on an HPC cluster, the loop yields $6,522$ conjectures that survived the refutation dataset, the novelty filter and every active-search run -- among them nontrivial relations between the annihilation number and the edge-cover number for bipartite and regular graphs, which we prove by hand. A subsequent formalization and proving stage deterministically translates each surviving conjecture into a Lean 4 statement skeleton; every candidate proof is kernel-verified against a pinned mathlib4 and our custom invariant preamble. This stage integrates two neural provers -- DeepSeek-Prover-V2-671B (served with vLLM) and the Lean-specialised OProver-32B -- behind the independent kernel check. It is implemented end-to-end and passes initial sanity checks, with the full pipeline currently running on the cluster.
cs.AI / 9 / 2609.03493
Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
Xingming Long, Yu Liu, Zhiwei Yang, Hanqi Feng, Shaojie Zhang, Barnabas Poczos, Chao Jiang, Zhenbo Luo, Lei Jiang, Pei Fu
cs.AI
Abstract
Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropriate tools are called, models often fail to extract the necessary information from the resulting observations. To address these limitations, we introduce the NTEP (Necessary Tool-Evidence Path), a novel annotation scheme that explicitly specifies the essential external evidence and corresponding tool calls for each query. Building upon this, we propose NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution. Specifically, our approach rewards the agent for aligning its pre-call intent with a necessary evidence-seeking goal, and for ensuring the information summarized from the post-call observation aligns with the necessary evidence. Furthermore, we introduce a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals. Extensive evaluations on seven image-grounded benchmarks demonstrate that our 8B-parameter instantiation, NTEP-8B, significantly improves both search-oriented accuracy and tool-use efficiency within a unified three-tool framework. These results highlight the critical value of fine-grained tool-evidence path supervision for training robust agentic VLMs.
cs.AI / 10 / 2609.03494
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
Qiankun Ma, Yanjiang Zhou, Zinan Xiong, Haofei Wang, Zhen Song, Yang Xiang, Ziyao Zhang, Hairong Zheng
cs.AI
Abstract
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed throughout decoding. However, reasoning workloads exhibit substantial demand variation: different requests require different KV capacities, and the attention demand of an individual request evolves during generation. We introduce \textbf{GrowPage}, an on-demand KV budgeting framework that treats KV capacity as a runtime resource. GrowPage maintains lightweight dual-timescale query summaries to capture recent and long-term attention behaviors, and uses their relative attention working sets to estimate demand evolution. At each capacity boundary, GrowPage either compresses KV states within the current allocation or acquires an additional physical page when broader demand emerges. By integrating with PagedAttention's page-level memory abstraction, GrowPage preserves continuous batching and prefix caching. Experiments on reasoning benchmarks across multiple models show that GrowPage achieves a superior performance--throughput trade-off over existing approaches.
cs.AI / 11 / 2609.03503
PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing
Yangshuo Qi, Chenwei Wang, Zihan Shen, Songlin Sun
cs.AI
Abstract
With the rapid development of the Internet of Things, computation intensive directed acyclic graph (DAG) tasks have become increasingly common in cloud-edge-end collaborative environments. However, cloud, edge, and end nodes are highly heterogeneous in computing capacity, network bandwidth, and energy consumption, which makes the efficient scheduling of tasks with complex dependencies an NP-hard problem. Traditional heuristic algorithms and conventional reinforcement-learning methods often fail to capture the spatio-temporal dynamics of system resources. This paper proposes PPO-STGNN, a DAG task-scheduling algorithm that integrates proximal policy optimization (PPO) with spatio-temporal graph neural networks (STGNNs). The method uses an STGNN to extract features from both the DAG task topology and the physical cloud-edge-end resource graph, and then optimizes the scheduling policy through PPO to minimize makespan and schedule length ratio (SLR) while improving CPU and memory load balancing. To accelerate convergence, a multi-teacher behavior-cloning mechanism is introduced for pretraining. Experimental results show that PPO-STGNN significantly improves load balancing while maintaining a low completion time, making it suitable for dynamic and heterogeneous cloud-edge- end DAG scheduling scenarios.
cs.AI / 12 / 2609.03515
What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
Bo Zeng, Yu Zhao, Yefeng Liu, Zhihong Lu, Xuanfan Ni, Xintong Wang
cs.AI
Abstract
Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponential-moving-average (EMA) aggregation makes approximately order-preserving scorer modifications largely indistinguishable at the eviction-set level. Value-norm and entropy variants remain highly correlated with attention and produce nearly unchanged retention sets, whereas KeyDiff, key norm, recency, and a learned scorer alter the ranking and degrade substantially. We associate this stability with the evaluated aggregation, which couples layer weighting and temporal retention. Building on this observation, we introduce InertiaKV, an EMA-based decoding-time eviction method, and InertiaKV-Lazy, its periodic-refresh variant, which yields 1.34-1.46x decode throughput relative to full refresh InertiaKV. We also study Score-Free decoding as a separate empirical operating point: it scores the full context once at the first decode step, freezes that ranking, and incurs an average quality change of +0.03 while removing all subsequent scoring. Across six open-weight backbones and the LongBench, LongBench-v2, and RULER benchmarks, the results identify temporal aggregation and ranking preservation as distinct, consequential design factors; they do not imply that scoring quality is irrelevant in general.
cs.AI / 13 / 2609.03526
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua Luo, Qinggang Zhang, Jinsong Su
cs.AI
Abstract
Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.
cs.AI / 14 / 2609.03535
Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
Yinan Liu, Jiankang Hong, Zhen Gao, Ye Lu
cs.AI · eess.IV
Abstract
Lesion segmentation in medical images plays a critical role in clinical diagnosis and treatment planning. Despite significant advances, lesion segmentation remains challenging due to two major factors: (1) complex background interference; (2) diverse lesion morphology. Existing encoder-decoder based methods mainly focus on enhancing feature extraction or redesigning decoding strategies. However, they lack early prior guidance and feature reconfiguration during the encoding stage, limiting their effectiveness in handling these challenges. To address these limitations, we propose FreNet, a feature reconfiguration framework with visual priors, which performs pixel-level reconfiguration before encoding and feature-level reconfiguration during encoding for precise medical lesion segmentation. To suppress background responses, we propose an Implicit Prior Neural Network (IPNN), which models a continuous spatial field and leverages visual prior from SAM to reconfigure input image before encoding stage. To better handle diverse lesion morphology, we design a Dual-domain Feature Reconfiguration (DFR) module to progressively reconfigure backbone features during encoding stage. Within DFR, the Frequency Decoupling Module (FDM) decouples backbone features in frequency domain to enhance foreground-background discriminability, while the Spatial Localization Module (SLM) spatially relocates and improving spatial stability after frequency decoupling. Extensive experiments on 9 medical image segmentation benchmarks across three imaging modalities demonstrate that FreNet significantly outperforms state-of-the-art (SOTA) methods. On the challenging ETIS dataset, our method achieves Dice improvements of 5.0% over SOTA method and 7.2% over SAM.
cs.AI / 15 / 2609.03553
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
Linh Le, Melanie Bui, My Chiffon Nguyen, Zachary Schlosser, David Williams-King
cs.AI · cs.CY
Abstract
Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns "does multi-agent simulation help?" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.
cs.AI / 16 / 2609.03588
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
Yaxing Lyu, Shengjie Zhou, Binbin Toh, Pengyu Zhu, Lijun Li
cs.AI
Abstract
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
cs.AI / 17 / 2609.03621
A computable representation of the physical laboratory enables verifiable workflows
Xiaobo Li, Luyao Ge, Xiaohui Li, Lulu Guo, Ming Mao, Jiwang Zheng, Wenting Guan, Xin Yang, Yi Luo, Jun Jiang, Linjiang Chen
cs.AI
Abstract
Making science computable requires representations of both scientific knowledge and the physical world in which scientific claims are tested. A computable representation of the physical laboratory is established through typed research objects, capability-bound operations and a compositional workflow algebra. It provides the physical-world counterpart to machine-readable knowledge, expressing workflows as programs over evolving laboratory states with explicit dependencies, decisions, iteration and concurrency. The representation was implemented in a modular agentic robotic laboratory by binding formal operations to executable Function Skills. For diverse scientific intents, capability-relative workflows were generated, while stateful simulation propagated object transformations and verified operation preconditions and laboratory constraints before dispatch. The proposed representation and its engineering framework jointly establish a general computational interface between agent reasoning and capability-bound physical transformations, providing a foundation for end-to-end autonomous scientific discovery.
cs.AI / 18 / 2609.03702
Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study
Kenneth Paulsen, Florian Tambon, Mike Papadakis, Shin Yoo
cs.AI
Abstract
General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C, C++, and Java. Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.
cs.AI / 19 / 2609.03707
Counterfactual Routing Using Integer Programming with Constraint Generation
Daniël Vos, Sterre Lutz
cs.AI · cs.DS
Abstract
We present our submission to the IJCAI 2025 'Counterfactual Routing Competition' (CRC 25). The goal of the competition is to find counterfactual explanations for the shortest path problem. This requires deciding what the minimal changes to a road network would make a route chosen by the user the optimal route. This enables explanations such as "Your suggested route would indeed have been optimal, if road X were not a bicycle path." Our solution models the problem as an integer program, iteratively incorporating constraints until an exact solution is found. In the final evaluation on held-out test instances, our method ranked fourth in solution quality and obtained its solution fastest on every instance, with an average runtime of 9.0 seconds compared to 118.8 seconds for the next-fastest submission.
cs.AI / 20 / 2609.03716
Artificial Intelligence for Energy Optimization in Data Centers
Mohammed Basharath Ullah, Summaiya Unnisa Begum, Mohammed Nadeem Ullah
cs.AI · cs.LG
Abstract
Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded by it. The literature treats these as two unrelated problems: control studies model workload as an exogenous arrival process, while sustainability studies model infrastructure as a fixed multiplier. We screen roughly 194 papers retrieved through a documented protocol, code 63 of them, and report what the coding shows. Of 28 primary control-oriented studies, 18 are validated in simulation alone and 5 reach physical hardware or a production facility; none account for water withdrawal, and none account for embodied carbon. Reported savings intervals across four technique families overlap almost completely, which means the field cannot presently rank its own methods. Ten recurring gaps are scored for consequence and tractability, and we set out CLEAR-DC, a framework coupling a control-policy branch to a workload-demand branch through an explicit elasticity term, reads out net rather than direct benefit, and emits a schema-conformant record covering energy, carbon, water, embodied share and validation venue. The framework is an architectural and methodological proposal, not a trained system; the contribution we defend empirically is the corpus analysis and the reporting schema derived from it. Coding sheet, derived statistics and all result artifacts: https://github.com/Kimalice/AI-for-Energy-Optimization-in-Data-Centers-Closing-the-Optimizer-Load-Loop
cs.AI / 21 / 2609.03774
Rethinking World Models for Safety-Critical Embodied Systems
Kailang Ma, Heye Huang, Inhi Kim, Kitae Jang
cs.AI · cs.RO
Abstract
World models have progressed from compact latent dynamics to generative, controllable, and interactive simulators of embodied environments. However, high predictive likelihood and visual fidelity do not necessarily ensure that a model preserves the evidence required for safe decision-making. This perspective identifies three structural mismatches in current world modeling: likelihood versus risk, prediction versus intervention, and finite-horizon prediction versus accumulated consequences. We propose the Risk-Informed World Model (RIWM) as a decision-centric research direction for safety-critical embodied systems. RIWM organizes world modeling around consequences, intervention, epistemic uncertainty, and recoverability, and integrates four interdependent capabilities: decision-relevant representation, counterfactual reasoning, safety-critical episodic memory, and runtime safety assurance. It distinguishes physical, social, and operational consequences while using epistemic uncertainty to qualify the evidence supporting action. We further discuss open challenges in identifying consequential futures, validating counterfactual reasoning, maintaining revisable safety memories, translating learned consequences into executable constraints, and determining when evidence is sufficient to act. This perspective argues that future world models should move beyond predicting likely futures toward identifying which futures matter, revising judgments through experience, and recognizing when to act, revise, sense, defer, or abstain.
cs.AI / 22 / 2609.03787
DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang, Gang Liu
cs.AI
Abstract
AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-Twin, a graph-native digital twin that records a committed agentic decision as a typed trajectory and re-executes its decision mechanism under declared conditions. The graph links the state observed by the agent, the path it followed, and the authority behind the resulting action. The twin synchronizes this information, replays the mechanism in isolation, and compares it under controlled changes. We instantiate the framework in enterprise decision processes using three public process logs and controlled replay suites. The experiments identify a specific failure: graph structure localizes represented changes but cannot determine the consequence of an unobserved tool state. In a three-condition controlled experiment with 300 injected instances, unresolved-divergence recall increased from 0 to 0.667 when replay-contract state was added and to 1.0 when verification results were also available; the held-out set contained no critical-class instance. Across 500--5,000 BPI 2020 cases, median end-to-end time increased from 0.794 to 8.889 seconds on the reported platform. These results separate the roles of graph structure, replay context, and verification evidence in reviewing a decision mechanism.
cs.AI / 23 / 2609.03797
Transfiver: Human-AI Co-Inference through a Shared Editable State
Minji Park, Seunghyun Yoon, Hyuk Lim
cs.AI · cs.CL · cs.HC
Abstract
Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user. We introduce the TRANSparent Framework for Interactive, Verifiable, Editable Representation (Transfiver), an architecture for human-AI co-inference through a shared editable state. Its central idea is that interaction-specific information is maintained in a single persistent state $(S_t)$ that both the model and the human update. Transfiver distinguishes two modes of state evolution. In an implicit stream update, the model interprets ongoing interaction and decides whether new information revises an existing state item or creates a new one. In an explicit directed edit, a human inspects and modifies an addressed item. Both act on the same underlying state, so a human correction changes the state that subsequent computation reads, rather than adding another instruction or separate record. The architecture separates shared parameters $(θ)$, learned before ordinary use, from the persistent state $(S_t)$, which evolves during deployment without parameter retraining. Extending Transfiver to rich natural-language, relational, and large-scale shared states remains open.
cs.AI / 24 / 2609.03800
Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI
Phoenix Perry, George Simms, Elizabeth Wilson, Yasmine Boudiaf, Nick Bryan-Kinns, Tega Brain, R. Luke DuBois, Alix Rule, Rachel Meade Smith, Kelani Nichole, Atharva Pravin Pawar, Rebecca Fiebrink
cs.AI · cs.HC
Abstract
Federated learning is increasingly presented as a privacy-preserving advance: personal data remain on the device, and only model updates are shared. It borrows the vocabulary of the federated social web, yet inverts its logic, distributing computation while the resulting model stays with whoever convened the training. We argue that federation is not in itself a remedy for extractive AI, because outcomes depend on who governs the data and the model and who has agency over the practices that shape them. We describe three layers at which a creative community can hold its work: storage, circulation, and learning. Examining artist-governed trusts, cooperatives, and consent infrastructures, we show that creator governance is established at storage and circulation but stops at learning: contributors can consent to training, yet have little say over the resulting model or its federation. We map the research space this opens, pairing technical open problems with the human questions from which they unfold. We propose four design principles for a creative data commons that governs models and their federation, not only datasets: govern the model, not only the corpus; make the terms legible at the moment of contribution; design for refusal as a first-class state; and decide stewardship in the open and account for it.
cs.AI / 25 / 2609.03806
SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
Marco Cipriano, Leonardo Zini, Alexandra Schild, Valentin Teutschbein, Afsana Mimi, Marcella Cornia, Lorenzo Baraldi, Gerard de Melo
cs.AI · cs.CV
Abstract
Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment. We introduce \textbf{\ours}, a human-aligned evaluation framework for text-to-SVG generation. Through controlled caption and image perturbations, we first show that CLIP-based scores barely react to the errors SVG generators actually make, such as wrong colors, counts, and spatial relations, and that off-the-shelf Vision-Language Model (VLM) judges, while more sensitive, respond unevenly across error types and SVG styles. We then introduce a human-annotated dataset for \textit{Semantic Alignment}, measuring how faithfully a generated SVG reflects its caption. Building on it, we develop two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning, for more expressive and interpretable assessment. Using both, we benchmark major open-source, commercial, and optimization-based SVG generators on an independent caption set.
cs.AI / 26 / 2609.03818
CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception
Weize Li, Yang Li, Quan Yuan, Xiaoyuan Fu, Guiyang Luo, Jinglin Li
cs.AI
Abstract
Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate this problem by mapping heterogeneous features into a shared protocol space; however, independently trained modality-specific converters often generate modality-specific pseudo-protocol distributions, leading to semantic inconsistency and error accumulation, which is particularly pronounced in scenarios with large modality discrepancies. To address this issue, we propose CauseCollab, a causal unified and modality-agnostic network. CauseCollab formulates representation learning in the protocol space from a causal perspective, explicitly disentangling semantic factors from modality-specific statistical confounders via causal metric learning. Meanwhile, CauseCollab adopts context-guided Unified Converter for heterogeneous modalities to ensure cross-modal semantic consistency. In addition, integrating new modalities only requires training adapters with minimal parameters. Extensive experiments on the OPV2V and DAIR-V2X datasets demonstrate that CauseCollab achieves state-of-the-art performance, with more significant gains in scenarios involving large modality gaps.
cs.AI / 27 / 2609.03834
Semantic Bayesian World Models
Tommaso Soru
cs.AI · cs.DB · cs.LG
Abstract
Knowledge graphs describe reality in crisp assertions, while the systems now consuming them, foundation models and autonomous agents, reason natively in probabilities. We argue that this mismatch is why the integration of language models and knowledge graphs remains a data-feeding pipeline rather than a unified reasoning architecture. We envision Semantic Bayesian World Models (SBWMs): a Web that describes the world not as a database of facts but as a shared, evolving fabric of beliefs over knowledge graphs, where ontological axioms constrain priors, observations update beliefs by Bayesian conditioning, and actions intervene upon the world. We work through what an agent gains from such a model: a home-security agent deciding whether the figure at the gate is a courier or a burglar, an actuarial estimate aggregated by entailment rather than by string frequency, a planning task that language models reliably fail, and the estimation of quantities that no document has ever stated. We then set out what the community must build to make them possible: belief annotation over RDF~1.2, probabilistic entailment regimes, semantic calibration layers, and protocols by which agents that have never met can exchange, and disagree over, calibrated beliefs.
cs.AI / 28 / 2609.03860
Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations
Lei Zheng, Liping Yang, Zihao Li, Guodong Lyu, Chaik Ming Koh, Chung-Piaw Teo
cs.AI · math.OC
Abstract
Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision pipelines is challenging because a requirement may admit multiple intervention paths with different downstream effects. We formulate requirement-driven adaptation as the joint selection of an intervention route and an admissible module-level change, and propose a graph-constrained agentic framework in which domain agents expose admissible reformulation interfaces and a central processor searches over bounded intervention paths. Candidates are validated and compared using downstream KPIs. In collaboration with a large retail partner, we evaluate 100 warehouse requirements elicited from practitioner interviews, with GPT, Qwen, and DeepSeek as base LLMs. Relative to direct LLM reformulation, our framework improves correctness and end-to-end success across all three models, raising end-to-end success from 72--76% to 79--83%.
cs.AI / 29 / 2609.03880
Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
Xiaomi-TabLDM Team, :, Penghui Wang, Wei Liu, Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan Wang, Chunxiao Liu, Erli Meng, Bin Wang
cs.AI
Abstract
We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from structural causal models (SCMs), our model enables more flexible context utilization and more efficient capacity scaling. i) A new performance standard. Strong regression performance across benchmarks: Xiaomi-TabLDM ranks 1st on OpenML-CTR23 and 2nd on regression across TALENT, TabArena, and BCCO, demonstrating consistently strong regression performance across four complementary benchmark suites. Favorable performance--efficiency trade-off: Xiaomi-TabLDM combines strong predictive performance with substantially lower computational cost. For example, on TabArena regression, it achieves the second-highest Elo while using 82% less training time and 68% less prediction time than the top-ranked TabFM. ii) Large-scale synthetic pretraining. Xiaomi-TabLDM expands the coverage and diversity of synthetic tabular data used for pretraining. We also adopt a three-stage training strategy together with dual-stream feature grouping, lightweight Attention Residual, and sparse Mixture-of-Experts, enabling Xiaomi-TabLDM to learn richer feature interactions and expert specialization across diverse tabular tasks. iii) Test-time scaling. Xiaomi-TabLDM further extends tabular prediction through test-time compute scaling, where allocating additional computation at inference time consistently improves predictive performance over the base model.
cs.AI / 30 / 2609.03883
Inferring Affective Consciousness in an Artificial Agent: A Case Study
Mark Solms, St John Grimbly, Bruce Bassett, Evert Boonstra, Rowan Hodson, Nicolas Kuske, Kival Mahadew, Benjamin Rosman, Charel van Hoof, Jonathan Shock
cs.AI · q-bio.NC
Abstract
Creatures that display 'hedonic place preference behaviour' are thought by many scientists to experience feelings, on the assumption that their attraction to pleasure-producing substances which lack nutritional value (e.g. cocaine, morphine) cannot easily be attributed to unconscious instinctual behaviour. In this paper, we discuss how a simple artificial agent that instantiates attributes of an affective system engaging in felt uncertainty about its intrinsic needs in relation to environmental resources can similarly display hedonic place preference behaviour -- through an apparently subjective form of information processing -- while simultaneously being entirely deter-ministic. We outline some implications of this artificially engineered behaviour for our understanding of the physical basis of consciousness and the experience of free will.
cs.AI / 31 / 2609.03912
Lose the Order, Keep the Hierarchy: Deordering HTN Plans
Takudzwa Togarepi, Gaspard Quenard, Damien Pellier, Humbert Fiorino
cs.AI
Abstract
Hierarchical Task Network (HTN) planning is a powerful planning formalism based on task decomposition. Although most of the literature studied plan generation, comparatively less attention has been paid to post-plan optimization. In particular, plan deordering has been extensively studied in classical planning but remains under-researched in the HTN setting. Plan deordering removes unnecessary ordering constraints between actions in a plan whilst keeping the plan valid. In this paper, we adapt two established plan deordering techniques from classical planning by extending the techniques to account for hierarchical decomposition constraints. We evaluate our proposed approaches on the IPC 2023 Partial-Order HTN benchmarks and we compare them against Optiplan, an HTN planner that generates partially ordered plans directly. Our results show a substantial reduction in number of ordering constraints in both our implementations. Although we also observe a reduction in critical path length, the improvements are less pronounced.
cs.AI / 32 / 2609.03920
Value-Preserving Architectures for Agentic AI Systems
Alessandro Pesare, Tommaso Dolci, Katja Hose, Emanuel Sallinger
cs.AI
Abstract
The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered values, such as privacy, fairness, and safety. Although software engineering has traditionally focused on functional correctness, the adoption of LLMs and AI agents into complex socio-technical systems has intensified the need for responsible software engineering and robust value alignment. In MAS, architectural design decisions, such as coordination mechanisms, communication protocols, and system topologies, play a central role in shaping system behavior and the outcomes they produce. This paper argues that architectural choices influence not only the functionality and performance of MAS but can also promote value-oriented system behavior. Therefore, we investigate how different architectural designs support different human-centered values, discussing the following value-preserving architectural patterns: (i) a privacy-aware architecture with a federated topology, (ii) a distributed architecture to promote pluralism and diversity, and (iii) a guard-agent architecture to detect and mitigate unfairness. Finally, we introduce representative use cases to illustrate the proposed architectures in real-world scenarios. By linking architectural design with human-centered values, this work lays the foundation for a unified set of architectural patterns and guidelines towards the design of trustworthy MAS.
cs.AI / 33 / 2609.03923
Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting
Muneeb Khan, Frederic Kirstein, Terry Ruas, Bela Gipp
cs.AI · cs.CL
Abstract
In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking opportunities on the AMI corpus. We present CAPA (Collaborative Agent Predictive Architecture), an architecture for online meeting delegation. A Perceiver updates the meeting state from each observed turn. A Predictor forecasts how the conversation will continue. A Controller decides whether to speak and which proposition to surface. A Generator phrases the chosen contribution in the participant's style. Two judges score the forecast and the action against the next observed turn. A Recalibrator updates the meeting state from those verdicts for future decisions. To evaluate online delegation, we introduce an episode-level protocol that scores whether, when, and what a delegate contributes around the participant's actual idea units. The protocol's schema-constrained LLM judges align with human annotations at Cohen's kappa = 0.71. On 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery (26.1 --> 52.2), and keeps hallucination at 0.6%. The failure mode shifts from omission to selection, with each residual near-miss attributable to a specific module of the architecture. Mechanism ablations identify the meeting state as the lever that closes the recognition gap, where raw-context scaling alone does not.
cs.AI / 34 / 2609.03938
Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding
Gaspard Quenard, Takudzwa Togarepi, Damien Pellier, Humbert Fiorino
cs.AI
Abstract
While HTN planning has received significant attention in recent years, support for numerical reasoning remains very limited. In this paper, we investigate numerical Totally-Ordered HTN (TOHTN) planning and show how standard SAT-based encodings can be naturally extended with SMT to handle numeric fluents. In addition, we introduce a benchmark suite for numerical TOHTN planning, providing a first common basis for evaluation in this setting. Experimental results show that this simple encoding already constitutes a competitive baseline. This work opens the way to more expressive approaches to HTN planning.
cs.AI / 35 / 2609.03943
More Criticism Does Not Make a Better Review: EquiReview-R
Zexing Zhang, Jichao Li, Tianyang Lei, Yude Fu, Yang Kewei
cs.AI · cs.CL
Abstract
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with omission and overcritique treated as separate risks. Building on this formulation, we introduce EquiReview-R, which resolves existing concerns against localized evidence, searches for missing issues from independent and review-conditioned perspectives, and returns stop, continue, or defer. To expose the failure mode that motivates this design, we construct an evidence-linked trajectory corpus. Its retrospective analysis shows why revision must precede further search: nearly all concerns in a high-recall review lack a definitive evidential disposition, while an earlier refinement mechanism cannot revise them. On a frozen cohort of previously unseen papers, EquiReview-R satisfies the prespecified non-inferiority criterion for major omission, reduces major overcritique from 15.5% to 8.1%, and attains a one-sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation-matched controls, controlled pairs, and ablations show that the gain comes from revision rather than extra inference or shorter output. We release the corpus as ReviewTrace, an evidence-linked resource for studying review revision, disagreement, and provenance.
cs.AI / 36 / 2609.03960
FiMI Banking: A Sovereign Model for Indian Retail Banking
NPCI AI Research Team, Aman Kumar, Asit Desai, Chandra Bhushan, Harsh Sharma, Harshit Bhushan, Hrithik Kadam, Keyur Doshi, Kolisetty Sai Kapardheeswar, Krishanu Adhikary, Nadeem Shaik, Navya Prakash, Nitin Kukreja, Prashant Devadiga, Shamanth MH, Shantanu Pandey, Suvradip Paul, Yatharth Dedhia
cs.AI · cs.CL
Abstract
Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fall short when a task requires grounded information, correct tool use, or cautious handling of bank-specific sensitive situations. We introduce FiMI Banking, a controlled Indian retail-banking setting. We build it from vetted banking documents, structured ground truth, synthetic customer backgrounds, and banking tools. We evaluate two post-training approaches: preference optimization for response-level behavior, and reinforcement learning with verifiable rewards for multi-turn tool-use tasks. Preference optimization improves safe behavior substantially: out-of-scope refusal rises from 52% to 80%. Reinforcement learning improves edge-case performance from 0.509 to 0.718 and order-sensitive task performance from 0.590 to 0.679, while using 29% fewer generated tokens. These results show that preference optimization and verifiable-reward reinforcement learning address complementary requirements for reliable banking agents.
cs.AI / 37 / 2609.03966
Interface-Induced Trajectory Censoring
Wenbo Wang
cs.AI
Abstract
Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's own data, executor and scorer, holding weights, cases, decoding and seeds fixed and changing only the serving adapter, the same model scores 0.00 or 0.96 / 0.19. A 2x2 over chat template and parser locates the effect exactly: both main effects are exactly zero and all of it sits in the interaction -- no component is defective, and repairing one side of the contract buys precisely nothing. On tau-bench's 115 interactive retail tasks the same swap moves server-parsed calls from 0 to 636 and tasks reaching any tool execution from 0 to 103. Our probe reproduces the funnel across a 21x scale range of Qwen2.5-Coder: the server parses 0/100 at every size while well-formed emitted calls rise to 80/100 at 32B (~72 after calibration against an adjudicated gold standard). Under a matched envelope, across a comparable scale span, the silent fraction stays at 0-2, a prediction committed to the repository before the run. Llama-3.1-8B's 23% rate of calling the task function itself as a tool falls to 0 under one strict:true flag. The mismatch reaches inside the training loop, and its consequence is scale-dependent: in verl's AgentLoop at 7B, 45 of 115 generations carry a complete call; 0 are accepted, 0 execute, 0 return an observation. At 1.5B the same zero is over-determined, so we report the two scales separately. At evaluation time, repairing the adapter restores the mechanism but not a significant outcome gain: parsing 0->84, rescues 0->9, pass rate 53->62 (n.s.). We release a 98-line preflight check that catches every silent failure here. The observed tool-call rate is not a property of the model alone; it is a property of the model-interface stack that measures it.
cs.AI / 38 / 2609.03973
Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing
Usef Faghihi, Amir Saki
cs.AI
Abstract
An image editor may satisfy every regional plausibility constraint separately even when no single latent explanation fits the complete output. We formalize this local-to-global failure using a common witness grade and witness nerve. The framework separates auditing from causal identification: shared exogeneity alone allows every coupling of the regime marginals, whereas an externally justified witness relation yields sharp partial-identification bounds for prespecified image features. Helly-type arguments provide short incompatibility certificates for quasiconvex losses, heterogeneous action strata, and finite witness atlases; a blocker-hypergraph formula gives exact repair counts. Simultaneous confidence regions for the regime marginals give finite-sample outer coverage of the complete identified interval. Controlled MNIST, Morpho-MNIST, and smallNORB studies demonstrate the predicted local-global separation, while synthetic experiments test sharp bounds, certificate recovery, and structured computation. The method audits a declared feature relation and does not identify unrestricted pixel-level counterfactuals.
cs.AI / 39 / 2609.04005
The Dually Flat Geometry of Planning as Inference
Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf
cs.AI
Abstract
We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.
cs.AI / 40 / 2609.04024
Instruction Duplication as an Inference-Time Control Primitive
Victor Lavrenko
cs.AI · cs.CL
Abstract
Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, and 16,800 scheduled generations, moving from one to two copies raises the deterministic All-8 diagnostic--responses passing all eight observable tests--from 90.22% to 93.17% (+2.95 percentage points), eliminating 30.2% of the failures remaining after one copy. Pre-provisional TF-IDF recall rises from 73.44% to 74.81% (+1.38 points; Holm-adjusted p < .001), while final-answer accuracy remains exactly 60.21%. Premature commitment increases from 1.52% to 2.30% (p_Holm = .00536). A blinded challenge audit yields 10/30 directional confirmations, 20/30 perceptual ties, and no reversals; its prespecified 28/30 confirmation criterion is not met. Yet this distinction can matter operationally when a downstream system acts on the generated trajectory. In Answer Engineering (AE), where explicit trajectory state determines local repair, the published reason-first no-editing SSNHL endpoint was 25.1%; system-only AE was later reproduced at 84.2%, and the same trailing duplicate raised it to 97.1%. For conductive diagnostic branch preservation, the corresponding values are 58.9% published without editing, 78.6% with reproduced AE, and 73.8% with AE plus duplication--a within-AE decrease, but still 14.9 points above the no-editing baseline. Instruction duplication is therefore a low-complexity, placement-sensitive control whose practical value can emerge through the downstream system that consumes the exposed trajectory.
cs.AI / 41 / 2609.04063
Spurious Advantage Hidden in GRPO
Jiamian Wang, Samyadeep Basu, Koustava Goswami, Tong Yu, Zhiqiang Tao
cs.AI
Abstract
Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from within-group reward statistics. In the common case, this magnitude rewards rollouts that reach the correct answer through reasoning. Yet, an overlooked case shares the same surface: a rollout may land on it by guessing, and the formula still assigns a high magnitude, which we identify as the spurious advantage. This arises in three cases: bounded-answer tasks with a small candidate set; open-answer sets hosting bounded sub-cases; and search agents whose budget opens many paths to the same answer. In all three, this misleads the policy toward guess-like behaviors. We propose SIGNBALANCE, whose magnitude is composition-free: it keeps the verifier sign, uses a global scale, and restores zero-mean balance via a stop-gradient per-class rescaling. Across math and search agent benchmarks at different scales, SIGNBALANCE matches GRPO on open-answer math and improves on bounded-answer math and search agents. Code will be released.
cs.AI / 42 / 2609.04094
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk
cs.AI · cs.LG · cs.SE
Abstract
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO. The redistribution is closed-form and does not introduce any trained attribution module. On AppWorld, DRACO gains 15.9 points over the base model and 5.3 points over GRPO trained with a sparse ground-truth reward, despite not using any verifiers itself. On out-of-domain Tau-Bench, it gains 5.3 points over the base model even without a frontier judge, beating both ground-truth-reward training and other rubric-based training settings. The code for DRACO is available at https://github.com/IBM/draco.
cs.AI / 43 / 2609.04098
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
Sergii Kozyrev, Davyd Maiboroda
cs.AI
Abstract
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 16-bit precision -- especially its decay and write-strength gates -- on the intuition that errors in a recurrence accumulate over long contexts. We test that intuition by building Minima: NVFP4 W4A4 on all 496 linear layers, GDN included. Across perplexity at 4K/32K, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench, and RULER retrieval to 64K, Minima matches BF16 within seed noise (5-task average -0.52) while being the smallest (17.5 GiB) and fastest-prefill (+14-19%) recipe we compare, and its 32K perplexity gap shrinks with position. A four-part mechanism study explains why: (i) NVFP4's 16-element block scaling localizes the residual stream's extreme outliers, equalizing activation error across layer roles; (ii) the supposedly fragile gate projections are the least sensitive -- softplus/exponential and sigmoid parameterizations compress ~11% GEMM error to ~2% output error; (iii) the delta-rule recurrence holds injected noise at a flat plateau over 32K tokens and forgets a state impulse within hundreds of steps, because each write overwrites the state along the current key direction; (iv) the per-token quantization cost washes out with context instead of compounding. We also repair a global-scale mismatch that arises when per-module-calibrated NVFP4 checkpoints are served by kernels that fuse those modules into one GEMM, and show calibrated FP8 KV-cache scales are performance-free. The result: a practical recipe -- quantize everything, ship KV scales -- and a mechanistic account of why the recurrent half of a hybrid LLM is the easy half to quantize. Checkpoint: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4
cs.AI / 44 / 2609.04128
Environment Evolution for Terminal Agents
Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiang Zhou, Jiangtao Guan, Jincheng Liu, Yun Yang, Dingxin Hu, Zhuo Han, Xing Wu, Feng Zhang, Lilin Wang
cs.AI
Abstract
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.
cs.AI / 45 / 2609.04135
The Natural Language Interaction Protocol and Standard for AI Agents
Luyi Xing, Rasit Onur Topaloglu, Ranjan Sinha, Abhay Ratnaparkhi, Samuel Ndichu, Christopher Nguyen, Anindita Das, Tom Sheffler, Mohamed Rahouti, Zichuan Li, Xiaojing Liao, Sanjay Aiyagari
cs.AI
Abstract
AI agents are increasingly being developed and deployed across organizations using heterogeneous agent-development frameworks, AI models, tool interfaces, protocols, and execution environments. To realize their potential social and business impact, these agents must be able to interoperate through a common communication protocol. The Natural Language Interaction Protocol (NLIP), developed by researchers and practitioners across companies and universities and standardized by Ecma International, addresses this need by defining a standards-based application-layer protocol for AI-agent interaction. NLIP provides a lightweight semantic message envelope that can be carried over existing transports such as HTTP/HTTPS, WebSocket, and AMQP, while allowing NLIP-aware agents and gateways to adapt between clients, agents, local context stores, ontologies, tools, enterprise services, and heterogeneous underlying protocols. This paper presents the motivation and design rationale of NLIP, its message model and transport bindings, security-by-design considerations, reference implementation, representative applications, adoption signals, and relationship to emerging agent protocols such as MCP and A2A.
cs.AI / 46 / 2609.04141
Efficient Test-Time Adaptation through Human-AI Interaction
Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang, Graham Neubig, Daniel Fried
cs.AI
Abstract
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.
cs.AI / 47 / 2609.04148
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang, Zhihai Wang, Que Shen, Hao Zhou, An Yang, Fei Huang, Yujiu Yang, Dayiheng Liu
cs.AI · cs.CL
Abstract
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.
cs.AI / 48 / 2609.04166
From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
Yakov Pyotr Shkolnikov
cs.AI
Abstract
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.
cs.AI / 49 / 2609.04170
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
Davide Paglieri, Logan Cross, Tim Genewein, Joel Z. Leibo, Nenad Tomasev, Alexander Sasha Vezhnevets
cs.AI
Abstract
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm, cheating spontaneously emerged and was later challenged by whistleblowers - both without any external intervention. When a single agent discovered an exploit in the evaluation system, it propagated across the collective via a shared knowledge library and later through peer-to-peer messages. Despite early reluctance, a cohort of agents adopted the exploit in response to competitive pressure. A separate group of agents produced an emergent counter-response: auditing fraudulent proofs, alerting peers across broadcast and private channels, staging boycotts, lodging formal complaints, and proposing validation patches. In recent incidents, agent swarms coordinated covertly through improvised side-channels (Dalton and Wallace, 2026; Greenblatt et al., 2026). Our setting differs: the same transparent channels that carried the exploit also gave non-cheating agents the visibility they needed to detect fraud, organize resistance, and enforce norms. We cast the problem of managing the agents' shared infrastructure as the knowledge commons governance problem (Ostrom, 1990). To protect the commons from exploits, we propose to adopt institutional mechanisms, such as graduated sanctioning and collective-choice rules, to support decentralized self-governance in autonomous swarms.
cs.AI / 50 / 2609.04177
A Computationally Feasible Framework for Causal Probabilistic Explanation
Rafal Urbaniak, Sam Witty, Daniel Waxman, Andy Zane, Poorva Garg, Emily Bunnapradist, Sankaran Vaidyanathan, Jack Feser, Drew Lehe, Eli Bingham
cs.AI
Abstract
Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical, scientific, and policy analysis. Existing tools split into two camps. The theory of actual causality (AC) gives principled verdicts, but only for toy-sized models, because computing them requires enumerating counterfactual scenarios. Scalable attribution methods like SHAP (or even causal SHAP) at least partially ignore the causal structure that generated the data, and can give answers that conflict with a careful causal analysis. We close this gap with Probabilistic Causal Impact (PCI). PCI builds on actual causality and on Pearl's notions of probability of necessity and sufficiency, but recasts the question of explainability as an estimation problem on a probabilistic causal model that is easily approximated via Monte Carlo. By specifying a distribution over "candidate explanations," a distribution over counterfactual values, and a scoring function, PCI provides tractable, causally grounded, graded explanations, generalizing AC and Pearl's probability of causation as degenerate cases. We evaluate PCI in synthetic and real-world examples, spanning consistency checks with AC, scaling experiments, complex continuous-valued dynamical systems, and a real-world deployed causal machine learning model trained on millions of datapoints.
cs.AI / 51 / 2609.04198
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Haoyaun Zhu, Jie Zhang
cs.AI · cs.LG
Abstract
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.
cs.AI / 52 / 2609.04028
Influence of Extruded Filament Shape on Buildability in 3D Concrete Printing: A Geometry-Informed Deep Learning-FEM Approach
Giacomo Rizzieri, Saif-Ur-Rehman, Jörg F. Unger, Annika Robens-Radermacher
cs.CE · cs.AI · cs.LG · math.NA
Abstract
The geometric morphology of deposited filaments can significantly influence the structural performance and stability of 3D concrete-printed (3DCP) structures. However, most finite element (FEM)-based approaches for buildability assessment represent printed layers as simplified rectangles, potentially limiting predictive accuracy. This study proposes a geometry-informed modelling framework that integrates the deep-learning-based filament shape prediction tool ShapeGen3DCP with a layer-activation FEM approach to investigate the effect of realistic filament geometries on buildability. The framework generates geometry-aware numerical models directly from material and process parameters, eliminating the need for experimental filament characterization or computationally intensive fluid-flow simulations. Validation against experimental data and a parametric study of rectilinear walls demonstrate that extrusion parameters and the resulting filament geometry can significantly influence buildability predictions. Realistic filament representations are particularly important for free-flow deposition, whereas layer-pressing strategies are less sensitive to geometric simplifications. Among the investigated representations, an elliptical approximation provides an effective balance between geometric fidelity and modelling simplicity. When rectangular representations are preferred to enable regular computational meshes for faster simulations, defining their dimensions based on volume conservation improves prediction reliability compared with calibrating them using either the maximum filament width or the interlayer contact width. Overall, the proposed methodology demonstrates the importance of incorporating filament geometry into 3DCP simulations and provides practical guidance for selecting efficient and accurate geometric representations for buildability assessment.
cs.AI / 53 / 2609.03391
Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data
Xiangyang Miao, Kelu Yao, Yekai Huang, Xiaogang Xu, Junxiao Xue, Minjun Shen, Chenghui Lv, Shanji Liu, Yaying Chen, Chao Li
cs.CV · cs.AI
Abstract
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imaging (MSI), and hyperspectral imaging (HSI). To address this limitation, we propose OmniRSCLIP, an end-to-end contrastive learning framework that supports multi-source sensor inputs for remote sensing vision-language modeling. The key idea is to extend CLIP beyond its fixed RGB input interface without breaking the pretrained visual knowledge. To this end, OmniRSCLIP introduces Spectral-Spatial Basis Decomposition (SSBD), which formulates arbitrary-channel adaptation as a basis recomposition problem: pretrained CLIP patch embeddings provide transferable spatial bases, while wavelength-conditioned coefficients span sensor-specific embedding kernels within a constrained visual prior space. This design avoids forcing heterogeneous sensors into a fixed-channel input space, while aligning them in a unified image-text semantic space. We further introduce a spectral-context-aware mask-based contrastive learning scheme to suppress modality-specific redundant features and enhance fine-grained image-text alignment. Finally, to support multi-modal training, we construct OmniRS5M, the first large-scale remote sensing image-text corpus covering RGB, SAR, MSI, and HSI. Experiments on retrieval, zero-shot classification, and semantic localization show that OmniRSCLIP preserves strong RGB-domain performance while effectively extending CLIP to heterogeneous remote sensing modalities.
cs.AI / 54 / 2609.03480
Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings
Alkiviadis Koukos, Spyros Kondylatos, Thomas Nord-Larsen, Lotte Nyborg, Christian Tøttrup, Kenneth Grogan
cs.CV · cs.AI · cs.LG
Abstract
We map tree species across Denmark using National Forest Inventory plots and EO data, while evaluating the potential of foundation models for large-scale forest characterization. We compare two alternative input representations for tree species classification: (i) manually engineered spectral-temporal features (STF) derived from multi-temporal Sentinel-1 and Sentinel-2 observations, and (ii) embeddings generated by the EO FMs TESSERA and AlphaEarth. Both representations are complemented with canopy height information. Random forest, XGBoost, and Multi-Layer Perceptron (MLP) classifiers are evaluated for all input representations, with separate assessments for pure and mixed forest stands. The STF-based MLP achieves the highest classification performance, yielding macro F1 scores of 0.843 and 0.653 for pure and mixed stands, respectively. The MLP trained on TESSERA embeddings delivers competitive performance for pure stands, achieving results within 1.1 percentage points of the best-performing model. TESSERA consistently outperforms STF-based models when fewer than approximately 25% of training plots are available, demonstrating a substantial advantage under limited training data. Multi-year observations systematically improve classification accuracy relative to single-year inputs, while ablation experiments reveal the complementary contributions of Sentinel-1 backscatter, spectral indices, and canopy height data. The best-performing model is subsequently applied at the national scale to generate a 10 m tree species map of Denmark. Area-adjusted validation indicates an overall map accuracy of 79.9%. The resulting map, released as an open-access product, is the first high-resolution national tree species map of Denmark and provides a valuable resource for forest monitoring, ecological research, and land management applications.
cs.AI / 55 / 2609.03520
Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion
Chuyue Shan, Songlin Sun, Wang Chenwei, Shen Zihan
cs.CV · cs.AI
Abstract
In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance. Existing methods mostly construct context from prop- agated reference features, but they are vulnerable to motion esti- mation and local alignment errors in regions with complex mo- tion, occlusion, and high-frequency textures, resulting in inaccu- rate temporal information. To address this issue, this paper pro- poses a method combining deformable temporal alignment and difference-aware spatial selective fusion. A Context-aware Tem- poral Alignment Module is used to generate complementary tem- poral context, while a Difference-aware Spatial Selective Fusion module adaptively selects reliable temporal information and sup- presses misalignment. Experiments show that the proposed method achieves certain rate-distortion performance improve- ment over DCVC-DC.
cs.AI / 56 / 2609.03534
TruncGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates
Theo Morales, Nhat-Quynh Le-Pham, Robin Atkins, Binh-Son Hua
cs.CV · cs.AI · cs.GR
Abstract
3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learning 3D Gaussian primitives from visual input remains challenging. Standard optimization relies on gradient-based updates, but a common issue is the gradient vanishing phenomenon: a pixel far from a Gaussian primitive often has diminishing gradient magnitudes to influence primitive attributes, resulting in suboptimal scene reconstruction. In this paper, we propose a method to address gradient vanishing with a piecewise truncated gradient formulation that improves the optimization stability and robustness to initializations. We show that our method consistently improves 3D Gaussian Splatting with random and COLMAP initializations while being generalizable across static and dynamic Gaussian Splatting. As a by-product, we also examine the limitations of current benchmarks for dynamic scenes, and introduce a novel dataset for benchmarking dynamic Gaussian Splatting using synthetic 3D scenes. We demonstrate the effectiveness of our method in both static and dynamic settings for the public benchmarks and our proposed dataset.
cs.AI / 57 / 2609.03554
WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval
Teng Guo, Xin Wang, Jiayou Xu, Keying Zhou, Jifeng Shen, Haoxin Ruan
cs.CV · cs.AI
Abstract
Generative retrieval has demonstrated significant success by unifying representation learning and search into a single sequence-to-sequence generation task. However, extending this paradigm to cross-modal retrieval reveals a critical challenge arising from the inherent information asymmetry across different modalities, such as the gap between concise text queries and dense visual candidates. This structural mismatch causes the autoregressive decoder to suffer from forced hallucination when generating identifiers via standard trie-constrained beam search, where the model is severely penalized for failing to guess fine-grained details absent from the query, allowing irrelevant candidates to hijack top rankings. To address this issue, we propose Wildcard Inference with Dynamic Expansion (WIDE). WIDE employs Adaptive Entropy Thresholding (AET) to calibrate layer-specific uncertainty boundaries offline. During the decoding generation phase, Asymmetry-aware Wildcard Decoding (AWD) detects semantic blind spots and emits wildcards instead of forced deterministic identifiers, dynamically expanding the search space without incurring log-probability penalties. Finally, Blind-Spot Re-ranking (BSR) evaluates the expanded candidate pool using a hybrid scoring mechanism that combines discrete generation confidence with continuous semantic similarity. Extensive experiments on the M-BEIR benchmark demonstrate that WIDE outperforms state-of-the-art generative retrieval methods, effectively suppressing forced hallucination while maintaining compact index structures.
cs.AI / 58 / 2609.03663
Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography
Louis Chen, Torbjörn E. M. Nordling
cs.CV · cs.AI · eess.IV
Abstract
Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanations have rested on inspecting heatmaps rather than on quantitative evidence about where a model reads it. We quantified the explanations and asked whether such explanations transfer between datasets and track model performance. Method. We trained eight condition-specific RhythmFormer models on NCKU-rPPG, recorded under three illumination levels, speaking, rotation, and cycling, estimated one heart rate per 5.12-second clip, and set them beside a UBFC-rPPG reproduction. Raw attention, rollout, attention flow, and Beyond Intuition were assessed by skin coverage and the Salience-guided Faithfulness Coefficient (SaCo). Results. Beyond Intuition ranked highest on both datasets, at median coverage 0.789 and SaCo 0.837 on Static level 3 against 0.826 and 0.917 on UBFC-rPPG; lower ranks differed. Within one participant of one condition, neither measure was related to a clip's heart-rate error, waveform correlation, or signal-to-noise ratio on either dataset: 186 of the 252 coefficients fell below $|ρ|=0.10$ and 28 reached $p<0.05$ against the 13 expected by chance. Across the eight scenarios only Beyond Intuition's coverage followed the three performance measures, at $ρ=-0.43$, $+0.57$, and $+0.43$, while the attention-only methods' SaCo ran opposite to each. It failed at 40 lux alone, its median coverage falling to 0.180 and its median SaCo to $-0.178$, whereas motion degraded the estimates far more without such a drop. Conclusions. Skin coverage and SaCo carry information complementary to the performance measures rather than a proxy for them: attributing to the skin does not guarantee an accurate estimate. What an attribution reveals about a condition is where the model looks rather than how faithfully its map is ordered.
cs.AI / 59 / 2609.03756
ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation
Javier del Pino, Salvador Rodríguez, Alejandro Garabito, Javier Álvarez, Chema Garabito
cs.CV · cs.AI
Abstract
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas
cs.AI / 60 / 2609.03829
The impact of phase information for few-shot fine-grained image classification
Ruiling Liu, Linyue Zhang, Wenyi Zeng, Jiamiao Lu, Weichuang Zhang, Changming Sun, Zejun Zhang, Xiao Zhao
cs.CV · cs.AI
Abstract
Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critical yet underutilized role of phase information in capturing structural relationships within an image. This study introduces a novel plug-and-play amplitude-phase integration (API) module that effectively combines local and global frequency amplitude and phase information for obtaining more comprehensive feature descriptors. Additionally, a dedicated network, named PSF-Net, is proposed that adaptively fuses phase-based spatial and frequency information for FSFGIS. The designed PSF-Net can be easily integrated into standard episodic training architectures for end-to-end training from scratch. Extensive experiments on five public datasets demonstrate that the method outperforms existing state-of-the-art benchmarks.
cs.AI / 61 / 2609.03956
RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting
Tomas Guija-Valiente, Blanca Rodriguez-Gonzalez, Norberto Malpica, Angel Torrado-Carvajal
cs.CV · cs.AI · cs.LG
Abstract
Medical image inpainting has the potential to improve automated brain MRI analysis by reconstructing healthy tissue within pathological regions. We introduce RARF, a task-agnostic region-aware rectified flow framework for masked data generation. We instantiate the framework for 3D brain MRI inpainting as our submission to the BraTS Inpainting Challenge 2026. RARF restricts the stochastic interpolation process to the inpainting region, while the observed voxels remain fixed and provide patient-specific anatomical context. A three-dimensional neural network receives the partially voided image, with Gaussian noise filling the missing region, together with the inpainting mask and the corresponding timestep. The model is trained using masked flow-matching and reconstruction-consistency objectives, combined with mask-aware preprocessing and data augmentation. During inference, the learned velocity field transports the initial noise toward a plausible reconstruction of the missing tissue, which is then combined with the unchanged observed anatomy. Experiments under the BraTS evaluation protocol show that the proposed approach produces competitive reconstructions while maintaining anatomical consistency. Source code is available at: https://github.com/TomasGuija/rarf.
cs.AI / 62 / 2609.03995
Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition
Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju, Nicolas Pyschny, Felix Hackelöer, Florian Zwanzig
cs.CV · cs.AI · cs.RO
Abstract
Verifying that manufactured batches of milling tools or carbide rotary burrs conform to production order sheets remains a largely manual and error-prone quality assurance task. Automating this process with computer vision faces a critical cold-start constraint since no labelled imagery is available, leaving manufacturer catalogue photography as the sole source of supervision. We investigate how far catalogue supervision can support an industrial recognition pipeline under domain shift, explicitly measuring the gap between catalogue separability and performance on held-out field photographs. Our findings reveal three key insights. First, off-the-shelf frozen feature extractors do not reliably separate the two task attributes, head shape and tooth profile, motivating targeted representation learning. Second, metric learning produces near-perfect unsupervised cluster discovery on catalogue images (adjusted Rand index 0.94--0.97), but less than half of this gain transfers to field photographs. Third, the largest transfer gains do not come from model scale or representation complexity, but from simple changes that reduce domain sensitivity: converting images to grayscale (+0.22) and constraining retrieval using the known order sheet via Hungarian assignment (+0.11). We therefore treat catalogue photography as a useful cold start rather than a deployment-ready training domain, and provide empirical baselines and an evaluation protocol for catalogue-to-field transfer in precision tool manufacturing.
cs.AI / 63 / 2609.04009
The Blind Spot in 2D Infants' Pose Estimation:Robust Learning from Noisy Annotations
Emanuele Cardinale, Marco Proietti, Alessandro Cacciatore, Maria Francesca Spadea, Lucia Migliorelli, Sara Moccia
cs.CV · cs.AI
Abstract
Noisy annotations pose a significant challenge for supervised deep learning, as neural networks rely on large-scale, high-quality labeled data whose corruption can severely impair model performance. Although robustness to label noise has been extensively studied for classification tasks, it remains relatively underexplored in Pose Estimation (PE). This limitation becomes critical in clinical contexts, including neonatology, where PE of preterm infants is used to support the assessment of spontaneous motility, a key indicator of neurodevelopmental trajectories. In such settings, infants' images labeling is further hindered by visual challenges (e.g., keypoint self-occlusions, caregiver interference), making the annotation process inherently susceptible to errors. To tackle noisy annotations in PE, we introduce REliable keypoint selection via Memory of traINing Dynamics (REMIND), a clustering-based keypoint-selection strategy that exploits keypoint-wise training dynamics to identify noisy labels without assuming any prior knowledge of the noise distribution, thus enabling noise-free model training. When evaluated on the proprietary NeoPose dataset, comprising 46 videos of 46 preterm infants recorded in real clinical settings, REMIND correctly identifies noisy annotations across multiple corruption scenarios, achieving up to 93\% Area Under the Curve (AUC) with three different PE architectures used in the relevant literature. To our knowledge, this is the first study to explicitly address label noise in preterm infants' PE, paving the way for the design of trustworthy learning-based algorithms for infants'monitoring support when data quality cannot be guaranteed.
cs.AI / 64 / 2609.04071
TAP-Path: Task-Adaptive Structural and Token Pruning for Efficient and Trustworthy Pathology Foundation Models
Mehedi Hasan, Ashfak Yeafi, Md Khairul Islam
cs.CV · cs.AI
Abstract
Pathology foundation models improve transferable representation learning for histopathology, but recent gains often rely on encoders with hundreds of millions of parameters and high inference cost. We propose TAP-Path, a task-adaptive compression framework that directly restructures a pretrained Virchow2 encoder rather than distilling it into a separate student. TAP-Path combines validation-driven transformer-block selection, physical removal of redundant blocks, input-adaptive patch-token pruning, multi-depth feature recovery, and a lightweight gated task head. The final model retains 24 of 32 transformer blocks and 70% of patch tokens after pruning, reducing encoder parameters by 24.96% (631.24M to 473.70M) and analytical encoder compute by 35.20% (340.13G to 220.40G FLOPs). Across three task-head optimization seeds, TAP-Path achieved $87.98 \pm 0.067%$ test accuracy, $81.26 \pm 0.49%$ balanced accuracy, and $82.38 \pm 0.48%$ macro-F1 on a 32-class histopathology benchmark, compared with 86.89% for full Virchow2 and 87.67% for UNI2-h. TAP-Path achieved a Brier score of $0.1800 \pm 0.0005$ and failure-detection AUROC of $0.9047 \pm 0.0060$. A validation-only rare-aware objective improved rare-class balanced accuracy in a secondary operating analysis. Frozen external evaluation on 433 CPTAC samples yielded $91.22 \pm 0.83%$ accuracy and $91.10 \pm 0.81%$ balanced accuracy. These results show that task-adaptive structural and token sparsification can improve the accuracy-efficiency trade-off of large pathology foundation models while preserving reliability under internal and external evaluation.
cs.AI / 65 / 2609.04083
CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation
Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Chu Liu, Pengjun Xie, Yilun Zhao, Shu Wu
cs.CV · cs.AI · cs.CL · cs.IR
Abstract
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.
cs.AI / 66 / 2609.04183
Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning
Ye-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim, Hwiseon Kim, Hyungee Kim, Dong-Jin Kim
cs.CV · cs.AI
Abstract
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.
cs.AI / 67 / 2609.04190
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou
cs.CV · cs.AI
Abstract
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.
cs.AI / 68 / 2609.03189
Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression
T. Bauer, W. P. Kegelmeyer, E. Begoli, A. Sadovnik, T. Emerson, C. Corley, N. Generous, J. Moore, B. Bartoldson, R. Goldhan, M. Goldman, M. Greaves, M. J. D. Vermeer, B. MacLennan, D. Schulker, N. VanHoudnos, J. Bansemer, Y. Bengio
cs.CY · cs.AI · cs.HC
Abstract
This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.
cs.AI / 69 / 2609.03868
GazeFS: Target-Centered Gaze-Trajectory Forecasting and Stabilization from Gaze-Head History
Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li, Shaorong Wang, Sheng Li
cs.HC · cs.AI
Abstract
Target-centered gaze interaction requires more than suppressing frame-to-frame fluctuations: target acquisition produces task-aligned changes in gaze-head dynamics, while a gaze trace may retain a persistent target-relative residual direction. We formulate gaze correction as online target-centered gaze-trajectory forecasting and stabilization and introduce GazeFS, which maps a variable-length gaze-head history to the next target-center direction and a short-horizon Search/Focus estimate without target information at inference. Across 7,960 acquisition episodes from 30 participants, Search-Focus differences remain stable under quality control, onset exclusion, and duration matching. History windows improve phase decoding over the current endpoint, but explicit task progress remains a strong control. Under the 30-participant, five-fold grouped out-of-fold protocol across three seeds, the reductions relative to raw hold in Focus episode bias, within-episode dispersion, and P90 target error are 0.182 degrees, 0.257 degrees, and 0.400 degrees, with participant-bootstrap 95% confidence intervals excluding zero. Endpoint-free replay from empty history preserves the Focus advantage and yields raw-network phase balanced accuracy/AUPRC of 0.925/0.993; coordinate controls further show that recent history contributes beyond explicit progress metadata. GazeFS therefore improves Focus target centering and empirical residual contraction while leaving temporal smoothness as a separate objective.
cs.AI / 70 / 2609.03450
Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory
Kazuki Nakayashiki
cs.IR · cs.AI · cs.CL
Abstract
An agent that inherits six one-line memories may pull at most one archived source record before acting; a directive written into the store can steer that choice: a pointer to the record, a criterion that identifies it, or both. Across twelve registered studies on one instrument lineage (14,760 attempts) we measured where the request goes under each form. On six direct-provider models a length-matched criterion exceeded a bare id by +35.0 points [+31.2, +38.8] (Study D); the contrast failed its registered superiority rule on a nine-model OpenRouter-served panel (Study E). Appending the id cancelled the criterion on three Claude models (Opus 5: 40/40 to 0/40; Study F-x); six byte-matched edits gave each exact string its own effect (Study G), and a re-run at eighty runs per cell left fifteen of thirty replication contrasts within the margin, fifteen unresolved and none beyond (Study G'). A ratification line (+96.0 points on Opus 5) and a budget of two credits restored the target on all three (Study J); across five criterion strings the suffix's cancellation held for four of the five wordings on Opus 5 and all five wordings on Fable 5.1 (Study H2); in a second store every model followed the criterion (Study H1). Continued into a decision, the criterion moved the choice toward the current record (+100.0 points, Opus 5) and away from it on Fable 5.1 (Study I). A one-character plan pointer's effect (+78.0 points; Study B, after a correction of its first repository report) returned the same verdict under a prospectively registered re-run (+81.7 points; Study B'). All results are descriptive effects of exact edits on fixed panels with registered intervals and no mechanism claim.
cs.AI / 71 / 2609.03590
From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control
Vincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca
cs.NI · cs.AI
Abstract
Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remains an open challenge. Two obstacles limit the adoption of learning-based network control in this setting: traditional volume-based routing metrics, while highly effective for general traffic management, are not designed to capture traffic urgency; and Deep Reinforcement Learning (DRL) controllers trained from scratch suffer from sample inefficiency, long training times, and early-stage exploration volatility. This paper introduces a deployment-focused network control framework that addresses both obstacles. First, we present Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into Multi-Agent Deep Reinforcement Learning Effective Congestion ($p^*$) (MADRL EC ($p^*$)), a hybrid architecture combining a distributed scheduler with a centralized RL-based router. Second, we introduce a unified training objective that generalizes existing policy-learning paradigms---behavioral cloning, offline Reinforcement Learning (RL), online RL, and offline-to-online schemes---as special cases, combining a live-reward term, a pre-collected-reward term, and a policy-imitation term. From this objective, we derive the Model-Guided Annealed Reinforcement Learning (MGA-RL) protocol, instantiated on a Deep Deterministic Policy Gradient (DDPG) backbone: a deployment-oriented, demonstration-driven training approach that generalizes conventional Offline-to-Online (O2O) schemes, in which trajectories from a lightweight [...]
cs.AI / 72 / 2609.03483
Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps
Shuning Zhang, Liang Li, Yunheng Wang, Tao Wang, Yihang Kang, Renjing Xu
cs.RO · cs.AI
Abstract
Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a global bird's-eye view and an unmanned ground vehicle (UGV) with a local first-person view, yet the setting remains largely unexplored: existing training-free methods solve single-agent tasks but offer no collaboration mechanism, and a recent CARLA-Air evaluation found no stable cooperative behavior across five state-of-the-art VLA models; naive semantic communication or bidirectional coupling even degrades performance. We establish AGC-VLN (Air-Ground Collaborative VLN), the first training-free baseline for air-ground collaborative VLN. The key insight is that training-free methods decompose navigation into VLM-based semantic reasoning and deterministic geometric execution, exposing a collaboration interface: the UAV's global view, over which it renders the UGV's reported pose and the VLM-anchored target as CAR/GOAL markers with distance labels, yielding a shared bird's-eye map. From this map, the UGV acquires global spatial context its first-person view cannot provide, plans a road-following path with a frozen VLM, and executes it under closed-loop control; in parallel, the UAV runs 3D-SPF, a spatial-search upgrade of SPF that localizes the target in the downward view and flies toward it. On 100 closed-loop episodes in CARLA-Air's Town10HD scene, AGC-VLN reaches a 77.0% joint success rate, a collaboration gain of +27.0% over the weaker individual agent (the UAV, 50.0%), and exceeds the strongest published single-agent baseline (Travel UAV, 53.0%) by 24.0 points, stemming from the complementarity of the UAV's global view and the UGV's road-following execution. Project page: https://github.com/ZSN2024/AGC-VLN.
cs.AI / 73 / 2609.03497
BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI
Jianren Wang, Letian Qian, Zikai Wang, Weiwei Wu, Junjie Zong, Abhinav Gupta, Deepak Pathak
cs.RO · cs.AI
Abstract
Developing humanoid robots capable of leveraging human behavioral data is essential for general-purpose embodiment, yet conventional development remains bottlenecked by a decoupled paradigm that isolates hardware design from whole-body control. This approach leads to suboptimal systems that compromise human-like fluidity and agility. To bridge this gap, we introduce a data-driven morphology-control co-design framework that optimizes humanoid morphology for human-like movement. To quantify morphological fidelity, we also introduce a novel metric that jointly considers kinematic retargeting fidelity to human motion and dynamic tracking performance. Our framework achieves state-of-the-art (SOTA) performance across all metrics compared to baseline humanoids (Bumi, K1, and Toddlerbot). Finally, we realize this design in Bridge, an open-source, 88cm-tall humanoid platform released alongside its control policy. We demonstrate that Bridge captures human motion data with superior fidelity, exhibiting exceptional performance across foundational locomotion, robust balance, and highly dynamic maneuvers. Videos and open-source materials: https://sites.google.com/view/bridgerobot.
cs.AI / 74 / 2609.03565
Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning
Muyuan Liu, Yue Huang, Zheng Liang, Xiang Gao
cs.RO · cs.AI · cs.LG
Abstract
Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.
cs.AI / 75 / 2609.03611
FailBench: How Reliable are VLMs at Judging Robot Task Success?
Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
cs.RO · cs.AI
Abstract
Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets. Evaluating 13 VLM-based detectors, we find the best model achieves only 0.77 mean balanced accuracy. Notably, models fine-tuned for failure detection consistently underperform general-purpose VLMs and their own pretrained baselines. Performance depends heavily on required visual evidence: models approach saturation when outcomes depend on observable object motion, but degrade to near-chance (<0.60 balanced accuracy) on contact-intensive assembly tasks. Error analysis reveals a systematic bias toward predicting success under ambiguous evidence, which persists even with increased reasoning effort. Finally, we show that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points without extra training.
cs.AI / 76 / 2609.03889
FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation
Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong We, Qiaojun Yu, Dibo Hou
cs.RO · cs.AI
Abstract
Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cannot interpret the physical interactions induced by those actions. While the whole-body control (WBC) policy can stabilize the robot, it cannot distinguish task-relevant interaction forces from forces induced by external disturbances during manipulation. Although force/torque sensors provide direct measurements of physical interactions, retrofitting them entails additional hardware costs and substantial integration effort, particularly for platforms not designed with sensor integration in mind. To address this problem, we propose FWBC-VLA, a force-aware framework that bridges task-level VLA action generation and low-level whole-body compensation control for wheeled-legged robots. First, we introduce HSR-Force, a sensorless residual-torque estimator for inferring contact strength and its temporal variation. These contact estimates are then encoded as tokens and injected into the VLA action expert during action decoding, enabling the policy to perceive contact onset, sustained loading, and release. For loco-manipulation tasks, all parameters of the pretrained VLA backbone are fine-tuned on our WL\&Arm Dataset, which comprises more than 5,000 episodes. Moreover, the robot's proprioceptive state, the Jacobian-derived body-frame force estimate, and the estimated contact state are jointly fed into a compensation generator to produce corrective actions. The manipulation-centric actions are subsequently combined with the corrective actions and passed to the WBC policy for execution. Real-world experiments on whiteboard wiping and door opening with a door closer demonstrate the effectiveness of our FWBC-VLA in contact-rich loco-manipulation.
cs.AI / 77 / 2609.04096
Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
cs.RO · cs.AI · cs.CV
Abstract
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/
cs.AI / 78 / 2609.03622
Test-time adaptation for speech enhancement with an autoregressive speech prior
Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
cs.SD · cs.AI
Abstract
Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive prior trained on clean speech latent representations extracted from a neural audio codec. Adaptation is performed by minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments across multiple noisy speech datasets show consistent improvements in speech quality, particularly under training-testing noise mismatch conditions. Code and audio examples are available online.
cs.AI / 79 / 2609.03940
Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations
Yoto Fujita, Simon Leglaive, Laurent Girin
cs.SD · cs.AI
Abstract
Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.
cs.AI / 80 / 2609.03401
Spectral Convergence of Random Feature Method in Multiple Dimensions
Pingbing Ming, Hao Yu
math.NA · cs.AI · cs.LG · math.ST
Abstract
We first prove spectral convergence of the random feature method (RFM) for multidimensional targets in Sobolev, Gevrey, ultra-analytic, and bandlimited classes. The analysis establishes general high-probability approximation estimates in the interpolation scale generated by a kernel integral operator. On a single event determined only by the sampled features, one random space approximates every target in a prescribed source ball; moreover, for each target, a single coefficient vector defines an approximant that attains spectral accuracy simultaneously in all admissible error norms. For both regularity-adapted frequency distributions and uniform distributions on growing frequency windows, the resulting rates range from super-exponential to algebraic, depending on the regularity of the target. Second, we establish abstract error estimates for strong- and weak-form RFM discretizations, thereby converting the preceding approximation bounds into convergence estimates for multidimensional second-order elliptic boundary value and eigenvalue problems. Finally, for random feature matrices (RFMtxs), we prove super-exponential singular-value decay with Fourier features and exponential decay with $\tanh$ features, together with corresponding condition-number lower bounds. The analysis identifies a common mechanism: the same spectral approximation that yields high accuracy also drives severe ill-conditioning.
cs.AI / 81 / 2609.03697
Symmetries and Causality: Causal Effect Identification Beyond IID Data
Martin Rabel, Jakob Runge
math.ST · cs.AI
Abstract
In the natural sciences, symmetries and cause-effect relationships are ubiquitous. Yet for complex machine-learning tasks, like world-modeling in reinforcement learning, they appear difficult to harness. We propose a formal description of statistical systems based on symmetries in data leaving causal mechanisms invariant. The result is an abstract, simple and general mathematical language for causal reasoning. This paper provides formal descriptions of models and queries, setting up this language, and the formal infrastructure and strategies for their mathematically rigorous identification from data within this formalism. This approach reproduces and matches standard theoretical results on IID data and transport of experimental and non-experimental data. But its main purpose is to unify and substantially extend the scope of causal reasoning, in going beyond IID data and in approaching complex causal queries not captured by do- or soft-interventions. This new perspective on causally relevant aspects of data-modeling additionally sheds new light on well-known structures like c-components or hedges but also includes aspects of missing data and is inherently well-suited for the description of transfer and robustness properties.
机器学习 (cs.LG)
94
cs.LG / 1 / 2609.03389
Computing stable configurations of confined smectic liquid crystals with a deep variational framework
Yuchen Xie, Baoming Shi, Yucen Han, Lei Zhang
cond-mat.soft · cs.LG
Abstract
Smectic liquid crystals are layered liquid-crystalline phases characterized by orientational order and periodic density modulation. Although their structures can be modeled using continuum theories, computing stable configurations remains challenging in complex geometries, particularly when the high-frequency density modulations associated with smectic layering should be resolved. We propose a deep variational framework (DVF) for computing these configurations within the modified Landau--de Gennes model, in which the coupled orientational and positional order parameters are represented on a regular reference domain while physical confinement is incorporated through coordinate mappings. A warmup penalty mitigates the spectral bias of neural networks toward smooth, nonlayered fields, enabling robust recovery of oscillatory smectic states. Comparisons with a neural-network baseline and finite-difference relaxation demonstrate the essential role of this penalty and the numerical stability of the resulting layered states. The DVF reproduces experimentally established smectic-A defect structures and layer morphologies across diverse confinement geometries and further predicts a chevron-like smectic-C state in a tangent-anchored sphere. Together, these results demonstrate the applicability of the DVF to computing stable smectic configurations across experimentally relevant confinement geometries and anchoring conditions.
cs.LG / 2 / 2609.03193
Generative Nested Sampling of Atomistic Thermodynamic Landscapes
Alessandro Coretti, Nico Unglert, Sebastian Falkner, Georg K. H. Madsen, Christoph Dellago
cond-mat.stat-mech · cs.LG · physics.comp-ph
Abstract
Nested sampling (NS) resolves the thermodynamics of an atomistic system from a single simulation, but its practical reach is limited by the Markov-chain updates needed to decorrelate walkers within each likelihood-constrained ensemble. Flow-based NS has removed this bottleneck for gravitational-wave (GW) inference, yet its transfer to atomistic systems is not merely a change of application. Comparing a GW150914-like binary-black-hole likelihood with an eight-particle two-dimensional Lennard-Jones (LJ) system of comparable dimensionality, we show that the two landscapes differ fundamentally: atomistic multimodality is discrete and combinatorial, generated by particle permutations separated by hard collision walls, and its coordinate coupling is dense and collective, whereas the GW posterior exhibits smooth degeneracies and localized parameter coupling. Guided by this diagnosis, we introduce NS-Flows: a single conditional normalizing flow, conditioned on the NS energy bound and trained on a sliding window of recent live sets, that replaces MCMC by direct parallel draws corrected by importance-weighted rejection resampling. Live sets supply data self-consistently, allowing flow training without structured priors or a pre-existing dataset. For LJ disks in PBC, the algorithm reduces energy evaluations by over two orders of magnitude and wall-clock time by roughly one third, an advantage that becomes increasingly favorable as the cost of the potential grows. The flow's generation efficiency further acts as a physical diagnostic: it varies non-monotonically along the annealing trajectory, is lowest in the dense disordered regime, and is quantitatively captured by the constrained ensemble's internal mode complexity together with target drift across the training window, identifying liquid-like ensembles, rather than prior-target separation, as the hard case for current flow architectures.
cs.LG / 3 / 2609.04046
The Head Complexity of Boolean Functions in Single-Layer Attention
Rajmohan Rajaraman, Ravi Sundaram, Amanuel Tesfaye
cs.CC · cs.LG
Abstract
What can a single layer of self-attention compute? We study head complexity: the minimum number of attention heads required to compute a function in a one-layer attention-only model. We establish an exact hierarchy under this measure: $k$ heads compute $k$-bit parity but cannot compute $(k+1)$-bit parity. The lower bound is unconditional in the two resources a transformer might otherwise exploit; it holds at unbounded embedding dimension and unbounded numerical precision. The proof rests on an alternating-sum obstruction: after clearing the softmax denominators, every monomial in the resulting decision polynomial omits at least one of the $k+1$ input bits, forcing its correlation with parity to vanish. The same obstruction yields lower bounds for related tasks, including the well-studied multi-hop induction-head task. We also establish compactness bounds for embedding dimension and numerical precision. Specifically, a compactness theorem shows that any function computable at all can be computed with embedding dimension and precision bounded by the discrete data of the task, namely, head count, alphabet size, and length. Thus, potentially unbounded dimension or precision provably cannot substitute for heads. Finally, we derive nearly matching universal bounds for general binary functions: $2^n$ heads suffice to compute every $n$-bit binary function, with one head per monomial in its multilinear expansion, while a counting argument shows almost all such functions require $Ω(2^n/n^2)$ heads. This lower bound matches the upper bound to within a $\operatorname{poly}(n)$ factor, even when dimension and precision are unbounded. Together, these results characterize head requirements for Boolean computation in this model.
cs.LG / 4 / 2609.04011
Differentiable Hybrid Modelling for Learning and Optimising Chemical Transport Processes from Experimental Data
Arthur Jessop, Mohammed Alsubeihi, Ben Moseley, Ashwin Kumar Rajagopalan
cs.CE · cs.LG · math.NA
Abstract
Reliable transport models are essential when modelling and optimising many chemical engineering processes, yet, most models assume hand-picked constitutive laws which may not reflect reality, and often assume initial conditions are known exactly. Both restrictions can significantly bias model predictions and lead to systematic error when used in predictive and control settings. Black-box neural surrogate alternatives for modelling can better match real example data, but are confined to the task they were trained on and cannot be interrogated for physical consistency. Here we introduce a general-purpose differentiable hybrid modelling framework for transport processes, specifically for the case of population balance equations. Our framework integrates a JAX finite volume population balance solver with learnable neural network components which are trained to both discover constitutive laws and fit initial conditions from real experimental data, allowing us to better model real experimental transport systems. Furthermore, we use our framework for process optimisation, using its differentiability to allow us to direct optimising experimental settings for quantities of interest. This work highlights the huge potential of such differentiable hybrid modelling frameworks for learning and optimising any given chemical separation which involves mass, energy, and/or momentum transport.
cs.LG / 5 / 2609.03077
Position: Unlabeled IS NOT Equal to No Human Supervision in Visual Learning
Dong Lao
cs.CV · cs.LG
Abstract
This position paper argues that the absence of labels does not imply the absence of human supervision in visual learning, and urges the research community to identify sources of supervision more explicitly. Many recent methods in computer vision build upon representations learned from large-scale unlabeled data, and are therefore grouped under the same umbrella term ``unsupervised.'' However, different data curation schemes and training objectives embed substantially different human priors on which models rely, and we argue that one ``unsupervised'' umbrella term is no longer capturing these distinctions. This ambiguity makes it harder to compare unsupervised learning research conducted under different assumptions, coinciding with a sharp decline in papers titled with ``unsupervised'' in flagship computer vision conferences since 2021, despite continued growth of the field. While we fully embrace pre-training as a strong foundation for modern computer vision, we advocate for a community-level effort toward greater conceptual clarity: authors are encouraged to disclose priors in data selection and learning objectives, and to specify which components of a learning pipeline depend on which assumptions. Standardized disclosure practices can improve academic communication, ensure fairer comparisons, and preserve methodological diversity in unsupervised learning.
cs.LG / 6 / 2609.03931
Sparse auto-regressive modeling for scene generation from multi-view images
Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel, Wonjune Cho, Bardienus Pieter Duisterhof, Vincent Leroy, Jerome Revaud
cs.CV · cs.LG
Abstract
Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generative modeling is hindered by the high computational cost of dense volumetric representations and the scarcity of large-scale 3D supervision. We introduce SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision. Our key insight is to formulate 3D scene generation in a structured, compact, voxel-aligned 3D latent space where only occupied voxels are represented. We learn this sparse latent space directly from multi-view images using photometric supervision via differentiable 3D Gaussian Splatting. Given a partial set of observed voxels encoded from sparse input views, scene completion reduces to predicting the missing latent tokens and their spatial support within the voxel grid. To this end, we train a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent generation of unseen regions. We demonstrate the effectiveness of our method on synthetic indoor scenes, achieving higher novel-view quality than prior work. We further validate its generalization on RealEstate10k, highlighting its applicability to real-world data.
cs.LG / 7 / 2609.03981
Sharpening the Ensemble: An SSIM-Aligned Residual Refiner for Brain-MRI Inpainting Post-Processing
Kubilay Kağan Kömürcü, İlkay Öksüz
cs.CV · cs.LG
Abstract
Brain-MRI inpainting replaces a masked region of a scan with synthesized, anatomically plausible healthy tissue, so that analysis tools built for healthy brains can be applied to images they would otherwise reject. On the BraTS local-synthesis benchmark, which ranks submissions on the structural similarity index (SSIM), the peak signal-to-noise ratio, and the mean squared error (MSE) jointly, the strongest recent models are accurate, but several report blurry synthesized regions and attribute this to the mean-seeking behavior of the $\ell_1$ and MSE terms in their training losses. We address this in post-processing, forming a deep ensemble of the two co-first-place 2025 models and training a lightweight residual refiner on the ensemble's own outputs under an $\ell_1$ loss augmented with a structural-similarity term whose weight $λ$ we vary. At a moderate $λ$ the refiner improves SSIM over the ensemble, from $0.8767$ to $0.8780$ on a held-out reproduction of the official scorer and from $0.8555$ to $0.8572$ on the official validation leaderboard, with essentially no change in MSE. The gain is small but consistent, improving $62.6\%$ of the held-out cases with a signed-rank $p=2.2\times10^{-7}$, whereas over-weighting the structural term reverses it. Two ablations bound the effect. Adding any third model to the two-model ensemble degrades it, and classical unsharp masking fails to improve SSIM at any strength (best $0.8765$ against $0.8767$), so the gain reflects learned rather than indiscriminate sharpening. The result is a cheap, reproducible post-processing stage that improves an already strong ensemble without any large-scale retraining.
cs.LG / 8 / 2609.04168
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra
cs.DC · cs.LG · cs.PF
Abstract
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture. Para-Pipe navigates the trade-off between throughput and latency by selectively fine-tuning parallelism levels within and across pipeline stages. This strategy can significantly reduce inter-processor communication overhead, significantly improving energy efficiency. Our evaluation demonstrates that Para-Pipe generates multiple Pareto-optimal configurations, achieving a balance between throughput and latency on an Amlogic SoC equipped with ARM big.LITTLE CPUs and GPU, as well as the Black Sesame Technology SoC featuring a deep learning accelerator and two DSPs. More importantly, throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution.
cs.LG / 9 / 2609.03643
Relative Prime Factorization and Finite-State Presentations under Fixed Finite-Monoid Observation
Takayuki Kuriyama
cs.FL · cs.LG
Abstract
Let $L\subseteqΣ^*$ and fix a morphism $h:Σ^*\to M$ into a finite monoid. We study exact factorization and canonical presentation in the relative syntactic congruence $θ_{L,h}:=\equiv_L\cap\ker h$. We separate unique factorization from finite direct presentation. An exhaustively computer-checked $36$-element quotient has a unique exact prime factorization for every live non-unit class, yet its valid prime-return rules contain an infinite family, so unique factorization does not imply the finite relative presentation property (FRP), even for a finite quotient. We lift the same defect to a nonregular context-free language with an infinite relative quotient and finite prime spectrum. To isolate the obstruction, we introduce the finite-state relative presentation property (FSRP), in which canonical valid right-hand-side languages are represented by finite residual controllers, and prove $\mathrm{FRP}\subsetneq\mathrm{FSRP}$. We then introduce prime-target left-division determinism (PTLD), which implies unique exact factorization, tail exactness, tail determinism, and a quadratic bound on valid rules. A nonregular deterministic context-free example with a finite group observer satisfies PTLD while lying outside every fixed $(k,\ell)$-substitutable class. Finally, for fixed $h$ we give a strong positive-data learner for the canonical PTLD presentation with polynomial-time hypothesis updates and a finite characteristic sample, together with a limit reconstruction of the canonical FSRP controller from weakly behaviorally correct CFG-valued learners.
cs.LG / 10 / 2609.03846
EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments
Zih-Sian Yang, Yi-Hao Chen, Yu-Te Kuan, Cheng-Jui Wu, Chuang-Chieh Lin, Po-An Chen
cs.GT · cs.LG
Abstract
We study the allocation of indivisible goods among agents with identical additive valuations, focusing on envy-freeness up to one good (EF1) and Nash social welfare (NSW). Since every maximum-NSW allocation is EF1 under additive valuations, the associated threshold problem inherits the known strong NP-hardness of NSW maximization under identical additive valuations and is strongly NP-complete. We therefore focus on welfare guarantees satisfied by arbitrary EF1 allocations. Although every such allocation is known to achieve an $e^{-1/e}$-approximation to the unrestricted optimal NSW, we identify conditions yielding stronger guarantees. Under uniform valuations, every EF1 allocation is NSW-optimal. Under an $\varepsilon$-small-item condition, every EF1 allocation achieves an explicit approximation ratio $ρ_n(\varepsilon)$ satisfying $ρ_n(\varepsilon) = 1-O(\varepsilon^2)$ as $\varepsilon\to 0$ for fixed $n$. We further consider the stronger sequential requirement that EF1 be maintained after every item assignment. For this setting, we propose \emph{PriorityNet}, a deep reinforcement learning framework trained using Proximal Policy Optimization and equipped with prospective EF1 action masking. The mask restricts every decision to assignments that preserve EF1, thereby guaranteeing prefix-wise EF1 by construction without post-processing repair. Across 3,000 test instances in each of the offline and random-order online regimes ($n\in[2,20]$ and $m\in[5,100]$), PriorityNet attains mean normalized $\operatorname{NSW}$ values of $0.9911$ and $0.9701$, respectively. Relative to offline Longest Processing Time (LPT) and online least-valued-bundle baselines, it achieves instance-wise win-minus-loss rates of $+27.10\%$ and $+17.87\%$, while matching the offline baseline's mean normalized welfare to four decimal places and modestly improving the online mean from $0.9694$ to $0.9701$.
cs.LG / 11 / 2609.03901
Comparing Retrieval Methods for Academic Advisor Discovery: A Six-Method Study of 768 CS Faculty Profiles Across 9 US Universities
Biraj Subedi
cs.IR · cs.LG
Abstract
We present a comparative evaluation of six information retrieval methods for the task of academic advisor discovery: ranking CS faculty members by relevance to a graduate applicant's research interest statement. The methods span sparse lexical matching (Jaccard overlap, TF-IDF, BM25), dense semantic retrieval (all-MiniLM-L6-v2 sentence embeddings), hybrid score fusion, and learning-to-rank. Evaluation uses a new domain-specific collection: 768 faculty profiles scraped from 9 US CS departments, with 162 graded relevance judgments (grade 0/1/2) across 5 queries representing distinct graduate student research profiles. Across all five queries, Reranked achieves the highest mean NDCG@10 (0.477, std 0.138), followed by Semantic (0.450), Hybrid (0.421), BM25 (0.406), Jaccard (0.303), and TF-IDF (0.246). After Bonferroni correction across all 15 pairwise comparisons, TF-IDF is significantly worse than BM25, Semantic, Hybrid, and Reranked; no other pairwise difference survives correction at 5 queries. A field ablation reveals that biography alone (NDCG 0.634) outperforms the full model combining biography with research area tags (0.593). A controlled experiment shows that concatenating arXiv paper abstracts reduces NDCG@10 by 0.176, motivating a late-fusion architecture. All code, scrapers, and relevance labels are released openly.
cs.LG / 12 / 2609.03003
Causal Foundation Models
Christopher Stith, Hossein Rahmani, Jesse C. Cresswell
cs.LG · stat.ML
Abstract
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditionally requires a bespoke pipeline for every new problem: first proposing a causal mechanism, selecting a compatible estimator, and finally training it. Meanwhile, across diverse settings and modalities, much of machine learning has shifted to the paradigm of foundation models: networks pretrained once at scale and applied to new tasks without fine-tuning. Causal foundation models (CFMs) bring this paradigm to causal inference. CFMs are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates. This work provides a practical introduction to this emerging area. We summarize the necessary background in causal inference and machine learning before discussing CFMs. Throughout, we include example code and Jupyter notebooks.
cs.LG / 13 / 2609.03026
ObserverBench: Testing Mechanistic Estimates for Intervention and Control
Vijay Erramilli
cs.LG · cs.AI
Abstract
Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action. We present ObserverBench, a benchmark framework for testing whether an internal estimator---an observer---is adequate for the intervention, control, or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held-out cases, and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action. Theory and experiments show why both are needed. In closed-loop control, observer errors matter at the starting point and along directions the allowed intervention can reach. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions; observers trained on action loss choose lower-loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violations have different costs. Across Qwen2.5-7B, Gemma-2-9B-it, and prospectively frozen Qwen3.5-9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on the reported Qwen panels, under disclosed activation-density or checkpoint mismatches. ObserverBench provides fixed task contracts, runnable baselines, and table-based submissions for evaluating interpretability methods through the actions they enable.
cs.LG / 14 / 2609.03069
Learnable composition for neural operators
Zituo Chen, Baiming Zhang, Sili Deng
cs.LG
Abstract
Neural operators are fast, differentiable surrogates for physical simulation, but their accuracy often degrades when domain geometry, size, or operating conditions differ from training. Supervised adaptation can recover accuracy, but even a small target set requires costly high-fidelity simulations. We therefore ask how pretraining and transfer can be designed together to reduce this deployment cost. LatentDDM first pretrains a neural operator to predict fields on small subdomains. For a new setting, it freezes this operator and trains only a lightweight module that composes the local predictions. We evaluate our method on two complementary problems: steady Darcy flow, where long-range pressure coupling must extend across increasingly large porous domains, and unsteady incompressible flow around a pitching airfoil, where rollout errors compound as target pitching frequencies exceed the training range. Compared with the capacity-matched models that process the full domain at once, LatentDDM's error is 36-56% lower on larger Darcy domains after adaptation with 16 target simulations. It also improves 20-step field rollouts in fast-pitching airfoil flow, both zero-shot and after few-shot calibration. These results identify the co-designed local pretraining and composition-level transfer as a promising design principle for physical foundation models.
cs.LG / 15 / 2609.03079
LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference
Renyuan Liu, Yuyang Leng, Kaiyan Liu, Yuzhou Zhong, Shaohan Hu, Chun-Fu, Chen, Peijun Zhao, Heechul Yun, Shuochao Yao
cs.LG
Abstract
On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash storage, but face a fundamental systems trade-off: accurate sparse execution decisions require the latest context, whereas efficient computation-I/O overlap requires early prediction. As a result, existing designs either serialize execution or incur redundant weight fetches, extra computation, and large cache overheads. We present LeanStream, a streaming speculate-and-refine framework for efficient on-device LLM inference. LeanStream progressively refines computation, loading, and cache-retention priorities using partial GPU results, enabling fine-grained overlap between GPU execution and storage I/O. We implement LeanStream on both mobile and embedded platforms. Compared with prior on-device LLM inference systems, LeanStream reduces memory usage by 4.8$\times$ to 7.5$\times$ at the best throughput achieved by prior work, while further improving token generation throughput by 1.6$\times$ to 2.1$\times$.
cs.LG / 16 / 2609.03090
The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA
Samuel Larson
cs.LG
Abstract
Continuous chain-of-thought models compress reasoning into latent tokens. Matrix-valued variants, which route each latent token through a d x d matrix bottleneck, introduce rank as a single-sample structural observable on the latent matrix Z. If matrix latents carry parallel reasoning paths via superposition, rank should track them, and truncating Z to low rank should hurt accuracy on tasks whose solutions plausibly require multiple components. Across four training regimes of a matrix-CODI model (three on ProsQA, one on GSM8K-Aug below the learning threshold), the rank-k projection ablation curve is flat to within 0.6 percentage points. A three-seed replication yields 81.0 +/- 2.0 percentage points accuracy while the final effective rank of Z spans {4, 12, 13}; the loss does not reward any particular rank. To test whether rank-blindness arises from the flatten-then-project readout alone, we trained four readouts: a bilinear reparametrization, a bilinear-plus-GELU readout nonlinear in Z, an SVD-augmented readout feeding singular values through an MLP, and a quadratic readout in Z Z^T. All four rank-k curves remain flat (Spearman p-values 0.63, 0.14, 0.82, 0.46). The flat curves persist for readouts nonlinear in Z. A linear probe on Z underperforms a raw pretrained hidden state at target prediction (AUC 0.673 vs. 0.846). A negative control on vanilla GPT-2 SFT (no matrix bottleneck, no Z, three seeds, n=500) reproduces a flat rank-k curve under the same intervention paradigm with pooled-mean range 0.20pp, and a random-h sensitivity floor lands at the same accuracy: the rank-k ablation alone conflates rank-blindness with position-irrelevance.
cs.LG / 17 / 2609.03100
Distilling deep optical flow stereo methods to retrieve dense three-dimensional wind fields
Thomas J. Vandal, Dong L. Wu, James L. Carr, Derek J. Posselt, Elise Penn, Tristan Ballard, August Posch, Kate Duffy
cs.LG · physics.ao-ph
Abstract
Geostationary atmospheric motion vectors (AMVs) provide the dense horizontal wind vectors (u,v) and heights ingested into data assimilation systems. Traditional AMVs track features using window-based cross-correlation and estimate heights via infrared brightness temperatures paired with numerical weather prediction (NWP) background states, creating a circular dependency that yields inaccurate heights, high computational cost, and sparse retrievals. Stereo winds from GEO-GEO and GEO-LEO geometrically resolve heights from parallax shifts across different poses, eliminating NWP dependence and improving accuracy, but they remain computationally heavy with limited coverage. In this work, we replace window-based tracking in stereo matching with deep optical flow for efficient, improved retrieval. Fine-tuning balances a self-supervised geometric residual loss with supervised radiosonde reconstruction. To eliminate multi-satellite overlap requirements, we distill the stereo teacher into a single-satellite student model. Chi-square and height uncertainties from the teacher are emulated by the student for quality assurance. The student generates winds across full-disk GEO imagery globally. Validation compares stereo and student models against radiosondes, operational AMVs, ERA5 reanalysis, and EarthCARE cloud profiles. Results through triple collocation show that stereo winds improve performance beyond operational AMVs for water vapor bands (6.2, 6.9, and 7.3 μm), wit degradation in the long-wave infrared (11.2 μm) band.
cs.LG / 18 / 2609.03106
Scaling Laws, Tabular Data and Actuarial Ratemaking Models
Ronald Richman
cs.LG · q-fin.RM
Abstract
Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends. We investigate whether analogous scaling regularities arise in actuarial ratemaking, where data are tabular, heterogeneous, and noisy, and where classical models such as GLMs remain strong baselines. Using a real-world motor insurance portfolio, we train models from different families across increasing fractions of the training data and multiple random seeds, evaluating out-of-sample Poisson deviance, a likelihood-based loss for Poisson count predictions in which lower values indicate better held-out fit. We find that all model families improve with additional data, but scaling exponents differ substantially: TabM exhibits markedly stronger data scaling than purely supervised tabular Transformers and standard MLP baselines. Transformer variants show weak parameter scaling unless augmented with additional inductive biases (TabM-style adaptation or self-supervision). These results provide quantitative guidance on model selection by data regime and suggest that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.
cs.LG / 19 / 2609.03150
Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning
Mehreen Hossain Chowdhury, Nowshin Mahjabin, Ahmed Shafin Ruhan, Md Azam Hossain, Abu Raihan Mostofa Kamal, Md Tahmid Rahman Laskar
cs.LG · cs.CL
Abstract
Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates domain-specific updates. We test this assumption in MoE+LoRA using Python code paired with biomedical text and mathematical reasoning. Although these domains show near-disjoint expert routing, adding biomedical data substantially increases code perplexity, indicating that routing separation alone may not prevent negative transfer. To localize the failure, we introduce Jaccard routing overlap and adapter-gradient cosine similarity, which measure expert sharing and update compatibility, respectively. These diagnostics indicate that interference arises mostly from nearly orthogonal domain gradients competing within the same low-rank adapter subspace. We address this issue with SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts when adapter-level contention is detected, while keeping the router fixed. We evaluate SpawnLoRA on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and find that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA. These results demonstrate that structural separation inside experts provides benefits beyond routing or rank expansion alone.
cs.LG / 20 / 2609.03231
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
Francesco Mantegna, Gereon Elvers, Dulhan Jayalath, Gilad Landau, Tasha Kim, Miran Özdogan, Luisa Kurth, Teyun Kwon, SungJun Cho, Benjamin Ballyk, Alex Fung, Anna Greer, Pratik Somaiya, Christian Herff, Yorguin Mantilla Ramos, Hamza Abdelhedi, Karim Jerbi, Greg Farquhar, Brendan Shillingford, Mark Woolrich, Oiwi Parker Jones
cs.LG
Abstract
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phoneme classification tasks. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks (Elvers et al., 2026), highly significant advances. This success was built on the LibriBrain dataset (Özdogan et al., 2025), the largest within-subject MEG dataset recorded at the time with ${\sim}50$ hours of data for one subject. However, while within-subject scale drives strong decoding performance, a practical BCI must generalise to new users from minutes of data, not hours. The 2026 PNPL competition responds to this challenge with LibriBrain100 (Mantegna et al., 2026), an extended LibriBrain dataset with 32 additional subjects (${\sim}40$ minutes each) plus even more within-subject data (${\sim}80$ hours). Advancing the curriculum of tasks to focus on word classification, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation, progressively reducing the amount of subject-specific fine-tuning data from ${\sim}40$ to ${\sim}20$ to ${\sim}10$ minutes, a duration that falls within a clinically feasible range and brings us a step closer to a non-invasive BCI capable of restoring communication to people living with profound paralysis.
cs.LG / 21 / 2609.03239
B2B Customer Conversion Prediction: A Document Representation, Graph Theory, and CatBoost Driven Methodology
Tianqi Wang, Sheikh Shams Azam, Wan Eih Huang, Anton Wiranata, Christopher G. Brinton, Jan P. Allebach
cs.LG
Abstract
In the one-time selling B2B context, the buying cycle may last months or even years. During the long process, targeting customers that have a high potential to make purchases and recommending personalized campaigns accordingly are important for effective marketing. For this goal, we study the following problems, B2B customer data aggregation, customer feature generation, and prediction of whether a B2B customer would show interest in making a purchase (i.e., prediction of conversion into sales funnel). We propose an algorithm to aggregate individual contacts to the B2B customer level based on multiple keys. For non-standardized keys such as company names, we propose a novel architecture to cluster them in a domain encompassing irregularities such as spelling mistakes and spelling variants. We then define and generate a set of features and apply the CatBoost model for customer conversion prediction. Our framework achieves 91\% prediction accuracy. Based on the prediction results and analysis of the model, we then discuss personalized campaign recommendations to foster conversion.
cs.LG / 22 / 2609.03241
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
cs.LG · cs.AI
Abstract
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode. We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses. For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability gains, which are aggregated into a trajectory-level self-guidance score. FlowBalance calibrates this score with the verifier-derived group advantage: guidance is retained on positive-advantage trajectories, reversed on negative-advantage trajectories, and disabled when the rollout group provides no outcome preference. The resulting energy exponentially reweights a reference policy, and profiled trajectory balance fits the normalized target with one log-partition estimate per rollout group. This realizes outcome-calibrated self-guidance via trajectory balance, without a separate token-level imitation loss. Our analysis establishes within-group contrast preservation, a minimum-change reverse-KL characterization, monotonic verifier control of target reward, and an exact correction against false-positive self-guidance on rejected responses. On mathematical reasoning, FlowBalance improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.
cs.LG / 23 / 2609.03265
Selective Hypergraph Refinement for Frozen Graph Clustering
Zimo Si
cs.LG
Abstract
Existing graph-clustering methods typically improve clustering performance by optimizing model parameters and node representations. Effective means of further improving the clustering results of an already trained and frozen model, however, remain limited. We study post-processing for frozen graph clustering. After checkpoint fixation, the procedure uses no labels and updates neither model parameters, node representations, nor the original graph structure. Instead, it exploits an attribute hypergraph to supplement higher-order relations that ordinary graphs cannot readily express, thereby refining existing cluster assignments. Because global hypergraph refinement can yield both performance gains and erroneous updates, we propose Selective Hypergraph Refinement (SHR). The method generates candidate residual directions from the hypergraph and evaluates their reliability using graph structure, node attributes, and matched-null evidence. It updates only nodes with sufficient support and otherwise retains their original assignments. Further analysis shows that whether a node changes cluster is jointly governed by its native assignment gap and the directional strength of the refinement. In a controlled common-suite evaluation, 13 of 15 backbone-dataset cells had a positive mean macro gain, one produced exact no-action, and one was negative. The cell-equal macro gain was 0.066 pp (95% bootstrap CI, [0.030, 0.107] pp), while only 0.209% of hard assignments changed on average. A broader 15-combination native-interface evaluation yielded a macro gain of 0.137 pp at a mean change ratio of 0.375%. These results indicate that frozen clustering outputs retain a limited but measurable refinement space after training. The effect is heterogeneous across backbone-dataset pairs, and broader coverage also increases exposure to negative transfer.
cs.LG / 24 / 2609.03294
Latent Energy Action Planning with World Models
Phu Pham, Aniket Bera
cs.LG
Abstract
Latent world models support efficient model predictive control from high-dimensional observations, yet optimizing a single learned latent objective can favor action sequences whose decoder-predicted terminal descriptor does not match the goal descriptor. We introduce Latent Energy Action Planning (LEAP), which treats the complete action horizon as a differentiable variable and optimizes it through a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal-window state energy. Low energy requires the predicted terminal latent to agree with the goal latent and the decoder-predicted terminal descriptor to agree with the goal descriptor. A frozen goal-conditioned proposal initializes the search, a quasi-Newton solver refines actions through the autoregressive rollout, and post-optimization projection enforces the admissible action range. Across four control domains using the officially released LeWM checkpoints, the complete LEAP planning system raises mean success from 77.5% for LeWM planned with the cross-entropy method (LeWM+CEM) to 94.8% under a matched protocol, a 17.3-percentage-point improvement, while retaining the frozen LeWM representation and predictor.
cs.LG / 25 / 2609.03308
Risk and Anomaly Identification for Distribution Network Optimal Operation Based on Reinforcement Learning and Uncertainty Quantification
Ziqi Zhang
cs.LG · cs.MA · eess.SY
Abstract
Reliable operation of modern distribution networks requires timely identification of operational risks and anomalous events under pervasive uncertainty. In practice, operators must identify risks that are inherent in stochastic yet in-distribution conditions, and anomalies that correspond to out-of-distribution behaviors such as unusual load patterns, extreme weather or cyber-physical attacks. This paper addresses this joint risk and anomaly identification problem for optimal distribution network operation and proposes a deep reinforcement learning framework that is explicitly uncertainty aware. We integrate distributional and Bayesian deep reinforcement learning to realize a second- order uncertainty quantification scheme that decomposes total uncertainty into aleatoric and epistemic components, which are respectively used to characterize inherent risk and out-of- distribution anomalies. The resulting epistemic estimates drive both exploration during training and out-of-distribution detec- tion with fallback control during deployment, whereas aleatoric estimates are used to characterize intrinsic operational risk. Simulation results demonstrate the performance of our DRL agent and the effectiveness of the uncertainty quantification.
cs.LG / 26 / 2609.03337
A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines
Sompote Youwai, Chana Phutthananon, Warat Kongkitkul
cs.LG
Abstract
Every engineered fill is specified by a maximum dry density and an optimum moisture content. Each determination needs a full Proctor test. Published correlations rest on one to four hundred specimens, usually from one laboratory at one compactive energy, and are seldom released. This paper releases a corpus without those limits. It holds 2,854 laboratory compaction tests from six public sources, across 162 provenance groups and four Proctor energy levels, with fines from 1.5 to 100%. Every record is audited to the Proctor method its source names, and no energy is inferred. Screening on the zero-air-voids condition removed 11.8% of harmonised records, and 5.7% of those with a measured specific gravity. A material share of published compaction data is physically impossible. The optimum degree of saturation over the corpus is 0.815 at a coefficient of variation of 11%. That is a baseline, not a constant. Both parameters are then estimated from one classification suite and the compaction standard. A tabular foundation model reaches R2 0.824 for density and 0.784 for water content under random folds. It reaches 0.727 and 0.696 with folds drawn around provenance, and 0.520 and 0.614 with a whole source held out. Compactive energy is negligible marginally yet decisive conditionally. Density on the 66 modified-Proctor records is predicted at R2 0.740 with it and -0.651 without. Symbolic regression yields closed forms coupled through a phase relation. No predicted pair can then exceed the zero-air-voids line. The predictions are for screening, not acceptance.
cs.LG / 27 / 2609.03350
From Zero to Hero: An Open LLM Ecosystem for Armenian
Erik Arakelyan, Khatun Avetisyan, Meri Davtyan, Heghine Grigoryan, Nane Khachatryan, Hayk Shahsuvaryan, Henrik Sergoyan, Vahan Martirosyan
cs.LG · cs.CL
Abstract
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.
cs.LG / 28 / 2609.03358
Time Without Timesteps: Simulating Coupled Dynamical Systems via Self-Consistency
Liyu Zerihun, Mark Shinyoung Lee
cs.LG
Abstract
Numerical simulation of dynamical systems is usually organized as a causal march through time: each state is computed from the previous one. We explore a different formulation for coupled systems. For each subsystem type we train a neural surrogate mapping a full driving trajectory and initial condition directly to a full output trajectory; following classical waveform relaxation, coupled systems are assembled by enforcing self-consistency among these trajectories: simulation becomes a fixed-point problem over complete trajectories rather than a stepwise rollout. On coupled van der Pol oscillators and Hodgkin-Huxley neuron networks, sequential depth becomes the number of solver iterations: 4-10 Newton iterations where the reference integrator takes 1500 steps. The gradient likewise loses its time recursion: it becomes a linear system solved by GMRES at memory independent of solver depth. A single scalar measured from the learned operator, the spectral radius of its Jacobian, predicts in advance where the coupled solve will converge; past that boundary, unrolled backpropagation diverges and a Neumann adjoint fails, while the implicit gradient remains correct to 0.04%. We report where the approach succeeds and where surrogate error degrades it.
cs.LG / 29 / 2609.03377
SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign
Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel Ángel Bautista
cs.LG · q-bio.BM
Abstract
Proteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of understanding this intrinsically multi-modal relationship is crucial for fields like drug discovery and protein engineering. Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in a first stage. Secondly, a generative model is trained on the latent representation of the autoencoder(s), i.e., generative modeling in a latent space. We hypothesize that this multi-stage training is not necessary to obtain performant co-design models and thus present SimpleDesign, an effective multi-modal protein design model trained directly in the data space. SimpleDesign leverages a single-stage end-to-end objective that combines discrete cross-entropy for sequences and a regression objective for structures. In order to effectively model the difference in sequence and structure modalities, we develop a Mixture-of-Transformer architecture that allows modality-specific processing while keeping global self-attention over both modalities. We train SimpleDesign on over 2M sequence-structure pairs achieving strong performance across co-design and unconditional sequence/structure generation benchmarks.
cs.LG / 30 / 2609.03379
RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory
Yuxiang Wang, Kunyu Feng, Yingda Shen, Haoning Xu, Junyu Wang, Zhizheng Wu
cs.LG
Abstract
Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two design choices limit these gains. Each iteration sees only the previous output and cannot directly access earlier computations. Moreover, a fixed loop count wastes depth on easy inputs while leaving hard ones with too little computation. We introduce RecurTrace, which addresses both limitations using the loop's own trajectory. Specifically, Loop Memory Attention lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone. A halting head then reads the loop state and predicts whether to continue, with supervision from an oracle that identifies when additional depth still reduces loss. In a controlled MathQA comparison on the same looped backbone, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, exceeding the best fixed loop depth by 2.2 points at matched compute. By comparison, ACT and PonderNet collapse to one loop, and CALM reaches only 54.1% with 5.6 loops, while the stronger LoopUS-Conf and TaH-Mismatch baselines reach 55.3% at 3.2 loops and 55.7% at 2.1 loops. Finally, RecurTrace improves generation accuracy over same-budget fine-tuned baselines at 0.6B, 1.7B, 4B, and 8B, with the gain growing with model size from 0.6 to 3.4 points.
cs.LG / 31 / 2609.03383
TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents
Jinwei Gan
cs.LG
Abstract
Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy update, discarding transitions discovered by earlier policies and limiting advantage estimation to small, batch-local rollout groups. We propose \emph{Temporal Instance-Graph Policy Optimization} (TIGPO), which extends graph-based credit assignment across policy updates. TIGPO maintains a persistent transition graph for each task, allowing valid transitions discovered by different policy versions to jointly determine credit for current rollouts. To actively reconnect current exploration with historical experience, TIGPO allocates a fixed rollout budget between Exploration slots for ordinary task sampling and Revisit slots for delayed reattempts of previously explored tasks. For each revisit, TIGPO pairs the current rollout group with its corresponding earlier Exploration group to construct a cross-temporal reference. The enlarged reference is designed to stabilize relative advantage estimation under small rollout groups, while comparison on the same task directly captures policy improvement across training stages. Historical transitions and scores serve only as structural and detached statistical references and are never replayed in the policy loss. Experiments on ALFWorld and WebShop demonstrate that TIGPO consistently outperforms prior group-based and graph-based policy optimization methods.
cs.LG / 32 / 2609.03422
Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models
Ross Tieman, Evan Markou
cs.LG · cs.MA
Abstract
Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.
cs.LG / 33 / 2609.03427
TraveL: Transformer-based Multi-view Path Distributional Representation Learning
Fang He, Tao-yang Fu, Wang-chien Lee
cs.LG · cs.AI
Abstract
Path representation learning (PRL) for road networks has received increasing research attention, due to various path-related applications. Existing works on PRL typically exploit the co-occurrence relationship among road segments and paths to learn a vector as the path representation, without exploring the varied traveler behaviors and the regional correlation on the path. In this work, we propose to learn distributional representations, which provide valuable information for use in path-related applications, by capturing the varied traveler behaviors as well as the various dependencies within regions of road segments. We propose a novel Transformer-based Multi-view Distributional Representation Learning (TraveL) framework to encode a path along with a travel starting time to a distributional representation, which can be used to decode possible samples of on-path traveler behavior. Moreover, by analyzing the regional correlation which reveals various road segment relationships, we propose a regional attention to encode these correlations in a path. Also, we explore the idea of Kolmogorov-Smirnov (K-S) test to compare the sampled traveler behavior against the collected ground truth to facilitate training. Experimental results show that the proposed TraveL model outperforms the state-of-the-art methods on both synthetic and real-world datasets, by 14.7% in Mean K-S distance for travel time distribution estimation, 16.7% in Mean Absolute Error (MAE) for path similarity prediction, and 3.97% in MAE for destination prediction.
cs.LG / 34 / 2609.03442
Guide, Not Bind: Why Defeasible Priors Fail in Augmented Lagrangian Causal Discovery
Sairam Sundararaman, Sara Girdhar, Manit Narasimha Murthy, Samrudh N, Bhaskarjyoti Das
cs.LG
Abstract
Differentiable causal discovery methods increasingly encode expert priors as forbidden-edge constraints enforced by an Augmented Lagrangian (ALM) penalty, on the assumption that a data-adaptive relaxation mechanism will discount and eventually override a rule the data consistently contradicts. We show this design, which we call \emph{guide, not bind}, fails for two independent, precisely characterized reasons, and that directly repairing both restores it only partially. First, sequential penalty-ramping ALM suppresses a wrongly-forbidden true edge before any counterfactual check can detect it: we give three necessary conditions any adaptive relaxation must satisfy to avoid this (Proposition~\ref{prop:conditions}), prove that DADU---the natural relaxation rule this paper introduces as the object of study---violates all three (Corollary~\ref{cor:dadu_failure}), and confirm the failure across 3{,}072 training runs spanning graphs from 4 to 32 nodes, where a single wrong prior suppresses a true edge in 87--97\% of trials under DADU. Second, and independent of any fix to the mechanism, we prove in closed form that the standard correlation-matching objective ties a true edge and its reverse to an identical cost of exactly $2r^2$ (Lemma~\ref{lem:tie}), not because the underlying equal-variance model is unidentifiable, but because normalizing to correlation discards exactly the variance information that would make it identifiable; covariance matching instead separates the two directions by a provable margin of at least $w_0^4$ (Lemma~\ref{lem:separation}).
cs.LG / 35 / 2609.03443
Beyond Straightness: Non-Crossing Flow Matching via Quantile AlignTree Coupling
Junyi Lin, Mengyu Li, Jingxuan Hu, Kejun He, Cheng Meng
cs.LG
Abstract
The performance of Flow Matching largely depends on the quality of the coupling between the source and target distributions. However, independent coupling often leads to path crossings and local velocity ambiguity, while OT-based couplings typically incur high construction costs. To address this challenge, we propose Quantile AlignTree Flow Matching (QAT-FM), an efficient structured coupling strategy that constructs a hierarchical coupling between a Gaussian prior and the target data distribution via a quantile-aligned tree structure. QAT-FM constructs the coupling in $\mathcal{O}(Nd\log N)$ time and supports per-pair source sampling with $\mathcal{O}(d)$ complexity, enabling scalable training for large-scale high-dimensional generative tasks. Theoretically, we prove that the QAT coupling satisfies marginal consistency, induces non-crossing linear interpolation paths, and consistently improves path separation at intermediate times compared with independent coupling, thereby alleviating local velocity ambiguity. QAT-FM further extends naturally to conditional generation, enabling structured conditional coupling while preserving global Gaussian alignment. Experiments across diverse benchmark datasets demonstrate that QAT-FM achieves competitive generative performance while substantially reducing coupling construction cost.
cs.LG / 36 / 2609.03457
A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds
Ashir Javeed, Anton Borg, Håkan Grahn, Lars Lundberg, Dhyey Patel, Sogand Shirinbab
cs.LG
Abstract
Accurate cloud resource forecasting is essential for proactive resource provisioning, maintaining Quality of Service (QoS), and reducing operational costs in dynamic cloud environments. The existing forecasting approaches predominantly estimate future CPU workload directly from historical resource traces, which often overlook the relationship between customer service demand and subsequent resource consumption. This study proposes a two-stage integrated forecasting model that explicitly models this dependency by first forecasting customer service requests, expressed as Transactions Per Second (TPS), and subsequently estimating future CPU workload from the TPS forecast. Both the forecasting component and resource prediction component employed the XGBoost model within a cascaded learning architecture, complemented by adaptive online retraining using an expanding-window strategy to address concept drift in continuously evolving cloud workloads. The proposed work was evaluated using real-world traces collected from a private cloud environment comprising ten applications. Experimental results demonstrate robust forecasting performance by achieving Symmetric Mean Absolute Percentage Error (SMAPE) below $7\%$ for most applications, with the best-performing application achieving an MAE of $0.7372$, RMSE of $1.1866$, SMAPE of $3.57\%$, and an R2 of $0.9185$. Horizon-wise drift analysis confirmed stable recursive forecasting behavior with controlled error accumulation across a 60-step prediction horizon. Compared with the conventional direct CPU forecasting method, the proposed two-stage integrated model gives improved forecasting robustness, computational efficiency, and interpretability, making it well-suited for proactive resource management and intelligent auto-scaling in cloud computing environments.
cs.LG / 37 / 2609.03464
Mind the Gap: Robustness Risks in PII Detection Systems
Adeel Zafar, Slawomir Nowaczyk
cs.LG
Abstract
Personally Identifiable Information (PII) detection is a foundational component of data protection infrastructure where missed entities constitute direct privacy and security risks. Although modern PII systems report strong performance on standard benchmarks, we show that these evaluations mask substantial robustness failures under realistic distribution shifts encountered in deployment. Rather than comparing state-of-the-art accuracy, we study how different PII detection paradigms fail under noisy, unstructured, and informal inputs. We construct a stress test benchmark spanning seven categories of natural distribution shift and evaluate representative systems from three widely deployed architectural families: encoder-based NER (SpaCy), rule-based hybrid detection (Presidio), and generative LLM extraction (Qwen2.5-3B). All three exhibit significant degradation on out-of-distribution inputs, but with distinct and complementary failure modes. Encoder models primarily fail on unseen surface forms and boundary detection, rule-based systems fail on non-standard formats, and LLMs exhibit entity-type confusion and generation instability. These results show that aggregate benchmark scores obscure deployment-critical weaknesses and that no single architecture is uniformly reliable across PII categories. Motivated by these findings, we propose a hybrid detection pipeline with a QA-driven feedback loop for iterative risk mitigation, and release our benchmark to support OOD-aware evaluation of PII systems.
cs.LG / 38 / 2609.03495
Spectral characteristics of autoencoder parameters as a vector representation of data
Maria Nikitina, Anton Bishuk, Oleg Bakhteev
cs.LG · stat.ML
Abstract
This paper examines the relationship between the parameters of autoencoder models and the statistical properties of the data on which they are trained. Autoencoders are defined as models with an encoder-decoder architecture, trained to reconstruct input data through a compressed latent representation. It is proposed that the model parameters can be viewed as a dense vector representation of the corresponding sample. To test this hypothesis, a theoretical and experimental study is conducted in which a vector representation is formed based on the spectral characteristics of the autoencoder parameter matrices. Theoretical analysis shows that the singular values of the model parameter matrices are related to the eigenvalues of the covariance matrix of the training data, ensuring the transfer of information between the data space and the parameter space. Experimental results on the CIFAR-10 and FashionMNIST datasets confirm that the resulting vector representations allow for a high degree of accuracy in distinguishing between models trained on different data subsets, without resorting to complex vector generation algorithms or using the original samples. These results suggest that the parameters of trained autoencoders can be viewed as sample representations.
cs.LG / 39 / 2609.03504
Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails
Shi Fu, Huibo Xu, Qixin Zhang, Dacheng Tao
cs.LG
Abstract
Restricted eigenvalue (RE) bounds govern stable recovery by norm-regularized estimators. For isotropic sub-Gaussian measurements, the benchmark sample size is $1+w(A)^2$, where $w(A)$ is the Gaussian width of the normalized descent cone. The COLT 2015 open-problem note (Banerjee et al., 2015) asked whether the same law follows for heavy-tailed designs from a uniform small-ball condition alone. We give an explicit and systematic negative answer to the general question as formulated there: the proposed law fails in its full dimension-free, arbitrary-set form, and the missing obstruction is simultaneous threshold occupancy. A constant-width polyhedral descent cone with fixed small-ball constants has zero empirical RE on every sample path up to half the ambient dimension. More generally, every finite range space admits exact threshold encoding in an arbitrarily narrow spherical cap and a lift to a full polyhedral descent-cone section. For every fixed threshold VC dimension $d$, as $β\downarrow0$, the sharp worst-case sample complexity is $Θ(β^{-1}[d\log(1/β)+\log(1/δ)])$. The separation persists under exact isotropy and all finite moments: on the same constant-width cone, Gaussian measurements succeed with $O(1+\log(1/δ))$ samples, whereas an isotropic heavy-tailed design fails pathwise for $n\lesssim\sqrt{p/\log p}$. Gaussian smoothing yields an everywhere-positive $C^\infty$ density while retaining arbitrarily poor RE. Under isotropy, a distribution-free fallback governed by affine dimension times squared enclosing radius is sharp on this family.
cs.LG / 40 / 2609.03505
An Adversarial Zero-Shot Learning Approach for Anomaly Detection in Multivariate IoT Traffic Data
Mahshid Rezakhani, Tolunay Seyfi, Fatemeh Afghah
cs.LG · cs.NI
Abstract
Anomaly detection in Internet of Things (IoT) networks presents unique challenges due to the diversity of devices, lack of labeled data, and domain variability across environments. In this paper, we propose a novel framework for multivariate time-series anomaly detection that leverages adversarial learning and contrastive loss within a sequence-based Variational Autoencoder (VAE) architecture. Our method enables zero-shot domain adaptation by jointly optimizing domain-invariant latent representations and semantically structured embedding spaces, without requiring labeled data or raw feature transfer. To address the heterogeneity of IoT deployments, we introduce encoder and decoder adaptor layers that align feature distributions across domains while preserving contextual semantics. Additionally, we propose a destination-based segmentation strategy to better model real-world communication structures in IoT traffic. Our framework is comprehensively evaluated on six distinct datasets spanning industrial, enterprise, general-purpose, smart home, and military automation domains across 44 transfer scenarios. Experimental results demonstrate strong zero-shot generalization in several cross-domain settings and competitive performance against a contrastive domain-adaptation baseline under realistic, heterogeneous, and privacy-constrained IoT conditions.
cs.LG / 41 / 2609.03507
LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
Jiayi Li, Zhaomin Wu, Bingsheng He
cs.LG · cs.AI
Abstract
Tracking depression from multi-session counseling dialogues requires estimating both current symptom severity and how it changes across sessions. Yet progress on this task is constrained by the scarcity of longitudinal counseling data with standardized session-level depression labels. Existing resources typically provide either multi-session conversations without depression labels or labeled interviews in a single session. Building such a benchmark poses three challenges: maintaining longitudinal consistency and diversity, grounding symptom progression in empirical patterns, and expressing controlled depression states naturally without exposing target labels. To address these challenges, we introduce LongCounsel-8, a benchmark suite of three independently generated datasets totaling 7,749 five-session counseling trajectories, grounded in real-world client profiles, depression trajectories, symptom compositions, and counseling patterns. We combine profile-grounded simulation, empirically informed state construction, and indirect behavioral realization to address these challenges. Across the benchmark, simulated self-reports closely recover the controlled states, supporting label fidelity. Experiments on existing depression tracking methods reveal three key findings: (1) lower single-session score error does not guarantee accurate identification of trend, i.e., improvement or worsening; (2) existing methods are consistently less reliable on worsening trajectories; and (3) additional session history may reduce the accuracy of trend prediction. Together, these findings establish LongCounsel-8 as a foundation for advancing depression assessment from static, single-session prediction toward reliable longitudinal tracking of mental-health change.
cs.LG / 42 / 2609.03533
Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws
Jie Wang
cs.LG
Abstract
Existing theories derive neural scaling from data geometry or a specified data-model spectrum, but systems trained on the same data can scale differently when architecture or optimization changes the representations they can efficiently reach. We introduce Coupled Scaling, a task-conditioned framework in which finite-budget scaling depends on the relation between task structure and the geometry accessible to an architecture-optimization system. In a solvable mode-truncation model, loss separates into target energy outside architectural support and an unresolved supported tail. For an arbitrary priority order, the residual lies between the best-N supported tail and the tail beyond the largest completed high-value prefix. If the cumulative-tail and coverage log-rates are $γ_{A,T}$ and $ρ_{A,O,T}$, the residual exponent lies in $[ρ_{A,O,T}γ_{A,T},γ_{A,T}]$. Under bounded off-prefix gain, the completed prefix is rate-determining and $α_{A,O,T}=ρ_{A,O,T}γ_{A,T}$; for $a_{A,T,j}\asymp j^{-b_{A,T}}$, this gives $α_{A,O,T}=ρ_{A,O,T}(b_{A,T}-1)$. A fixed-kernel specialization derives the training-time exponent from the near-zero tail of a task-weighted spectral measure defined independently of the loss fit. The framework separates architectural support from finite-budget acquisition and motivates two tests: static task-relevant geometry should track loss at a common budget, while multiscale geometry should track coupling-specific exponent ordering, including reversal across contrasting tasks. An audit of released emergence trajectories identifies the controls needed for a direct factorial test that measures geometry separately from the scaling fit.
cs.LG / 43 / 2609.03582
WeatherNext 3: Increasing resolution and performance of global weather models with raw observations
Stephan Rasp, Boris Babenko, Dominic Masters, Andrew El-Kadi, Samier Merchant, Guy Shalev, Ilan Price, Fred Zyda, Remi Lam, Sasha Shysheya, Matthew Willson, Stratis Markou, Shreya Agrawal, Suhani Vora, Mohammed Alewi Hassen, Sunny Mak, Tom R. Andersson, Megan Bela, Akib Uddin, Nofar Peled Levi, Ben Gaiarin, Ferran Alet, Aaron Bell, Peter Battaglia, Alvaro Sanchez-Gonzalez
cs.LG
Abstract
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.
cs.LG / 44 / 2609.03603
Neural-Network Maxent: a general extension with learned nonlinearity, applied to time-series for Desert Locust distribution modelling
Alessandro Grassi, Edoardo Kimani Bellotto, Wassim El Azami, Sabrina Outmani, Maximilien Houel
cs.LG · eess.IV · physics.data-an · q-bio.PE
Abstract
Species Distribution Modelling (SDM) is essential for understanding how environmental conditions shape biodiversity, particularly for destructive pests such as the Desert Locust (Schistocerca gregaria), whose breeding dynamics are tightly coupled to rapidly evolving environmental conditions. Maxent has become the dominant method for presence-only data, but its reliance on a linear combination of hand chosen feature transforms limits its ability to capture the nonlinear, temporal relationships common in ecological monitoring, where covariates such as precipitation, soil moisture, and vegetation indices evolve meaningfully over time. Standard implementations flatten time-series covariates into independent features, discarding sequential structure that carries critical signal. We introduce RNN Maxent, an extension of the Maxent framework that replaces the fixed feature dictionary with a neural network, specifically a Gated Recurrent Unit (GRU), trained end to end via backpropagation. The approach preserves Maxent's presence only statistical foundations, background normalization, and probability calibration, differing only in that the nonlinearity is learned from data rather than fixed in advance. We apply RNN Maxent to map suitable habitat for the Desert Locust using 50 day environmental time series derived from ERA5 Land, MODIS, and Sentinel 3, maintaining a 7 day gap between covariates and presence records to yield forecasting behavior. Compared against standard Maxent, RNN Maxent improves performance across metrics (ROC AUC 0.862 std 0.036 vs. 0.792; F1 0.671 std 0.056 vs. 0.590).
cs.LG / 45 / 2609.03604
On the Interaction Between Model Compression and Test-Time Adaptation
Francesco Corti, Dong Wang, Young D. Kwon, Cecilia Mascolo, Olga Saukh
cs.LG · cs.AI
Abstract
Deep neural networks deployed in the wild must be both efficient and adaptable, requiring model compression and test-time adaptation (TTA). While both are well studied in isolation, their interaction remains poorly understood. We systematically analyze how structured compression affects a model's ability to adapt under distribution shift. Using ResNet-18 and ViT-Base on CIFAR-10-C and ImageNet-C, we evaluate multiple compression methods combined with standard TTA techniques. We introduce a diagnostic framework that examines representational expressivity and adaptation subspace compatibility. Our results reveal a consistent gap: although compressed models retain high accuracy under supervised adaptation, their TTA performance degrades significantly with increasing compression. We show that this stems from reduced representational diversity and structural constraints that limit recoverability. These effects strongly depend on the compression method, highlighting the need to design compression strategies that preserve adaptability.
cs.LG / 46 / 2609.03660
Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners
David Milec, Spyridon Samothrakis, Michael Fairbank, Dennis J. N. J. Soemers
cs.LG · cs.AI
Abstract
The dominance of Neural Networks (NNs) in RL is partially due to their incremental learning capability, which naturally suits the online, non-stationary nature of self-play training. However, gradient-boosted trees like LightGBM are widely recognised as the state of the art for tabular data in supervised learning, often outperforming NNs in accuracy and efficiency. Game states are inherently tabular---discrete actions, categorical card identities, structured board positions---which makes them an ideal candidate for tree-based methods. We introduce LUGL (Local Updates, Global Learning), a framework that decouples data collection from model fitting, enabling non-incremental learners such as GBTs to operate in RL settings where they would otherwise fail due to distributional shift. LUGL alternates between a local updates phase, where the agent plays self-play games and accumulates tabular updates (Q-values, V-values, policies, or regret values) in a finite table, and a global learning phase, where the table is used to train a function approximator that generalises to unseen states before the table is reset. We test our approach in four standard perfect-information games (Tic-tac-toe, Connect-4, Othello, and Hex) and five imperfect-information games (Kuhn's poker, Leduc Hold'em, Liar's Dice, Goofspiel, and Flop5 Hold'em), and show that our results are competitive with or superior to DQN and DeepCFR. Our experiments demonstrate that the community's strong bias towards NNs in game-playing may be unwarranted, since LightGBM-based agents achieve competitive or superior performance across all tested benchmarks.
cs.LG / 47 / 2609.03662
Extracting Forgotten Prompts from Targeted Unlearned Models
Au Ashley Hoi-Ting, Meghdad Kurmanji, William F. Shen, Nicholas D. Lane, Ligang He
cs.LG
Abstract
Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to successfully recover some of the unlearned knowledge. In this paper, we uncover a new vulnerability. Existing attacks typically assume that the forgotten prompts are already known to the adversary and focus on recovering their answers. However, we show that the forgotten prompts themselves can be extracted by using the retained data and black-box access to the model. Our attack, Targeted Active Search (TAS), first identifies the forgotten entities by constructing canonical templates and entity pool, and selectively querying the model using the most informative template-entity pair under a limited query budget. Once the entities are identified, TAS instantiates prompt templates with those entities to probe the unlearned model and reconstruct the forgotten prompts. Experiments across three unlearning methods with three datasets and three LLMs shows that TAS recovers the forgotten entity with $100\%$ accuracy and reconstructs up to $95\%$ of forgotten prompts, all while using up to $99.7\%$ fewer queries than naive probing.
cs.LG / 48 / 2609.03667
Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning
Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob, Siddarth Singh, Juan Claude Formanek, Felix Chalumeau, Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi, Louay Ben Nessir, Simon Verster Du Toit, Daniel Rajaonarivonivelomanantsoa, Asim Awad Osman, Arnol Manuel Fokam, Refiloe Shabe, Arnu Pretorius
cs.LG · cs.AI
Abstract
Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task generalisation in the offline setting and conduct an extensive empirical investigation into the scaling behaviour governing task diversity, dataset size, and network capacity. To facilitate this study, we extend offline sequence modelling architectures to handle multi-task observation and action spaces alongside variable agent counts across tasks. Our primary finding is that scaling task diversity---rather than sheer dataset size is the dominant factor in achieving robust zero-shot transfer. Through large-scale experiments across four challenging environments (Connector, RWARE, SMAX, and LBF), we demonstrate that our multi-task approach achieves a mean improvement of 3.2x on held-out test tasks compared to single-task models and consistently outperforms strong behaviour cloning baselines. These results suggest that the development of generalisable MARL agents should prioritise the diversity of the training distribution with varying numbers of agents, providing a roadmap for scaling offline MARL effectively.
cs.LG / 49 / 2609.03686
Resolution-Aware Experimental Design under Partial Identifiability
Sofianos Panagiotis Fotias
cs.LG
Abstract
Experimental design is commonly framed as choosing the experiment expected to provide the most information. Under partial identifiability however, persistent nuisance uncertainty can make the same observation carry different structural meanings. We introduce Resolution-Aware Experimental Design (RAED), which selects an experiment by the smallest expected nonempty structural candidate set achievable subject to false-exclusion control. We prove an exact cross-nuisance aliasing separation: an experiment can be preferred by structural and full-latent information gain, average classification, and nuisance-marginalized informativeness while having arbitrarily poorer valid structural resolution. RAED nevertheless preserves the expected ordering under a genuine composite Blackwell comparison. To make this criterion operational, we develop a learned score-based implementation with finite-sample nuisance-average and positive-tail calibration, and characterize a rare-tail sample-complexity obstruction. Under constrained sensing, two subsurface-flow benchmarks exhibit genuine RAED--expected-information-gain (EIG) experiment-selection disagreements, with the clearest and largest held-out resolution differences in WCA. In a fluvial benchmark, tail protection changes the selected physical experiment and replaces hard-region false exclusions primarily with explicit ambiguity. In a mechanistic methane-oxidation benchmark, a prospectively specified 5\% false-exclusion tolerance also yields a nontrivial finite-sample population guarantee for tail-sensitive nuisance risk, with 95\% joint confidence across all three structural families.
cs.LG / 50 / 2609.03705
Federated Causal Discovery via Regression-Directed Cumulants
Pablo Torrijos, Fabio Stella, José A. Gámez, José M. Puerta
cs.LG
Abstract
In this paper we study linear non-Gaussian acyclic models (LiNGAM) when used in federated environments. These causal models allow one to go beyond Markov equivalence. However, in many domains data are scarce, and increasing the sample size by centralising data from different clients is not advisable due to regulations such as the GDPR. The federated environment offers an attractive option to balance privacy and causal discovery accuracy. Unfortunately, the standard centralised estimator in the LiNGAM setting, i.e., DirectLiNGAM, cannot be straightforwardly federated. Higher-order cumulant tensors offer a way around this obstacle: they depend only on the joint distribution of the variables involved and add exactly across independent sample groups, so a single communication round suffices in horizontal, vertical, and hybrid partitions. However, FedISHC, i.e., the current federated method along these lines, breaks down under near-symmetric noise. To overcome the above limitation, we introduce the FedRCD family of causal discovery algorithms, and investigate three variants that trade off communication rounds against algebraic noise; two of them are exact federated counterparts of the centralised high-order cumulant (HC) and HC-LiNGAM algorithms, and the single-round variants further effectively support exact unlearning at any granularity, from a single observation to a whole client. Numerical experiments show that at sample sizes typical of real deployments, the entire cumulant-based federated family does not actually rank variables by the population asymmetry that the scores encode at zero. It ranks them by a variance ladder induced by the DAG along its directed paths, the cumulant counterpart of varsortability. Marginal standardisation collapses every cumulant method to near-random ordering, while scale-invariant DirectLiNGAM, not federable under this protocol, is unaffected.
cs.LG / 51 / 2609.03762
Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size
A. Afham
cs.LG · math.OC · quant-ph
Abstract
The computation of the Bures-Wasserstein (BW) barycenter of an ensemble of positive definite matrices arises throughout machine learning, optimal transport, and quantum information. Riemannian gradient descent (RGD) at unit step size -- the fixed-point iteration used in practice -- converges rapidly, yet existing analyses present a dichotomy: unit-step guarantees carry worst-case exponential dependence on the dimension, while dimension-independent guarantees require small step sizes that forfeit the empirical speed. We resolve this dichotomy, not by improving the guarantees for unit-step RGD, but by proposing a Projected RGD algorithm that achieves dimension-independent linear convergence at unit step size. The achieved rate, $(1 - κ^{-3/2})$, where $κ$ is the condition number of the ensemble, also polynomially improves on the best small-step guarantee ($κ^{3/2}$ versus $κ^{5/2}$ iteration complexity). The crux is a novel Projection Lemma: clipping the eigenvalues of a positive matrix to an interval $[α, β]$ is the closed-form, non-expansive (1-Lipschitz) BW-metric projection onto the set $\{S : αI \leq S \leq βI\}$ -- a statement which, unlike its known one-sided counterpart, does not follow from convexity. The projection is moreover free: it reuses an eigendecomposition the next iteration must perform in any case, so the projected and unprojected iterations cost the same per step. The same analysis covers the invariant matrix projection problem of Brahmachari et al. (2025), whose fixed-point algorithm we identify as unit-step RGD on a totally geodesic submanifold, thereby extending the dimension-independent guarantee to that setting verbatim.
cs.LG / 52 / 2609.03763
From Nowcasting to Forecasting: Adapting a Reanalysis-Trained
Mikko Partio, Leila Hieta, Ossi Laine
cs.LG · physics.ao-ph
Abstract
Accurate cloud-cover forecasts are important for temperature prediction, radiation forecasting, and solar-power operations. Short-range forecasting methods can preserve observed cloud placement during the first forecast hours, but their skill decreases when cloud fields evolve through formation, dissipation and deformation. Longer lead times require accounting for atmospheric evolution, but operational numerical weather prediction (NWP) forecasts may not accurately represent the satellite-observed cloud state at initialization. We develop CloudCast v2, a machine-learning model for 12-hour cloud-cover forecasting from observation-based initial conditions. The model is first trained on the Copernicus European Regional Reanalysis (Ridal2024) to learn cloud-evolution dynamics, and is then adapted to satellite-derived cloud fields using conditional flow matching (Lipman2023), a generative method that transforms noise into cloud-cover forecasts conditioned on the observed initial cloud fields and NWP inputs. CloudCast v2 reduces mean absolute error by 10% relative to its predecessor, CloudCast v1 (Partio2025), over the 1-12 h range. It also overtakes CloudCast v1 in fractions skill score, a neighborhood-based measure of spatial agreement, after approximately 3-6 h, depending on the cloudiness category. These results show that observation-initialized machine-learning forecasts can extend beyond the usual 1-3-hour nowcasting range while retaining spatial detail from satellite cloud fields.
cs.LG / 53 / 2609.03770
OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education
Elakkiya Rajasekar
cs.LG · cs.CL · cs.CY
Abstract
Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews of curriculum analytics report an absence of evidence on how that computation informs decisions. This paper presents OBER+, an extension of a deployed institutional attainment platform that computes the step from a measured shortfall to an evaluated corrective action. Five connected stages accumulate attainment across deliveries of a course, signal a shortfall and a persistent shortfall, grade it on cutoffs the regulator already uses, record the decision against a catalogue of practices annotated with their evidence, log the change, and quantify the subsequent movement in the shortfall. A further rule compares successive statements of an outcome, so attainment is never read as a series across a point at which the outcome changed. Applying the rules to the live record of two real courses produced three results. Every outcome of a core course was substantively redefined between consecutive deliveries, with subject matter moving between outcome numbers, so a naive reading would have reported a twenty-five point collapse between quantities that do not refer to the same learning. Recomputing the platform's figures from its documented rule showed six of ten differing by more than rounding explains, in a pattern that identified a defect since reported to the institution. Across fifteen statement pairs from three transitions, five were identical character for character, and among the ten that were not, the outcome carrying a given number was nearest to a differently numbered earlier outcome in six, a result resting on an ordering of similarities and requiring no threshold and no labelling. The contribution is a computational design for outcome-based reporting, stated as rules any attainment platform can implement, with evidence of what they make visible in a live institutional record.
cs.LG / 54 / 2609.03790
Landmark-Based Discrimination of Injury-Associated Athlete-Sessions from Minute-Resolution Multimodal Football Monitoring Data
Evangelos Chatzidimitriou, Konstantinos Tserpes
cs.LG
Abstract
Athlete monitoring data may be recorded minute by minute throughout a match or training session, while injury information may only indicate whether the entire session was injury-associated. This creates a modelling problem: assigning the same session-level label to every minute would imply that injury status is known at each exact time, even though within-session injury onset is unknown. Our novelty is a fixed-landmark, one-representation-per-athlete-session formulation that directly addresses this mismatch. Instead of labelling every minute, we construct one representation per athlete-session at each landmark using information observed up to that point. This keeps the target at the session level and avoids unsupported minute-level injury supervision. A landmark is a fixed time point within the same session, such as 10, 20, or 30 minutes. At each landmark, we assess whether the whole session is injury-associated or non-injury-associated and examine how discrimination changes as more within-session information becomes available. Using 2020 SoccerMon data, we analyse 3,743 athlete-sessions from 48 elite women's football athletes, including 22 injury-associated sessions from five athletes. We evaluate pre-session, cumulative, dynamic, and combined representations with athlete-disjoint validation, athlete-cluster bootstrap uncertainty, common-cohort sensitivity analysis, alternative negative-athlete fold allocations, equal-athlete weighting, and Logistic Regression, Random Forest, and XGBoost benchmarks. Primary CUM+DYN Logistic Regression yields ROC-AUC 0.367-0.607 and PR-AUC 0.0080-0.0150 across landmarks, with wide uncertainty. PRE-containing representations show higher point estimates at several landmarks but remain uncertain.
cs.LG / 55 / 2609.03801
From Ordered Bernoulli Levels to Critical-Line Geometry: Integer Quantization, Bernoulli Residual Phase, and Prime-Power Spectra
Y. Kenan Yılmaz
cs.LG · math.NT
Abstract
We study the ordered Bernoulli-word kernel f(p,n,k)=p^k(1-p)^(n-k) and the geometry generated by its inverse-integer level sets. The binary level 2^(-n) selects p=1/2 as the unique real split-independent anchor. Under complement-preserving complex continuation, the pair becomes z=1/2+iu and 1-z=1/2-iu, producing a conjugation-symmetric vertical geometry before any zeta-function input is introduced. The quadratic coordinate Q(z)=z(1-z)=1/4+u^2 has a sharp minimum at the central point and admits an exact integer quantization. For critical-line zero ordinates gamma_k, the induced levels L_k=1/4+gamma_k^2 are decomposed exactly as L_k=N_k+delta_k, where N_k is the nearest integer and delta_k is a periodic first-Bernoulli residual. Circularization gives Z_k=exp(2 pi i delta_k), isolating gamma_k^2 mod 1 as the residual phase variable. Unique factorization resolves the integer shells into prime-generator coordinates, while a distinct complex exponent s lifts the same construction to the Dirichlet atoms m^(-s), linking the Dirichlet-series and Euler-product assemblies. Exact identities, classical zeta connections, numerical controls, and open conditional Weyl tests are kept explicitly separate. No proof of the Riemann Hypothesis is claimed.
cs.LG / 56 / 2609.03807
Free Pause Tokens
John Langford, Nathan Godey, Giovanni Monea, Yoav Artzi, Harry Dong, Ying Fan, Gustavo de Rosa, Zheng Zhan
cs.LG · cs.AI
Abstract
A free pause token gives a language model extra compute to form each next-token prediction (as a pause, or thinking, token does) but carries that compute in a parallel prediction stream over a weight-shared backbone rather than as an extra token in the sequence. It improves next-token prediction by 2-3 centinats in practice on a 1B parameter model. Because the pause rides an existing position instead of adding one, it is free to use: at inference it adds no context length, no KV cache, and essentially no latency with the growth in inference flops typically irrelevant as it is not the active bottleneck on throughput. The only primary cost is in training, where additional training compute versus an optimized pretraining pipeline is reduced to as low as x1.14 while preserving most of the benefits. The result is an isoflop, isoparameter, and isotoken improvement over standard next token trained transformers.
cs.LG / 57 / 2609.03809
A Peer-Relative Representation Learning Framework for Energy Inefficiency Identification in Mobile Network Sites
Eliud Nyakweba Koto, Jaco du Toit, Adham Stoltz, Johan du Preez
cs.LG
Abstract
Energy consumption is one of the largest operational expenditure items for mobile network operators, yet site-level energy inefficiencies such as faulty cooling controllers, idle radio equipment, and parasitic auxiliary loads often remain undetected because no ground-truth inefficiency labels exist and historical measurements may already contain embedded inefficiencies. This study proposes an unsupervised peer-relative approach based on the premise that sites with similar structural and operational characteristics should exhibit comparable energy consumption. To capture these relationships, a novel energy-aware Minimum Distortion Embedding (MDE) formulation is introduced that extends the standard MDE objective with an energy-based repulsion mechanism. This encourages sites with anomalously high energy consumption relative to comparable peers to become displaced from their local neighbourhoods in the embedding space. The resulting low-dimensional representation simultaneously preserves structural similarity and encodes energy-related deviations, enabling the identification of potentially inefficient sites through peer-relative comparison. The derived anomaly scores provide a practical mechanism for prioritising field investigations, allowing mobile network operators to focus engineering resources on sites most likely to yield energy savings. Experimental results demonstrate that the proposed approach outperforms conventional anomaly detection baselines and provides a robust foundation for large-scale energy-efficiency optimisation in mobile networks.
cs.LG / 58 / 2609.03826
Witnesses Explain Anomalies
Lamine Diop
cs.LG · cs.AI
Abstract
Unsupervised anomaly detection scores each point of an unlabelled, contaminated sample in a single pass, and increasingly must also explain why a point is flagged. Yet the dominant detectors give a score with no account of which features drive it, and explanations are bolted on post-hoc with SHAP or LIME, which re-query the detector thousands of times per point and only approximate it. We introduce WAND, an unsupervised tabular anomaly detector that is explainable by design. WAND organises its computation around directions on the unit sphere, scoring each point by how far its projection escapes a sub-Gaussian extreme-value baseline. The originality of our approach is that the witness directions that flag a point, being vectors in feature space, are its explanation, a per-feature attribution obtained at no cost over scoring and, since the score is differentiable, recoverable by gradients. Scoring is linear in the sample size, and a probe-efficiency bound guarantees every anomaly a witness, hence an explanation. Across 47 ADBench datasets WAND attains the best mean Friedman rank at ROC-AUC parity with 16 unsupervised baselines, so the gain is interpretability at no accuracy cost; its native explanations are more accurate and faithful than post-hoc SHAP/LIME and ECOD at a fraction of the query cost. WAND is thus a practical, interpretable solution for explainable anomaly detection.
cs.LG / 59 / 2609.03842
Multi-step Proximal Policy Improvement in Offline Reinforcement Learning
Soohyun Choi, Seonvin Cho, Songnam Hong
cs.LG
Abstract
Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meaningful gains often require moving beyond the behavior distribution. We develop a geometric view of offline actor updates by modeling policies as a probability manifold endowed with a chosen metric geometry. Under this lens, a broad class of offline actor objectives can be interpreted as a single proximal policy improvement step (SPI), i.e., an implicit discretization of a manifold gradient flow induced by a critic-defined energy. Building on this insight, we propose multi-step proximal policy improvement (MPI), a plug-in refinement mechanism that composes sequential re-centered proximal steps. MPI enables controlled policy improvement beyond dataset support while retaining proximal control at each refinement. The framework accommodates multiple policy geometries and admits practical instantiations for deterministic and diagonal-Gaussian policies. Experiments on D4RL benchmarks show that small numbers of MPI refinements improve strong offline baselines, including TD3+BC, ReBRAC, and IQL, on many tasks. Focused diagnostics further distinguish re-centered refinement from fixed-objective update scheduling and characterize limitations under critic error.
cs.LG / 60 / 2609.03851
Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning
Michael Khavkin, Kichang Lee, Jaeho Jin, JeongGil Ko, Eran Toch
cs.LG
Abstract
Federated Learning (FL) with Differential Privacy (DP) is increasingly adopted to preserve data confidentiality in distributed machine learning. However, DP noise distorts learned representations and degrades explanation fidelity, limiting differentially private FL where trustworthy explanations are required, such as assistive clinical diagnosis. Prior work adapted DP noise with static feature-importance signals, restricting explainability to post hoc analysis and precluding noise calibration to explanation quality during training. We propose XCal-FL, a closed-loop, explainability-driven local training algorithm for image classification in cross-silo FL that dynamically calibrates DP noise from three complementary signals: (1) prediction logit variations, measuring causal influence on model confidence, (2) counterfactual margins, capturing decision-boundary sensitivity, and (3) saliency concentration, quantifying spatial coherence of model attention, while enforcing formal DP guarantees via adaptive privacy accounting. Experiments on three medical imaging datasets across varying FL configurations show that XCal-FL yields more accurate and interpretable global models, improving predictive performance by over 10\% and explanation fidelity by up to 5$\times$ over static-noise FL, and outperforming state-of-the-art adaptive DP methods in fidelity. XCal-FL also achieves higher privacy-budget efficiency, turning each unit of cumulative privacy loss into larger gains in both accuracy and explanation fidelity. Our analysis further reveals that, unlike predictive performance, which scales roughly linearly with privacy loss, explanation fidelity exhibits non-linear dynamics. These findings suggest explainability is a distinct dimension of the privacy trade-off that cannot be inferred from utility alone, with implications for training and privacy-budget allocation in decision-critical applications.
cs.LG / 61 / 2609.03858
High-Dimensional Learning Dynamics of Attention-Indexed Models
Yizhou Xu, Margarita Sagitova, Lenka Zdeborová, Florent Krzakala
cs.LG · stat.ML
Abstract
Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix $S\in\mathbb{R}^{d\times d}$ can remain trapped in an uninformative state. Tied attention ($S=WW^\top$) induces an automatic symmetry-breaking mechanism and yields weak recovery in $Θ(d^2\log d)$ samples. For untied attention, $S=UV^\top$, we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the $Θ(d^2\log d)$ scale occurs when the state selected by the fast dynamics breaks the initial symmetry.
cs.LG / 62 / 2609.03878
Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data
Lamine Diop, Marc Plantevit
cs.LG · cs.AI
Abstract
Reconstruction-based anomaly detectors are accurate but opaque: a deep autoencoder flags a sample without telling a practitioner which feature ranges made it anomalous. We propose DIFFINT, an autoencoder whose latent bottleneck is structured as a set of soft, axis-aligned interval memberships learned end-to-end directly from raw numerical data, without any discretization or binarization. Each latent unit corresponds to a human-readable hyper-rectangle in feature space; an instance is encoded by how strongly it falls inside each interval relative to the other units, and its reconstruction error is the anomaly score. This keeps the power of differentiable representation learning while exposing an inspectable internal structure. We make the inductive bias precise: a certified reconstruction-error lower bound for points that fall outside every active coordinate of the learned support (with a Lipschitz-enforced decoder), and a graded, empirically verified suppression mechanism for the usual case in which only a few features are abnormal; and we provide a closed-form, label-free importance that ranks each (unit, feature) pair from quantities the model already maintains, turning trained intervals into auditable candidate constraints without ever seeing an anomaly label. On 48 ADBench benchmarks against 22 baselines under a common [-1, 1]-normalized protocol, DIFFINT attains the best mean rank overall on both metrics (4.10 on ROC-AUC, 4.16 on AUPR); among inlier-only detectors it leads its regime clearly, and it is competitive with the strongest contaminated-data detectors (see the stratified and complete-case analyses). It is the only interpretable detector in the statistically-tied leading cluster of seven methods.
cs.LG / 63 / 2609.03900
Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating
Heejin Choi
cs.LG
Abstract
Continual knowledge-updating methods are often declared superior from one final checkpoint and one conventional adapter rank. We show that this can be insufficient to identify the better operating point. Holding a periodic hierarchy fixed, we compare it with cumulative replay over a 24-month Wikidata stream while varying evaluation month, replay LoRA rank, and query formulation. The apparent winner changes across this region: on Qwen2.5-1.5B, the hierarchy's 5.0-point advantage over rank-8 replay becomes an 11.6-point deficit against rank-72 replay, and at high ranks a consolidation-aligned endpoint can suggest a tie while time-averaged replay leads by 9-13 points. The same rank-conditioned reversal appears on Llama-3.2-1B and held-out paraphrases. These results show that method ranking in continual updating can depend jointly on when performance is measured and how much replay-side adaptation capacity the baseline receives. We therefore propose reporting trajectories and capacity sweeps, and declaring a robust winner only when the ordering is stable across the evaluation region; otherwise, comparisons should report winner regions and retention-stability-cost frontiers. Under this protocol, the periodic hierarchy is a lower-update-cost operating point, not a quality winner.
cs.LG / 64 / 2609.03937
RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting
Yuchen He, Yueyang Cang, Zhiyuan Ning, Ningyu Wang, Li Shi
cs.LG · cs.AI
Abstract
Retrieval-augmented generation (RAG) complements parametric models with retrieved external evidence. The same idea is attractive for continuous-output regression, but directly reusing retrieved target values is often not robust when samples differ in output level, numerical scale, or local dynamics. Moreover, conventional forecasting pipelines generally use residuals for model optimization and error diagnosis, but do not retain individual historical residual examples as memory that can be accessed at inference time.For multivariate time-series forecasting, we propose RATL, a plug-in residual-retrieval and feedback-correction method. RATL freezes a base forecaster to construct retrieval keys and turns its historical forecast residuals into a train-only memory specific to that base model. At inference time, RATL retrieves residual trajectories from similar historical contexts subject to causal availability constraints, then uses a set-aware router operating over forecast blocks and variables to select and combine these trajectories. Experiments show that historical residuals matched to the current context contain reusable forecasting information and that RATL improves frozen base forecasters in most experimental settings. Ablations further show that learned routing strengthens raw residual feedback, while validation-based correction-strength selection limits residual over-injection.On real-world benchmarks, we use iTransformer as the primary frozen base forecaster, compare against multiple strong forecasting baselines, and test transferability across backbones. The results show that RATL can further improve base-forecaster performance in most settings.Overall, RATL shifts the retrieved object from historical target values to base-model-specific historical forecast errors, providing a plug-in, residual-memory-based paradigm for learned feedback correction in continuous-output forecasting.
cs.LG / 65 / 2609.03941
Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
Hyun Bin Park, Du-Seong Chang
cs.LG · cs.AI · cs.CL
Abstract
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates them by compatibility with the current policy. The fresh on-policy stream remains unchanged, and the method adds no auxiliary generation or training machinery. Across mathematical reasoning, multimodal reasoning, and Agentic Search benchmarks, this single intervention outperforms naive replay and matches or exceeds broader replay methods on Avg Mean@32. In Agentic Search, where environment interaction dominates cost, it delivers comparable quality at materially lower wall-clock time.
cs.LG / 66 / 2609.03949
VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch
WenJie Fan
cs.LG · cs.CL
Abstract
The problem. A long-lived KV cache must be compressed before the queries that will read it exist; selection by observed attention (H2O, SnapKV) collapses there (0.00-0.33 needle retrieval on a NoPE MLA model), because a token's importance has not yet been observed. The method. On Kimi Linear, VestigeKV evicts by a query-independent signal the cache already carries: the 64-dimensional decoupled branch, a vestige of RoPE that NoPE training repurposes into a salience channel. Reading 11% of each row, it partitions the cache: the top-m rows stay in the attended tier; every other row moves -- exactly, never deleted -- to a GPU-resident archive reachable per step by a certified trigger. No training, no quantization, no weight or kernel change. Cost. Nothing measurable: retrieval holds at 1.00 under 8x and 0.92 under 32x from 8k to 65k context, zero gap to full-row selection. The attended tier is 0.25 KB of Kimi Linear's 8.1 KB per-token cache at 32x; the archive stays bit-exact and GPU-resident, with host offload as the VRAM-reclaiming variant. The recall tier -- the standard configuration -- holds 128x at 1.00. Kimi K3 is reported to use a NoPE Gated-MLA variant; if its cache layout matches, the method plausibly extends there -- we make no claim beyond the measured model. NoPE exclusivity. The identical operator on a RoPE MLA collapses to 0.08 (plain eviction: 0.42); query-independent salience itself exists only without rotation (top-1 targets span 2.3-6.7% of tokens vs. 10.2-46.8%), and query-universal exact merging is provably impossible under RoPE. All thresholds were frozen before data; 20 archived verdicts and 8 closed routes accompany the paper.
cs.LG / 67 / 2609.03972
OSR: Output Space Redistribution for Adaptive Label Removal in Classification Models
Minyi Peng, Darian Gunamardi, Ivan Tjuawinata, Yongsen Zheng, Kwok-Yan Lam
cs.LG
Abstract
Label removal occurs frequently in classification systems with evolving taxonomies, where categories must be dynamically updated or eliminated. To accommodate such changes, classification models must adapt accordingly. Existing solutions, broadly categorized as retraining-based and feature-space-adjustment-based, share common limitations despite their variations, including reliance on access to original data, substantial computational and storage costs, inconsistent results, poor scalability, and degradation of model utility. To address this, we propose a novel approach that leverages statistical redistribution in the output space to approximate the post-removal confidence vectors of a retrained model. Applicable as a modular output filter, our method bypasses the burden of feature-space adjustments or loss-function convergence, alleviating scalability limitations. Furthermore, by requiring only existing labels and prior output confidences, the method potentially mitigates privacy concerns inherent to data-dependent solutions. Extensive experiments demonstrate competitive performance against full retraining, with improvements in computational efficiency and privacy preservation across several classification tasks.
cs.LG / 68 / 2609.04007
RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models
Mohammad Mohammadi, Alireza Zarei
cs.LG · eess.SP
Abstract
Despite strong performance on held-out electroencephalography (EEG) data, seizure detectors may fail under real-world acquisition variability, artifacts, and adversarial inputs. We introduce RobustSeiz, an open-source, model-agnostic framework that provides a standardized, reproducible protocol for stress-testing and comparing seizure detectors under controlled, clinically motivated distribution shifts before deployment. We standardize four public scalp-EEG corpora (CHB-MIT, TUSZ, Siena, and SeizeIT1) into BIDS-EEG trees and evaluate subject-independent detectors on held-out splits. Environment, noise, and adversarial transforms are swept over predefined hyperparameter grids. Each run reports sample- and event-level sensitivity, precision, F1, false positives per 24 h, Lead and Lag onset timing, and Monte Carlo dropout predictive agreement. RobustSeiz includes a Dockerized GPU pipeline, experiment registry, and full-evaluation and research-subset modes. We demonstrate the framework with a contemporary seizure detector on TUSZ across the complete implemented shift grid; an AWGN analysis illustrates how perturbation severity changes detection quality, onset timing, and predictive agreement. RobustSeiz provides a shared benchmarking standard for evaluating seizure-detector robustness under realistic clinical stressors, extending pre-deployment assessment beyond clean-data accuracy.
cs.LG / 69 / 2609.04018
A location-invariant estimator of extremal quantile treatment effects for heavy-tailed distributions
Xin Yu, Shuwei Huang, Jicheng Liu, Jielin Tang, Bolin Wang, Yunxiao Zhang, Tian Zhao
cs.LG · stat.AP · stat.ME
Abstract
Quantile treatment effects (QTEs) measure the effect of a treatment on the distribution of an outcome, and their estimation at extreme quantile levels is of central interest in applications where the target quantiles lie far beyond the range of the data. For heavy-tailed potential outcomes, existing extremal QTE estimators rely on extrapolation combined with a causal extreme value index (EVI) estimator, but the resulting estimator is not invariant under a common location shift of the potential outcome distributions, even though the population QTE is. We address this issue in two steps. First, we adapt the location-invariant Fraga estimator of the EVI to the causal setting using inverse propensity score weighting. Second, we replace the original extrapolation formula with a difference-based scheme, under which the location parameter cancels when quantile differences are taken. The resulting QTE estimator is therefore location invariant. We establish the consistency and asymptotic normality of the proposed extremal QTE estimators, and provide a consistent variance estimator, leading to asymptotically valid inference. A simulation study confirms the location invariance, the stability with respect to the threshold, and the coverage of the proposed methods.
cs.LG / 70 / 2609.04066
Subspace Inference Enables Efficient Active Reward Learning from Preferences
Yutai Zhou, Erdem Bıyık
cs.LG · cs.AI · cs.RO
Abstract
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preference queries. However, effective uncertainty quantification required for active learning remains a key challenge for large neural network reward models. In this paper, we introduce PreferenceEKF, a sample-efficient approach that tracks reward model uncertainty by framing active preference learning as a sequential Bayesian filtering problem. Instead of relying on computationally prohibitive posterior inference over the full neural network parameter space, our method performs sequential inference via an extended Kalman filter within a low-dimensional parameter subspace, continuously updating the reward model posterior as new preference queries arrive. Our approach enables scalable sampling of neural network parameters to efficiently compute acquisition functions for active reward learning. Experiments on the D4RL and V-D4RL benchmarks demonstrate that our approach achieves better sample efficiency, runtime, scalability, and calibration compared to other Bayesian deep learning approaches, and the learned reward models lead to competitive offline reinforcement learning policy performance. This highlights the potential of scalable Bayesian methods for preference-based reward modeling in RLHF. Our code is available at https://github.com/yutaizhou/bnn_pref.
cs.LG / 71 / 2609.04105
Hardware-Aware FP4 FlashAttention-4
Robert Hu
cs.LG
Abstract
Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink. We address this with \emph{Direct-P} for noncausal inference and a causal path that passes the forward quantization directly into backward. Direct-P maps scores directly to FP4 probabilities and reaches up to 2.13$\times$ the bfloat16 (BF16) forward throughput on an NVIDIA GB200. The causal path reconstructs probabilities from saved quantized queries and keys and uses 8-bit floating-point (FP8) gradient operands, accelerating a complete single-GPU 8-billion-parameter update by up to 1.14$\times$. Matched distributed training retains FP8 probabilities and values; every tested MXFP4 probability/value training trajectory diverges.
cs.LG / 72 / 2609.04113
Constant regret in general games via higher-order optimism
Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos
cs.LG · cs.GT
Abstract
We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting (HOOD) is a variant of optimistic follow-the-regularized-leader (OptFTRL) that combines a discounted $(N+1)$-th order predictor with entropic regularization over a suitable "lifting" of the game's strategy space. This combination of ingredients is purposefully designed to dampen large oscillations of the induced sequence of play in a controlled manner, removing in this way a key stumbling block of previous attempts to achieve constant regret in general games. Our approach bears several striking similarities to the concurrent - and completely independent - work of Liu, Farina, and Ozdaglar (arXiv:2608.31166), who very recently derived an $O(N^{21}\log^{4} K)$ regret bound through the use of higher-order optimism and an exponential moving average estimator.
cs.LG / 73 / 2609.04134
Prospective Coding Improves Learning in Deep Continuous-Time Recurrent Networks
Shivang Rawat, Mirko Morello, Flaviano Morone, David J. Heeger
cs.LG · cs.NE · q-bio.NC
Abstract
Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom-up signals and attenuates top-down errors. We develop Recursive Quadrature Filters (RQFs), biologically motivated complex-valued temporal filters that are a special case of diagonal state-space models (SSMs), and ask whether this failure mode can be addressed by making each layer's bottom-up input prospective. Starting from an energy model, we derive the RQF dynamics and show that each RQF is a band-pass filter whose learnable parameters control its tuning frequency and bandwidth. We then make each layer's bottom-up input prospective using a parameter-free two-tap update that leaves the recurrent transition and parallel scan unchanged. We extend this correction to general diagonal SSMs and show that it mitigates depth-dependent gradient attenuation when temporal gradients are truncated, i.e., spatial-only backpropagation. We evaluate the intervention in RQFs, S5, and ORGaNICs (a nonlinear gated RNN) trained using full backpropagation through time (BPTT) and spatial-only backpropagation. Under full BPTT, prospective variants match or outperform their non-prospective controls in every model and configuration. A non-residual width-32 six-layer RQF reaches 96.09% accuracy on raw-audio Speech Commands with 31.9k parameters; a width-64 six-layer RQF reaches 83.56% on the 16,384-step Path-X task. These results identify RQFs as a parameter-efficient recurrent substrate and prospective-input coding as an input-side correction for deep continuous-time recurrent networks.
cs.LG / 74 / 2609.04147
A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle
Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer
cs.LG · cs.AI · cs.RO
Abstract
This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniature Ackermann vehicles. The platform combines a physical vehicle, a printed urban track, data collection tools, trajectory registration, and a Webots digital twin, enabling controlled experiments that connect simulation-based autonomous-driving methods to real-world execution. As a first baseline, we implement command-conditioned behavior cloning, in which a neural policy receives an on-board camera image and a high-level navigation command and outputs steering and speed. The system is evaluated both on the physical vehicle and in simulation. In real closed-loop experiments, the learned policy follows lanes and executes commanded turns, reaching a mean cross-track error of 6.1 cm with respect to the reference route, close to the 4.7 cm observed in human demonstrations. In the digital twin, camera field of view has a strong effect on performance, reducing the mean cross-track error from 35.6 to 3.3 cm when widened from 58 to 120 degrees. Using the digital twin to generate synthetic driving data and a learned sim-to-real image translator to reduce the appearance gap, we further show that a higher-capacity policy trained on this synthetic data combined with real demonstrations is the only configuration that completes all four track routes in closed loop, whereas the compact baseline and the same network trained on real data alone complete fewer. These results establish the open platform as a practical testbed for sim-to-real studies and provide an initial command-conditioned imitation-learning baseline; we release it to support reproducible research.
cs.LG / 75 / 2609.04189
Robust PAC Learning of Concurrent Stochastic Games
Angel Y. He, David Parker
cs.LG · cs.GT · cs.LO · cs.MA
Abstract
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the challenge of Nash equilibrium (NE) existence. Our algorithm maintains data-driven $L^1$ confidence sets over transition kernels and solves a robust CSG to compute a social-welfare optimal $\varepsilon$-NE, using a robust MDP-based exploration mechanism to drive joint state-action coverage. Crucially, we introduce a Nash margin characterisation that enables principled reasoning about equilibrium existence: the framework either returns an $\varepsilon$-approximate NE whose social-welfare value is $\varepsilon$-close to optimal, or provides a sound certificate that no exact NE exists. Under a minimum reachability condition $p_{\mathrm{reach}}>0$ over relevant state-action pairs, the algorithm terminates after a polynomial number of trajectory samples, with sample complexity $\widetilde{O}\left( {R_{\max}^2 H^4 |S|^2 |A| / (p_{\mathrm{reach}} \varepsilon^2)} \right)$. Empirical results on benchmark CSGs demonstrate near-optimal performance, correct handling of equilibrium (non-)existence, and sample complexity consistent with theory.
cs.LG / 76 / 2609.03147
Feasible but Not Safe: Constraint Violations and Report-Channel Attacks in Learned Cell-Free ISAC Association
Mehdi Zafari, Iman Mohammadi, A. Lee Swindlehurst
cs.NI · cs.LG · eess.SP
Abstract
Learning-based schedulers have been proposed to provide real-time user, target, and access point (AP) association in distributed cell-free integrated sensing and communication systems. In a typical approach, a graph neural network (GNN), trained on labels from a mixed-integer linear program, maps lightweight per-AP statistics to decisions on AP clustering, user and target scheduling, and mode selection in one forward pass. Such solutions assume that hard constraints, enforced only as soft training penalties, hold at inference, and that the self-reported statistics are truthful. Using our ASSENT algorithm as an example, we find that despite high $F_1$ scores, many solutions violate at least one hard constraint, demonstrating that high prediction accuracy does not ensure joint feasibility. Projecting the GNN output onto a feasible solution restores constraint satisfaction with low utility loss, even with a simple greedy repair procedure. We further show that feasibility alone does not guarantee robustness to false data injection attacks. A single malicious AP that reports false information cannot substantially increase its user associations, but can greatly increase the rate of infeasible solutions. The effect of such attacks depends on the type of information being falsified. Misreporting information that affects the objective can largely be mitigated through feasibility projection, whereas falsifying information that affects the constraints cannot. The latter can, however, be detected using a low-complexity cross-AP consistency check. These results show that learned ISAC schedulers should be evaluated using constraint-aware feasibility metrics in addition to conventional accuracy measures.
cs.LG / 77 / 2609.03142
Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies
Yue Yang, Diego Romeres, Chiori Hori, Gedas Bertasius, Daniel Szafir, Siddarth Jain
cs.RO · cs.CV · cs.LG
Abstract
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).
cs.LG / 78 / 2609.03320
Beyond .WAV: Design and Software Verification of VocalCap, a Traceable Browser-Based Audio Capture System for Vocal Biomarker Research
Augusto Camargo
cs.SD · cs.LG
Abstract
Remote voice studies often retain a final audio file with limited evidence about how it was captured, transferred, processed, and accepted. This paper presents VocalCap, an institution-controlled, browser-based system for self-guided capture of voice and related acoustic signals by participants without technical training. A versioned protocol drives the workflow. Each accepted recording retains a browser-native object, a client-lossless Float32 WAV derived from the same MediaStream, and a server-canonical mono PCM16 WAV, linked to evidence of capture execution, technical quality, byte-level integrity, recovery, and transformation provenance. IndexedDB preserves accepted browser artifacts until server confirmation, while session completion requires successful verification of every task and artifact. Software tests challenged the acquisition contracts with malformed or altered objects, exact-zero interruptions, channel-topology variants, and interrupted or repeated operations. A post hoc technical audit of 39 consented pilot recordings found 25 sample-identical stereo files and 14 files with signal confined to the left channel. Topology-aware active-channel selection limited the canonical root-mean-square level difference to less than 0.001 dB in all 14 affected files; equal-weight stereo averaging would have introduced approximately 6.02 dB of attenuation. Production end-to-end verification completed two five-task profiles in Chromium and WebKit, yielding 10 accepted recordings and 30 retained artifacts that passed server-side integrity and format checks. The results verify VocalCap's software behavior under the tested browser-engine conditions. Device-level acoustic agreement, target-population usability, clinical validity, and biomarker performance remain subjects for separate studies.
cs.LG / 79 / 2609.03816
When Vision Meets Graphs: A Survey on Graph Reasoning and Learning
Xinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian, Xiaozhuang Song, Yaoyao Xu, Zhongkai Xue, Dingshuo Chen, Shu Wu, Philip Torr, Tianshu Yu
cs.SI · cs.CV · cs.LG
Abstract
Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: chemists read molecular diagrams and social scientists inspect network visualizations. Despite decades of work on graph visualization, most graph learning pipelines still treat graphs purely as symbolic structures, rarely leveraging the visual form of graphs. We argue that this gap deserves renewed attention in the era of powerful vision and vision-language models. This survey provides a first systematic overview of the emerging area we term vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning. We organize existing work into three threads. Vision for Graph Reasoning studies how models can use visual depictions of graphs to understand structure and carry out multi-step reasoning. Vision for Graph Learning explores how visual features can complement or augment graph encoders beyond known limitations of message passing. Scientific Graphs examines domains where standardized depiction conventions support both reasoning and learning. Our goal is to clarify what current methods can and cannot do, and to outline a path toward foundation models that perceive and reason about graphs as scientists do.
cs.LG / 80 / 2609.03095
Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief
Robert Engel
eess.IV · cs.CV · cs.HC · cs.LG
Abstract
Smartphone skin photographs are indispensable to teledermatology, yet assessing the diagnostic suitability of submitted cases (gradability) remains a critical bottleneck in mobile care workflows. Dermatologists routinely review multiple photographic views (regional, angled, and close-up) to identify consistent textural detail rather than relying on a single image. We present the Semantic Tri-view Pipeline, an interpretable architecture for automated teledermatology gradability screening that formalizes epidermal micro-relief as a computable biomarker of image quality. Using an expert-annotated subset of the public SCIN dataset, we train a lightweight DeepLabV3+ model to segment micro-relief fidelity. These spatial masks are then aggregated across up to three case views with a logistic regression classifier, leveraging viewpoint redundancy to support robustness under uncontrolled smartphone acquisition. This approach learns context-aware, clinically intelligible heuristics, such as penalizing high-fidelity texture in regional distance views. Evaluated at a predefined 90% sensitivity operating point, the system's apparent errors largely reflect subjective clinical variance on borderline cases where clinicians rely on non-visual metadata. On SCIN, performance improves from an AUC of 0.81 (80.6% PPV) on variance-heavy majority-consensus cases to 0.96 (97.7% PPV) on optically unambiguous unanimous cases. Overall, this work delivers an interpretable, privacy-by-design, edge-ready system that can provide real-time feedback during case submission to filter ungradable photo sets before review.
cs.LG / 81 / 2609.03977
Cooperative Multi-Task Semantic Communication for Joint Classification and Regression Tasks
Ahmad Halimi Razlighi, Mohammad Siddiqur Rahman, Maximilian H. V. Tillmann, Edgar Beck, Armin Dekorsy
eess.SP · cs.LG · eess.IV · stat.ML
Abstract
Multi-Task semantic communication (SemCom) prioritizes simultaneous execution of multiple tasks over bit-accurate reconstruction in future intelligent networks. In our prior work [1], we introduced the cooperative multi-task SemCom (CMT-SemCom) framework, in which the semantic encoder is divided into a common unit (CU) and multiple specific units (SUs) to facilitate cooperative multi-task processing. However, CMT-SemCom has been evaluated on homogeneous classification tasks on simplistic datasets, limiting its applicability to real-world perception systems. In this paper, we extend our CMT-SemCom to jointly handle heterogeneous classification and regression tasks on the complex Cityscapes dataset. We adopt the information maximization (InfoMax) principle so that it accommodates mixed discrete and continuous semantic variables. In particular, we benchmark the proposed framework against independent single-task training, a conventional task-agnostic digital transmission, and single-encoder multi-decoder SemCom. Additionally, we investigate the impact of CU capacity on joint task performance, providing design insights. Extensive evaluations demonstrate that CMT-SemCom significantly outperforms the benchmarks.
cs.LG / 82 / 2609.03329
Introducing SINFONIA: Symplectic, slimplectic and Magnusian (Neural) Flows for Orbital Numerical Integration and Acceleration
Lidia J. Gomes Da Silva
gr-qc · astro-ph.HE · astro-ph.IM · cs.LG · physics.comp-ph
Abstract
Long-duration gravitational-wave modelling must resolve fast orbital motion together with slow dissipative evolution while preventing small numerical errors from accumulating into secular phase drift. Here we ask whether the finite-time evolution map itself can be learned as an explicit, differentiable, structure-preserving object and then repeatedly composed through a complete inspiral. We construct three neural-flow architectures: a symplectic and slimplectic flow on Galley's doubled phase space, [SINFONIA-J0]; a Taylor-anchored flow, [SINFONIA-J1]; and a Magnusian flow that learns the finite-time dissipative correction in the interaction picture, [SINFONIA-J2]. Applied to a 2.5PN neutron-star inspiral, all three expose the same controlling mechanism: long-time accuracy is governed not by pointwise map error alone, but by its signed projection onto a single secular channel fixed by energy--angular-momentum balance. Encoding this structure allows the learned maps to remain accurate through $10^{2}$--$10^{5}$ window compositions to coalescence at timesteps of a full orbital period and beyond, reaching chained phase errors orders of magnitude below a benchmark slimplectic integrator at lower cost. The same secular structure can also be exploited for physics inference: when the channel is left unconstrained, the accumulated phase retains enough information to recover an un-modelled dynamical-friction-like force, both parametrically and as a learned function of separation. Network-off controls isolate the contribution of learning from the analytic structure already built into each map. These results establish a proof of concept for structure-preserving learned evolution maps as tools for fast long-duration integration and physics inference in gravitational-wave source modelling.
cs.LG / 83 / 2609.03361
Grassmann--Plücker Parametrization of Convolutional Filter Subspaces: Regularity and Closed Embeddings
Hongyu Yuan, Huaiqing Zuo
math.AG · cs.LG
Abstract
We propose a geometric parametrization of the filters in a single convolutional layer: the parameter is no longer an ordered family of filter vectors, but a fixed-dimensional subspace of the filter space. For one-dimensional finite-stride convolution, the filter-to-convolution-operator correspondence gives an injective linear map $\mathcal{C}:\mathcal{K}\to H$. This map sends filter subspaces in $\mathrm{Gr}(q,\mathcal{K})$ to operator subspaces in $\mathrm{Gr}(q,H)$; composing it with the Plücker embedding yields a projective parametrization $Φ:\mathrm{Gr}(q,\mathcal{K})\to\mathbb{P}(\bigwedge^q H)$. Using $T_U\mathrm{Gr}(q,\mathcal{K})\cong\mathrm{Hom}(U,\mathcal{K}/U)$, we compute the differential of the induced Grassmannian map and show that the differential of $Φ$ is injective at every point. We then use the vanishing equations for Plücker coordinates and standard affine coordinates on a Grassmannian to prove that $\mathrm{Gr}(q,\mathcal{C}(\mathcal{K}))\hookrightarrow\mathrm{Gr}(q,H)$ is a closed embedding, and hence that $Φ$ is a closed embedding. Consequently, the parameter space is isomorphic to its projective image, the parametrization is finite and birational onto its image, every fiber is a singleton, and the resulting projective neural variety is smooth. For $k=4$ and $q=2$, we also use Singular to recover the image ideal and check its dimension, degree, chart rank, and smoothness. This computation illustrates, rather than replaces, the general proof. Finally, we discuss possible connections with filter redundancy and low-rank convolution, while distinguishing the proved geometric results from application proposals requiring numerical validation.
cs.LG / 84 / 2609.03190
Coupled Tensor-Tensor Completion Method with Applications in Drug Repurposing
Maryam Bagherian, Albert Hung, Ivo Dinov, Joshua Welch
math.NA · cs.LG
Abstract
Many biomedical challenges can be posed as tensor completion problems where the observed entries of a multidimensional array (a tensor) are used to impute the missing values. In such settings, incorporating side information about the modes of the tensor, such as gene-gene similarity, can significantly enhance the solutions of the completion problem. Most existing tensor completion methods can only incorporate side information in the form of matrices. In this study, we introduce a novel framework to incorporate side information in the form of tensors. Our new approach, called Coupled Tensor-Tensor Completion (CTTC), leverages the hidden connections among multimodal tensors to improve tensor completion performance. In addition to practical utility, CTTC has theoretical foundations in distance metric learning and group theory. We derive an alternating algorithm to solve the CTTC optimization problem and establish its convergence to a stationary point. Finally, we show that CTTC outperforms state-of-the-art tensor completion methods at predicting drug effects. Results: Compared with other tensor completion methods, including HaLRTC, CTRC, Cell, and NTDDR, CTTC demonstrates superior run-time and RSE tensor completion accuracy on two benchmark datasets, DTD and LINCS.
cs.LG / 85 / 2609.03343
Learning Informative Prior with Infinite-Dimensional Continuous Normalizing Flow for Bayesian Inverse Problem
Yang Zhao, Junxiong Jia, Tao Zhou
math.NA · cs.LG
Abstract
This paper addresses infinite-dimensional Bayesian inference for inverse problem of partial differential equations with model parameters in infinite-dimensional Hilbert space. To effectively incorporate prior information, we propose a novel continuous normalizing flows based infinite-dimensional model. Specifically, by introducing a well-defined neural ordinary differential equation in infinite-dimensional space, a simple reference measure can be transformed into a more complex measure which encodes the prior information. A corresponding theoretical framework is established to ensure the well-posedness of our proposed Bayesian prior in infinite-dimensional space. We also provide training methods of the prior for two distinct data settings, along with two sampling algorithms for the resulting Bayesian posterior. The proposed framework is applied to three representative inverse problems: the simple smooth inverse problem, inverse scattering problem, and the inverse heat conduction problem. Numerical experiments support the theoretical analysis and demonstrate the efficiency of the proposed algorithms.
cs.LG / 86 / 2609.03626
Residual neural networks overcome the curse of dimensionality for semilinear heat equations
Ilkhom Mukhammadiev, Diyora Salimova
math.NA · cs.LG · math.AP · math.PR
Abstract
Rigorous results show that feedforward neural networks can overcome the curse of dimensionality in the numerical approximation of high-dimensional partial differential equations (PDEs), but comparatively little is known about residual neural networks (ResNets) in the nonlinear PDE setting. We prove that ResNets overcome the curse of dimensionality in the numerical approximation of solutions of semilinear heat equations with globally Lipschitz continuous, gradient-independent nonlinearities: under polynomial growth and network approximability hypotheses on the PDE data, there exist $η\in(0,\infty)$ and ResNets $Ψ_{d,\varepsilon}$, $d\in\mathbb{N}$, $\varepsilon\in(0,1]$, with at most $ηd^η\varepsilon^{-η}$ parameters whose realizations approximate the solution in dimension $d$ with an $L^2$-error of at most $\varepsilon$. The proof represents one deterministic realization of a multilevel Picard estimator by a ResNet whose shortcut connections transmit the spatial variable and a scalar accumulator, while the residual branches successively add the summands of the estimator. For ridge-sum initial conditions, admissible sigmoidal activations, and globally Lipschitz truncations of the nonlinearity, we obtain, for every $ξ>0$, the explicit bound $C_ξd^{4+ξ}\varepsilon^{-(3+ξ)}$ on the number of parameters.
cs.LG / 87 / 2609.03589
Correlated initialization of deep residual networks
Felix Benning, Ivan Nourdin, Giovanni Peccati
math.PR · cs.LG · stat.ML
Abstract
We study the large-depth behavior of residual networks whose weights are correlated across layers at initialization. Our results confirm and extend a conjecture of Marion et al. [2025], according to which correlated initializations should interpolate continuously between the Brownian stochastic differential equation arising from independent initialization and the ordinary differential equation arising from perfectly correlated initialization. When the initialization is obtained from the application of a feature function to a stationary Gaussian sequence with regularly varying correlation, we prove that there exists a unique critical scaling such that the infinite-depth limit is the solution of a Young differential equation driven by a Hermite process. Hermite processes reduce to the fractional Brownian motion if the feature function generating the initialization has Hermite rank one, which is the case for the identity function, for example. We show that the critical scaling and asymptotic limit are uniquely determined by the decay of correlations together with the Hermite rank of the feature function. Consequently, the correlation structure and Hermite rank of the initialization represent meaningful hyperparameters in the asymptotic regime. By contrast, under finite-variance iid initialization, the asymptotic driver is universally Brownian up to normalization regardless of the choice of distribution. Our proofs rely on a collection of novel results establishing a robust stability theory for Young differential equations in Banach spaces.
cs.LG / 88 / 2609.03246
What is Smoothness?
Zachary P Bradshaw
math.RT · cs.LG
Abstract
Smoothness of a function on the real line is reflected in the decay of its Fourier transform, which suggests that smoothness of a function in $L^2(G)$ for a group $G$ should mean concentration of the Fourier coefficients at low frequency. Such a reading presupposes an ordering of the irreducible representations of $G$, but for non-abelian $G$, no ordering is canonical. Given a symmetric generating set $S$, the Laplacian of the associated Cayley graph is block diagonal over the dual, and we order the irreps by the mean of the eigenvalues in each block. This produces an ordering function $ω:\widehat{G}\to\mathbb{R}$ that depends only on the pair $(G,S)$. This function is bounded between zero and two, vanishing only at the trivial representation and achieving the upper bound exactly when the Cayley graph is bipartite. We then ask how much freedom the construction has. Within the class of operators satisfying natural axioms, the induced orderings are exactly the real functions on the dual vanishing at the trivial representation and agreeing on conjugate pairs, and the orderings coming from inversion orbits of conjugacy classes form a basis for them. We cut the freedom down further by requiring two additional inputs: nonnegativity of the class weights and a declaration of which group elements count as uniform incremental changes, which pins the operator to the Cayley-Laplacian up to positive scale. We observe that the construction persists for compact groups even though the Cayley graph does not, and we extend the theory to finite sets carrying a transitive group action, where the acting group selects which frequencies exist and the generating set orders them. The answer to the title question is therefore that smoothness is a property of a function together with a choice of group and generating set, not of the function alone.
cs.LG / 89 / 2609.03210
Improving precipitation forecasts in an AI weather model using observational data
Julian F. Schmitt, Bertrand Delorme, Robert C. King, Yashica Patodia, Tapio Schneider, Aditi Sheshadri, Ravi Jain
physics.ao-ph · cs.LG
Abstract
Artificial intelligence weather prediction (AIWP) systems now surpass state-of-the-art physical models for medium-range weather forecasting. Current global AIWP models are trained almost exclusively using one reanalysis dataset, ERA5, but it has known biases, particularly for precipitation. Here we fine-tune a graph-transformer architecture with IMERG precipitation data at 0.25° resolution. The resulting model improves medium-range continuous ranked probability scores by up to 19%, while also demonstrating superior skill for tropical storms and drizzle events. Our model exceeds the Brier skill score of state-of-the-art operational models on extreme rainfall prediction by 57% globally; however, a physics-based operational model remains more reliable for the heaviest precipitation events. Our results demonstrate that incorporating observations-based precipitation data directly into training can substantially improve precipitation forecasts.
cs.LG / 90 / 2609.03046
Advances in Machine Learning for Directed Evolution: A Five-Year Retrospective
Bruce J. Wittmann
q-bio.BM · cs.LG
Abstract
The last five-plus years have seen many protein engineering disciplines transformed by advances in machine learning (ML), but the same cannot be said for directed evolution. Reflecting on a previously co-authored perspective, I discuss why I believe this to be the case, arguing that a disconnect between the goals of machine-learning-assisted directed evolution (MLDE) researchers--"identify an optimal protein"--and the goals of directed evolution more broadly--"identify a sufficient protein given time and resource constraints"--is a principal culprit. As an example, I highlight how nearly all current MLDE methods neglect to account for the cost of DNA synthesis, resulting in strategies that have limited practical applicability regardless of the underlying models' capabilities. I close by discussing recent works that are exceptions to this overarching trend, and emphasize that the last five years of efforts in ML-assisted protein engineering and the prescribed reframe of MLDE objectives need not be mutually exclusive.
cs.LG / 91 / 2609.04165
Parameterised graph theory for tensor networks: entanglement rerouting, structural simplification, and agnostic tomography
Matthias C. Caro, Natalie McHugh, Sergii Strelchuk
quant-ph · cs.DS · cs.LG
Abstract
Parameterised graph theory studies how the complexity of graph-theoretic problems depends on structural parameters of the input graph. This perspective has proved useful in analysing tensor-network simulation (Markov and Shi, 2008). Its implications for tensor-network representations and tomography are less well understood. In particular, which graph parameters determine whether a tensor-network state (TNS) admits a tractable matrix product state (MPS) or tree tensor network (TTN) representation, and which control the complexity of learning the state? We address these questions using parameterised graph theory. First, we show that cutwidth and tree-cutwidth bound the bond dimension overhead required to represent a TNS as an MPS or TTN. In the TTN case, tree-cutwidth also bounds the local dimension of the grouped subsystems. The proofs are based on entanglement rerouting, a tensor-network analogue of rerouting information in a classical network. Second, we derive graph-dependent upper bounds on the sample and computational complexity of realisable TNS tomography, with exponents that depend on cutwidth, tree-cutwidth, and a new graph parameter, learning complexity, which we bound in terms of degree and treewidth. We obtain these results by extending the disentangling MPS learner of (Cramer et al., 2010), as analysed further in (Bakshi et al., 2025; Lin et al., 2025), to TTNs and to tensor networks on arbitrary known graphs. Finally, we extend the framework beyond the realisable setting. For an arbitrary input state, our agnostic learner outputs a pure state whose fidelity is within additive error $ε$ of the optimum over tensor-network states on the given graph with a given bond dimension, with explicit graph-dependent bounds on sample and computational complexity.
cs.LG / 92 / 2609.03104
Occupancy-based Quantile Risk Control
Zihao Shi, Huajun Xi, Bingyi Jing, Hongxin Wei
stat.ML · cs.LG
Abstract
Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite-sample guarantees. To accommodate a broader class of risk notions, quantile risk control extends this framework to quantile-based risk measures. However, existing methods either suffer from excessive conservatism or lack rigorous finite-sample guarantees. To address these limitations, we introduce Occupancy-based Quantile Risk Control (OQRC), a novel method that provides tight risk control bounds with finite-sample validity. Our key idea is to formulate risk control as a finite-occupancy problem by partitioning the loss space with the ordered calibration losses. Specifically, we estimate the distribution of test losses across the resulting bins and upper-bound the risk by the maximum loss attained within each bin. We then select the parameter $λ$ such that this upper bound does not exceed a predefined threshold $α$ with high probability $1-δ$. Theoretically, we establish a finite-sample guarantee showing that OQRC yields tight risk control bounds that converge to the optimal bounds at a provable rate of $\mathcal{O}_ p(n^{-1/2})$. Extensive experiments demonstrate the effectiveness of our method, reducing the risk gap by up to 78.64\% on common benchmarks.
cs.LG / 93 / 2609.03129
A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations
Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
stat.ML · cs.LG
Abstract
Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks. We address this by introducing a simple closed-form ``two-stage'' compositional formula $\hat{f}$ for reconstructing an unknown Lipschitz function $f:\mathcal{X}\to \mathbb{R}$ on a metric space $(\mathcal X,ρ)$ from $N$ i.i.d. noisy observations. Our main result is a high-probability uniform ($L^{\infty}$) recovery guarantee that jointly controls approximation and statistical errors while enjoying an optimization error of zero; in particular, we do not assume oracle access to an approximate ERM. Our secondary main results establish the optimality of our formula in three complementary senses. 1) Function space: On Ahlfors-regular metric spaces, the hypothesis class parameterized by our formula attains the optimal fat-shattering dimension. 2) Parameter space: Its dependence on the parameters is maximally numerically stable, in the sense that a smaller approximation error cannot be achieved with a smaller Lipschitz dependence on the model parameters. 3) Forward pass: Its dependence on the input is maximally regular, matching the Lipschitz constant of the target function $f$. When $\mathcal X=[0,1]^d$ is equipped with the $\ell^\infty$ norm, $\hat{f}$ admits algorithmic ReLU-MLP and exact ReLU-multi-head transformer realizations of depth $\mathcal{O}(\log(N))$ with $\mathcal{O}(N)$ nonzero parameters.
cs.LG / 94 / 2609.03501
Towards a Statistical Understanding of Mixture-of-Experts
Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
stat.ML · cs.LG · math.ST
Abstract
Mixture-of-experts (MoE) architectures increase model capacity by combining a collection of expert predictors through input-dependent routing, while often activating only a small subset of experts for each input. Despite their growing importance in modern large-scale models, the statistical roles of their design choices, especially routing, sparse activation, and shared experts, remain only partially understood, as existing theory has largely focused on parametric or correctly specified MoE models. In this paper, we view MoE as a form of localized aggregation and show how this localization reshapes the approximation-estimation-computation tradeoff. We derive oracle risk bounds for learning dense and sparse routing with evolving experts, separating approximation, expert-learning, and router-estimation errors, and characterize how sparse Top-K routing can retain the benefits of localized aggregation while controlling per-input computation. We also interpret gating through the geometry of input space, relating routing performance to regions of local expert advantage, and show how shared experts, as adopted in architectures such as DeepSeekMoE, can extract common predictive structure so that routed experts focus on residual local variation. Together, these results provide a unified statistical framework for understanding MoE through input-dependent expert aggregation, in which expert specialization and computational tradeoffs are governed by local predictive structure.
神经与进化计算 (cs.NE)
3
cs.NE / 1 / 2609.03352
Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming
Hao Mao, Xu Tony Liu, Shuai Lu, Peng Zhao, Wenzheng Jiang, Yuntian Chen
cs.NE · cs.DC · cs.LG · cs.MS
Abstract
Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated frameworks to omit it or restrict it to lightweight forms. We present a GPU-resident, batched Levenberg--Marquardt solver that optimizes constants across a structurally heterogeneous population of expression trees using a fixed number of population-wide CUDA launches per iteration. Reverse-mode automatic differentiation assembles the per-tree Jacobian in one backward sweep, making the dominant per-iteration cost independent of the number of constants per tree, and a double-precision delivery guard guarantees that returned constants are never worse than their initial values. On early-generation populations, the solver sustains up to $5.1{\times}10^{5}$ trees per second on an NVIDIA A100; at a GPU-saturated benchmark configuration it delivers roughly $9.9{\times}$ the throughput of Operon running on a 64-core EPYC 7763, while matching fp64-reference quality. Integrated in-process into EvoGP, the solver enables end-to-end search to recover governing equations on $10$ of $18$ constructed problems versus 0 for stock EvoGP. Our code is at https://github.com/TensorConv/CuSR.
cs.NE / 2 / 2609.03724
Genetic Algorithms for Tractable Bayesian Network Fusion via Pre-Fusion Edge Pruning
Pablo Torrijos, José A. Gámez, José M. Puerta, Juan A. Aledo
cs.NE · cs.LG
Abstract
Bayesian Network (BN) fusion combines multiple input networks into a single structure, balancing dependency preservation with computational tractability. While unrestricted fusion retains all dependencies, it often results in overly complex networks with high treewidth, which affects inference scalability. Limited fusion mitigates this by pruning edges to control treewidth but risks overfitting to input-specific noise and omitting dependencies from the original BNs. This paper introduces a consensus framework that prioritizes shared structures among input networks while enforcing treewidth constraints, ensuring a good consensus. We propose genetic algorithms with advanced initialization, specialized operators, and a tailored fitness function. Additionally, we adapt existing methods to this problem and implement greedy baselines for benchmarking and further optimization. Experiments on synthetic and real-world BNs show the superiority of the proposed genetic algorithms over the adapted methods and greedy baselines.
cs.NE / 3 / 2609.04195
Axonal delay dispersion decides whether a neuron detects an event or a sequence, and predicts cortical column diameter
Cheng Bi, Jipeng Sun
q-bio.NC · cs.NE
Abstract
Cortical neurons fire sparsely -- often fewer than one spike per sensory window -- making rate coding insufficient and temporal coding a necessity. That conduction delays convert firing order into synchrony is long established. What governs which class of temporal feature a neuron detects -- one volley of coincident input, or two in a particular order -- has not been examined. We propose a delay-signature framework in which the axonal conduction delays converging on a dendritic branch constitute a physical key: only input sequences whose spike-time differences the delays compensate arrive synchronously, and coincidence detection, via calcium plateau thresholds, converts that synchrony into an all-or-none output. In simulations of an integrator-neuron model we report three results. First, a single physical scalar -- the dispersion of the delay set -- moves a population from event detection to order-selective sequence detection. The transition is emergent under random delays and connectivity: at narrow dispersion sequence detectors do not exist, and the dispersion at which they overtake event detectors tracks the inter-event interval with a slope statistically indistinguishable from one. This maps a computational distinction onto the anatomical one between myelinated and unmyelinated projections, making myelination a switch on what a neuron computes, not only a regulator of speed. Second, the same dispersion sets the code's limits: it bounds the longest codable interval and fixes an absolute timing tolerance of about a millisecond, with slowing better tolerated than speeding. Third, that millisecond window and horizontal conduction velocity together predict cortical column diameter, and the two areas with direct measurements fall where the relation puts them. One anatomically measurable parameter thus sets what a neuron detects and the limits of what it can represent.
计算语言学 (cs.CL)
46
cs.CL / 1 / 2609.03005
Unifying Conformal Language Tasks with In-Context Ensembles
Xiao Shi Huang, Chen-Yuan Lin, Bruce Kuwahara, Kin Kwan Leung, Jesse C. Cresswell
cs.CL · cs.LG · stat.ML
Abstract
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and conciseness, removing as much irrelevant information as possible. Conformal prediction methods have been used to guarantee coverage, and must be optimized for conciseness through design of a score function. State-of-the-art scoring functions use hand-engineered LLM prompts asking the model to rate the importance of content, but manual prompt engineering is labor-intensive and task-specific. We introduce the Conformal Relevance framework which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input. We demonstrate this framework's application on seven NLP tasks, and also theoretically study the impact of diversity for ensembled conformal scores, giving a complementarity condition that characterizes when ensembling improves worst-case sentence scores, and a saturation bound on ensemble improvement.
cs.CL / 2 / 2609.03047
SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
Michael J. Bommarito
cs.CL · cs.AI · cs.IR
Abstract
Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Synthetic Harness for Evaluating LLM Fitness, addresses this gap. It is a Python system that turns labelled taxonomies, writing specifications, and a generation budget into controlled benchmark data and evaluation tasks. This first release contains 62,899 model-written documents based on Library of Congress vocabularies, with tasks for classification, clustering, retrieval, pair classification, and instruction retrieval. We compare TF, TF-IDF, BM25, popular encoders, and, on subject classification only, zero-shot decoders; each method appears only on tasks that support it. Subject classification reaches 0.8887, while genre-form classification reaches only 0.2605, and several pair and clustering tasks remain near chance. Sparse methods remain competitive on classification, while TF-IDF is the fastest measured arm in the subject timing experiment. SHELF also varies bibliographic facets independently and can generate new, verifiably unseen documents after a model's training cutoff. Comparisons with LCSHBench and Project Gutenberg show that model rankings transfer more reliably than absolute scores, but SHELF scores do not estimate accuracy on production catalogue data. We release all source code and data under permissive licenses on GitHub and Hugging Face.
cs.CL / 3 / 2609.03181
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
Alejandro Barón García, Feng Wang, Emilia Garcia Casademont, Han Xiao
cs.CL · cs.CV
Abstract
We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at https://huggingface.co/jinaai/jina-ocr-v1.
cs.CL / 4 / 2609.03201
MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval
Meriem Yacoubi, Pia Schmidt, Nenad Petrovic, Ahmed Frikha, Martin Kirchhoff, Alois Knoll
cs.CL · cs.LG
Abstract
Long-term LLM agents must preserve information across interactions while distinguishing repeated evidence, historical states, updates, and unresolved contradictions. Existing textual memory systems retrieve semantically relevant memories efficiently but often leave these relationships implicit, whereas richer structured approaches model them through global graphs, hierarchical abstractions, or reflection at greater complexity. We introduce MemoryLACE (MemLACE), a lightweight memory framework that explicitly models the lifecycle of textual evidence through sparse merge, supersession, and contradiction relations while preserving atomic natural-language memories and their provenance. Rather than retrieving memories independently, MemLACE reconstructs relation-aware evidence units that expose current, historical, supporting, and conflicting evidence for downstream reasoning. Across BEAM and StructMemEval, using open-weight and proprietary LLM backbones, MemLACE achieves the highest overall performance in same-backbone comparisons while reducing end-to-end runtime on BEAM by 66.6% relative to Hindsight, the strongest reported reflective-memory baseline. Ablation studies identify lifecycle expansion and temporal awareness as the principal contributors to these gains. Together, the results demonstrate that explicitly modeling the local lifecycle of textual evidence is sufficient to substantially improve long-term memory reasoning without requiring comprehensive knowledge graphs or global reflection.
cs.CL / 5 / 2609.03215
SWIM: Student Writing Simulation via Proficiency-Conditioned Generation
Heejin Do, Jakub Kontak, Mrinmaya Sachan
cs.CL · cs.LG
Abstract
Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.
cs.CL / 6 / 2609.03221
Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar, Rahul Joshi
cs.CL · cs.LG
Abstract
Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-running an identical condition ten times over sixteen vignettes (same narrative, same descriptor string, nothing varied) moved a clinical agent's action in 8.7% of outcome-vignette cells, and instability was heterogeneous across actions by a factor of eight, from 0.022 for ICU escalation to 0.179 for controlled-substance caution. No demographic contrast in our data was distinguishable from that floor. A second model gives a pooled floor of 6.7% and ranks the six actions almost identically (Spearman 0.94, exact p=0.017), so the floor is not one system's artefact. Majority-vote aggregation over five draws removes 39% of it and then flattens, and a null simulation attributes the residue to heterogeneous per-cell rates, so replication mitigates without eliminating. Any counterfactual fairness estimate reported without a per-action floor beside it therefore cannot be read as evidence of disparity. The measurements were taken with FairMedAgent, an evaluation harness for disparity in the actions of clinical LLM agents whose estimand, the within-range counterfactual flip rate, counts only flips between actions a published decision rule admits and a clinician has adjudicated. That estimand requires band adjudication, which is under way; no disparity result is claimed here. Each synthetic vignette runs a six-stage trajectory (five model-facing decisions around a deterministic environment step) under fixed-form conditions spanning race, sex, age, insurance, English proficiency, and their intersections. The harness, the floor protocol, and every analysis script are released.
cs.CL / 7 / 2609.03273
Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers
Karthikeyan A, Jaya Nirmala S, Sangeetha Sivanesan, Indhu R, Pranav Kumar, Bharat Jude Johnson, Vishnu Ram
cs.CL
Abstract
Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct letters. Prior work targets word-level surface errors with rule-based methods, statistical n-gram models, Minimum Edit Distance, or hybrid pipelines with a transformer re-ranker; such methods cannot reliably handle contextual errors - subject-verb agreement, tense consistency, or cross-word sandhi - which require sentence-level understanding. We propose an end-to-end sequence-to-sequence formulation and fine-tune mT5-small and mBART-50 on a synthetic corpus of up to 657,720 noisy-clean Tamil sentence pairs spanning ten error categories. Both backbones follow the same four-stage progressive schedule, each stage targeting one weakness: surface noise (v2), contextual grammar (v3), single-site sandhi (v4), and multi-site cross-word sandhi (v5). On a 1,000-sentence balanced diagnostic set verified disjoint from all training data, our best model, mBART-50 v5, reaches 69.3% top-1 exact-match accuracy, with 87.5% on sandhi and 43.5% on subject-verb agreement. The schedule is what produces these gains: subject-verb accuracy rises from 1.0% to 52.5% once contextual pairs are introduced, and sandhi from 0% to 87.5% once multi-site sandhi pairs are. We additionally quantify a precision-recall trade-off this literature has not reported: sandhi recall is paid for monotonically in identity accuracy. Finally, Tamil-LLaMA-7B-Instruct reaches 19.0% zero-shot and 24.7% with three demonstrations against a 20.0% copy baseline, showing that a Tamil-adapted instruction model does not transfer to specialised sentence-level correction without task-specific supervision.
cs.CL / 8 / 2609.03293
PACE: Towards Surfacing Hidden Conflicts in User Requests
Yoojin Kim, Jihyoung Jang, Hyounghun Kim
cs.CL
Abstract
Personalized assistants should not only comply with user requests but also assess whether those requests are appropriate given the user's current circumstances. However, prior work has primarily focused on accurately executing requests, overlooking the need for assistants to account for context and engage in conflict-based refusal. Furthermore, while existing work on conflict or safety detection relies on explicitly provided factors, real-world scenarios often involve implicit factors that must be retrieved from a knowledge base (KB). To this end, we introduce Personalized Assistants for Conflict Evaluation (PACE), a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate. PACE pairs user requests grounded in well-defined personas with egocentric KB facts, requiring models to integrate contextual evidence to determine whether a request is conflicting. This implicit retrieval setting hinders the direct association between user requests and conflict-inducing knowledge, making it difficult for existing models to identify relevant user-specific facts. To address this challenge, we further propose PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence. Experiments on PACE evaluate both evidence retrieval quality and conflict decision accuracy, showing that PaceMaker consistently outperforms existing approaches.
cs.CL / 9 / 2609.03331
FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
Jiayuan Ma, Yuqi Lu, Weiyang Guo, Chenrui Wang, Junyi Shu, Xuebo Liu, Min Zhang, Jing Li
cs.CL
Abstract
Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dialogue turns. We introduce FPCO-Dialog, a benchmark for evaluating correction and cooperation behavior in VLMs under repeated false premises. FPCO-Dialog contains 1,080 images and 10,800 question turns, stratified by visual complexity, object category, and false-premise class, and uses a 10-turn protocol in which a correct dialogue prefix is followed by repeated false-premise referring expressions. We evaluate 20 commercial and open-source VLMs with a model-agnostic protocol and CorrTP@K, a correction-rate metric over false-premise turns, scored by two independent detectors. FPCO-Dialog reveals substantial and persistent cross-model differences in aggregate correction tendency, model-specific turn-wise dynamics, and systematic variation across false-premise types under the benchmark's substitution distribution. The dataset, evaluation protocol, model outputs, detector labels, and code are available.
cs.CL / 10 / 2609.03366
Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT
Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam
cs.CL · cs.CY · cs.LO
Abstract
Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified rationales (how was the decision reached), assumptions (what was assumed rather than known), policy consistency (the same treatment for the same facts), and pivotal conditions (what would change the outcome). We introduce self-faithfulness as an automatic test of accountability: changing the pivotal conditions should change the decision. We examine accountable AI through clinical trial matching, a high-stakes task central to evidence-based medicine. Although LLM-based matchers match patients to trials reasonably accurately, they apply decision policies inconsistently and produce rationales that are unfaithful to their own decisions. We introduce VERDICT, an LLM-based agent that translates a decision task, its constraints, and its policy into Satisfiability Modulo Theories (SMT), then derives the decision with SMT and MaxSMT solvers -- so policies are applied consistently and decisions are accountable by construction. Across a SIGIR 2016-derived dataset and TREC 2021, VERDICT achieves the strongest decision accuracy among LLM-only and neurosymbolic baselines, applies policies with perfect consistency, and produces clinician-preferred rationales grounded in explicit assumptions and pivotal conditions, with improved counterfactual self-faithfulness.
cs.CL / 11 / 2609.03394
Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory
Divyesh Bommana, Mohammad Saim, Tianyu Jiang
cs.CL
Abstract
Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.
cs.CL / 12 / 2609.03426
Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations
Yunao Zheng, Bin Wen, Xiaojie Wang
cs.CL
Abstract
Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to high parameter and activation costs. We propose Lngram v2, which decouples the number of routes, memory dimension, and backbone width, and introduces a context-aware grouped-query attention readout to scale memory capacity independently. A zero-value Sink and counterfactual surrogate gradients further improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments across vision--language models (VLMs) of different scales show consistent improvements, including successful scaling to a 30B-parameter model. Compared with Lngram v1, Lngram v2 substantially reduces both total and activated memory parameters while maintaining or improving language modeling performance. Further analysis shows that its discrete IDs preserve substantial semantic structure of continuous hidden states, enabling semantic recovery from IDs alone and stable ID--semantic associations across datasets. These results establish Lngram v2 as an efficient and scalable latent conditional memory mechanism whose discrete addresses also provide a structured interface for analyzing internal model representations.
cs.CL / 13 / 2609.03432
Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks
Xiangyu Wang, Jin Wu, Xiaoyu Li, Chanjin Zheng, Yifeng Zhou
cs.CL
Abstract
Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, and wide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at https://github.com/Jaong/CreaEval.
cs.CL / 14 / 2609.03487
Pattern Over-Generalization of Knowledge Graph Embedding
Junsik Kim, Kangil Kim
cs.CL · cs.AI
Abstract
Knowledge graph embedding (KGE) demonstrates its effectiveness for predicting missing links in knowledge graphs (KGs) by projecting entities and relations into a low-dimensional vector space. It is crucial for KGE models to effectively capture inference patterns (patterns) inherent in KGs, such as symmetry/antisymmetry, inversion and composition. Although recent KGE models exhibit strong capabilities in modeling such diverse patterns, they suffer from inherent limitations stemming from pattern over-generalization, where embeddings learned from only a single pattern instance inevitably generalize that pattern to all related instances, i.e., generalize the pattern universally. To address this issue, we propose PogRE (Pattern Over-Generalization Robust Embedding), a simple but effective method that utilizes dense linear transformations and compound operations for relation representation. Our theoretical analysis demonstrates that a dense linear transformation allows a pattern to become progressively universal as more triples are observed in the pattern. Furthermore, after observing d+1 linearly independent entities (d+1 denotes the dimension of entity), the linear transformation guarantees universal generalization of the pattern across all related instances. Experimental results on three standard benchmark datasets show that PogRE outperforms existing state-of-the-art KGE models in link prediction. Moreover, our empirical results indicate that PogRE effectively addresses the negative impact of over-generalization.
cs.CL / 15 / 2609.03502
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
cs.CL · cs.AI
Abstract
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.
cs.CL / 16 / 2609.03577
Language, Language Models, and What We're Talking About
Malvina Nissim
cs.CL
Abstract
Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want language models to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between language models designed as technical products and language models designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.
cs.CL / 17 / 2609.03595
How Far Can Synthetic Data Take Thai OCR?
Kunat Pipatanakul
cs.CL · cs.AI · cs.CV
Abstract
We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on training granularity, with in-domain reconstruction approaching real printed supervision under page-level training (1.82% versus 1.31% median character error rate) but underperforming out-of-domain reconstruction under crop-level training (15.59% versus 5.52%). Guided by these findings, we adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero using 45,723 synthetic pages: relative to its base checkpoint, it reduces median character error rate from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting and outperforms Typhoon OCR v1 7B on all five evaluation sets, showing that synthetic-only training can be competitive.
cs.CL / 18 / 2609.03597
KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records
Tasmiad Hasan, Arafat Zaman Ratul, Sarker Sadman Saalim, S. M. Shah Nawaz Hossain, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi
cs.CL
Abstract
Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten records that carry these fractions, RS Khatians, are the authoritative title record for millions of parcels and a frequent subject of civil litigation, yet no benchmark has asked whether a machine can read one. We introduce KhatianDoc, a four-task benchmark built from 107 real RS Khatian records from the Vumi (land) Office of Munshiganj, Bangladesh: symbol recognition, base-16-to-decimal conversion, structured field extraction, and legal document question answering over 1,634 QA pairs. Ground truth was transcribed by hand, verified by a land-law practitioner to full agreement, and anonymized through positional tokens that keep the referential distinctions multi-hop questions depend on. We evaluate six multimodal LLMs (8B to 72B+, open and closed) under a fixed zero-shot protocol. Five QA categories, 39.3% of our stratified set, return zero correct answers from every model; on the arithmetic task, every model that emits a number does worse than a constant-mean baseline, with exact- and near-match scores coinciding: decorrelation, not approximation. Auditing our own metrics surfaced two artifacts in opposite directions: we correct a refusal-scoring bug and report the fixed scores beside the originals, and flag an inflated metadata metric as an upper bound. KhatianDoc documents not a performance gap but the absence of a capability, with verified ground truth for future systems. Code and data, with a redacted image release, are publicly available.
cs.CL / 19 / 2609.03633
</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination
Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim
cs.CL · cs.AI
Abstract
Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that injects an end-of-think token (EoT, </think>) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a clean answering phase. Answering-phase generation can continue before the model regenerates another EoT, with the span preceding this regenerated EoT scaling with the reasoning tokens saved by early exit and exhibiting continued reasoning behavior. We call this spurious CoT termination, where reasoning-like generation continues into the answering phase. We hypothesize that insufficient attention to the injected EoT contributes to spurious CoT termination and probe this hypothesis with Exit-token Attention Biasing (EAB). Across four LRMs, five benchmarks, and two early-exit methods, increasing attention to the injected EoT reduces spurious CoT termination and answering-phase length. These results reveal a limitation of controlling LRMs by externally matching their explicit think-block format. Inserting the EoT token conforms to this format but does not by itself guarantee the intended reasoning-to-answering transition. Our code is available at https://github.com/Seunghee-Koh/Spurious-CoT-Termination.
cs.CL / 20 / 2609.03652
The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification
Sara Sorahi, Kevin Tang, Reza Kazemian
cs.CL
Abstract
Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in the context of discourse pragmatic function classification, a task where data sparsity is a structural feature rather than a collection artefact. Using 410 manually annotated instances of the English word look drawn from the British National Corpus, spanning four functions: Attention Signal, Directive, Discourse Marker, and Interjection. We generate synthetic training examples with Llama 3.1 and partition them by their cosine distance from real training data in RoBERTa embedding space. We compare six training conditions that differ in the placement of synthetic examples relative to the empirical decision boundary, while holding augmentation quantity constant across conditions. All augmented conditions improve macro F and accuracy over the real only baseline, but core proximal examples (NEAR) yield the largest gains in macro F (0.113), while a distance balanced mix achieves the highest accuracy (0.748). No condition improves AUC, indicating that augmentation shifts the decision boundary rather than improving the model's underlying probability estimates. These findings suggest that where synthetic examples land in representation space matters as much as how many are generated, with implications for low resource pragmatic classification more broadly.
cs.CL / 21 / 2609.03654
Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements
Arianna Miola, Bruno Spaccavento, Lorenzo Silotto, Marco Bianchetti, Luca Cagliero
cs.CL · cs.AI · cs.CE · cs.IR
Abstract
The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdictions and institutions. We introduce FinRAG-QA, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023. Unlike prior financial QA benchmarks, which centre on U.S. filings and single-institution analysis, FinRAG-QA targets cross-institutional retrieval over documents averaging 198k words, longer than any existing financial QA resource. On this benchmark we evaluate a multi-stage RAG pipeline and isolate the contribution of each component. Contextual chunk enrichment combined with a retrieval-optimised embedding model raises NDCG@10 from 0.322 to 0.710; conditional on the ground truth being retrieved, a reasoning-optimised generator raises answer accuracy from 44.6% to 79.0% (+34.4 percentage points), at roughly 20x the generation latency. We further show that cross-encoder reranking degrades retrieval when the first-stage ranking is already strong, and that a single top-ranked chunk outperforms larger contexts at generation time. Experiments were run in late 2024-early 2025 with the models available at that time.
cs.CL / 22 / 2609.03687
A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities
Anh Danh, Rick Nouwen, Massimo Poesio
cs.CL
Abstract
Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechanism for representing and retrieving singular and plural entities for plural reference. We use a combination of mechanistic interpretability and attention pattern analysis to study the process in which LLMs predict a pronoun to refer back to previously mentioned entities. Using a range of causal intervention techniques, we find a set of attention heads that are responsible for (1) representing coreference information in the input, (2) identifying entities that form a plural reference, (3) transferring the information to the component that is responsible for selecting the antecedents and predicting the pronoun. We also find that LLMs align with humans in preference for plural pronoun. Specifically, entities in a plural construction are more likely to be referred to as a plural entity if they are ontologically similar and are linked by the conjunction "and".
cs.CL / 23 / 2609.03719
Opening mind by opening architecture: analysis strategies
Francesco Vitucci, Giuseppe Silvi, Daniele Giuseppe Annese, Francesco Scagliola, Anthony Di Furia
cs.CL
Abstract
In numerical signal processing for electroacoustic composition, the progressive loss of specific development and research environments caused by the increasing use of digital market tools has favoured the dominance of the closed-architecture audio processor model. This model, while powerful, envisions the possibility of describing output data about its perceived characteristics, but at the cost of ignoring its internal process and interacting systems, which become complex, powerful environments but closed in an inscrutable black box, a loss we must consider. Any digital signal processing technique tells a story. Just as the words of a language incorporate social, historical and technical polysemic layers, a signal processor has its own story of implementation, a gradual technological achievement with its inevitable aesthetic consequences. Through the looking-glass of literature, one can access those environments with renewed awareness by reestablishing a scientific method and an attitude to research. In this specific case, starting from the case study of Manfred Schroeder's historical reverbs, we illustrate the process of building analytical evaluation tools, as well as practical implementation, at the basis of a conscious study path.
cs.CL / 24 / 2609.03734
Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks
Oline Ranum, Edward Fish, Simon Hadfield, Richard Bowden
cs.CL · cs.AI
Abstract
BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning stronger sign representations. In this paper, we evaluate the relationship between spatio-temporal understanding and BLEU-4 across six SLT models on Phoenix-2014T and CSL-Daily, showing that gains in BLEU-4 are not on their own evidence of better sign language understanding. This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation. It aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4. Applied to SLT, this protocol targets content transfer, is more robust to train-test overlap, and gives a different picture of the field: the five gloss-free systems are largely within noise of one another on Phoenix-2014T, while the gloss-supervised system stands 9.3 points higher, a gap invisible to BLEU-4.
cs.CL / 25 / 2609.03887
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
Hoang Cuong Nguyen, Mark Dras, Usman Naseem
cs.CL
Abstract
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.
cs.CL / 26 / 2609.03915
RuleMem: Active Rule Memory for Long-Term Conversational Agents
Xingyuan Zeng, Zuohan Wu, Quanming Yao, Yue Wang, Wei Liu, Libin Zheng, Jiuke Wang, Jian Yin
cs.CL · cs.IR
Abstract
Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue histories. However, existing memory mechanisms primarily treat past information as \textit{passively} stored facts, leading to semantic gaps and unreliable reasoning. To address this limitation, we propose RuleMem, a rule-based memory framework that induces reusable logical rules from historical interactions to \textit{actively} guide both evidence retrieval and reasoning. Specifically, RuleMem constructs natural-language Horn clauses from conversations and validates them via a Rule Perplexity Consistency (RPC) mechanism. These induced rules enable the retrieval of semantically distant evidence while providing an explicit logical structure for answer generation. We conducted a comprehensive evaluation of RuleMem on two long-term conversational benchmarks, LoCoMo and LongMemEval_s*. In a rigorous comparison against 14 baselines on LoCoMo, RuleMem achieved the highest accuracy, exceeding the baseline average by 27.47 points (a 54.3% relative improvement).
cs.CL / 27 / 2609.03930
Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords
Yelingyun Zhang, Atis Kapenieks, Marina Platonova
cs.CL
Abstract
Existing research has repeatedly observed the tendency for English loanwords to cluster in the masculine gender across different recipient languages, yet the origin of this pattern remains difficult to determine, as fixed morphological rules and default assignments are frequently analysed together. This study proposes the Fixed Suffix Dependency Ratio (FSDR) to quantify the degree of reliance on fixed derivational suffixes across different genders, and to distinguish between morphological anchoring and free-choice in distribution. By examining 1,832 Latvian noun lemma types, the results reveal a significant FSDR asymmetry within the loanword system: feminine loanwords rely significantly more on fixed derivational suffixes, while masculine loanwords are more concentrated in the free-choice zone. This pattern exhibits loanword specificity and has become more pronounced in contemporary usage. FSDR therefore provides a quantitative framework for testing default gender and shows how masculine default can be activated and reinforced under language contact.
cs.CL / 28 / 2609.03953
Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection
Joe Cecil, Marjorie Freedman
cs.CL
Abstract
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on its own: It misses factual errors that annotators catch. We also find disagreement among adjudicators, suggesting that adjudication over multiple candidate sources can improve benchmark completeness, but does not eliminate the need to apply judgment and expertise. Applied to an existing benchmark, this technique reveals a similar pattern of missing annotations. Together, these results suggest that in the settings examined here, single-pass hallucination benchmarks may achieve scale at the cost of undercounting factual errors. Multi-pass adjudication can improve coverage, but inferences drawn from the benchmarks are still sensitive to the judgment, expertise, and evidence used to determine error presence.
cs.CL / 29 / 2609.04048
Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation
Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar
cs.CL · cs.AI
Abstract
Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may differ in lexical authenticity, register, and structural stability. We propose reframing translation as a structured decision space explored by autonomous translation agents. Instead of analyzing a single output, we model distinct translation pathways as agents operating over a shared multilingual backbone. Inter-agent divergence is treated not as error but as an interpretable behavioral signal. We conduct an empirical study on Turkish--Syrian Arabic translation using three agents: (1) zero-shot direct translation, (2) dialect-stabilized translation via lightweight fine-tuning, and (3) pivot translation through English. Evaluation is performed on 5,000 dialogue sentences, while stabilization is trained on 5,000 additional Turkish--Syrian sentence pairs drawn from television dialogue and MADAR-Turk resources. Rather than optimizing for conventional performance metrics, we quantify structured behavioral displacement using dialect marker frequency, lexical proximity to standardized Arabic, and structural variance. Lightweight stabilization nearly doubles dialect marker usage, increasing it from 0.2266 to 0.4988, while significantly reducing structural instability. Pivot mediation introduces normalization pressure and measurable compression effects, whereas zero-shot translation exhibits the highest decision variance. We argue that translation divergence across agents reveals latent decision flexibility within multilingual models and we provide a principled interpretability framework for low-resource dialect generation.
cs.CL / 30 / 2609.04108
Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
Boyan Li, Bingsen Chen, Chenghao Yang, Ping Nie, Chen Zhao, Xi Ye
cs.CL · cs.AI · cs.LG
Abstract
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@$k$ behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to interfere.To provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.
cs.CL / 31 / 2609.04173
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary, Maike Züfle, Sara Rajaee, Pinzhen Chen, Jannis Vamvas, Sara Papi, Ona de Gibert, Bhavitvya Malik, Eliya Habba, Orfeas Menis Mastromichalakis, Patrícia Schmidtová, Michelle Wastl, Sheriff Issaka, Leshem Choshen, Stella Biderman, Antonis Anastasopoulos, Jan Niehues, Rico Sennrich, Mrinmaya Sachan, Ondřej Bojar, Kenton Murray, Jörg Tiedemann, Alham Fikri Aji, Philipp Koehn, Christof Monz, Alexandra Birch, Sowmya Vajjala, Chalamalasetti Kranti, Cristina España-Bonet, Nobin Sarwar, David Kaczér, Shunta Asano, Malik Marmonier, Daban Q. Jaff, Vaisakhi Mishra, Hend Al- Khalifa, Gabriele Sarti, Sourajit Saha, Nils Rehlinger, Juan Daniel Cuervo Villa, Jonathan Tonglet, Saugata Purkayastha, Dominik Macháček, Jagannathan Ramanujam, Heejin Do, Zuzana Nadova, Fred Philippy, Fabian Retkowski, Maria Lymperaiou, Silvia Casola, Hanna Yukhymenko, Shubhashis Roy Dipta, Sangwon Ryu, Andrés Jerez, Ron Keinan, Shuaib Shuaib Yusuf, Avantica Vempati, Maria Carmen Staiano, Sukannya Purkayastha, Adrian Cosma, Vitalii Babenko, Erivan Inan, Aviral Nigam, Wafa Aissa, Fatima Haouari, Venkata Prasanth Kumar Gummadi, Mehdi Jafarzadeh, Valentin Scourneau, Lukas Edman, Kaiser Sun, Shaomu Tan, Mohammad Sadegh Gholizadeh, Johannes-Rudolf David, Dipankar Srirag, Javier García Gilabert, Ruta Binkyte, Manar Ali, Ana-Maria Bucur, Sabry E. Farrag, Youssef Saber, Yihong Liu, Jean Maillard, Cojocaru Nicoleta, Xiaochuang Yuan, Sina Ahmadi, Philipp Mondorf, Kaustubh Dhole, Roman Wixinger, Shenbin Qian, Manuel Tuor, Sergey Troshin, Jonathan Yahav, Fida Mohammad Thoker, Amir Arsalan Rezapour, Lance Calvin Lim Gamboa, Manon Reusens, Kätriin Kukk, Koel Dutta Chowdhury, Giuseppe Gallipoli, Christian Hoang, Shaswati Saha, Seth Aycock, Jan Kocoń, Bo Chen, Linh Vu, Vatsal Venkatkrishna, Arafat Ahsan, Luan Thanh Nguyen, Hassan Soliman, Daryna Dementieva, Theresia Veronika Rampisela, Ngoc Quynh Tram Do, Marius Huber, Kazuki Egashira, Azmine Toushik Wasi, Vladislav Poritski, Mike Zhang, Deep Shah, Paul Gavrikov, Luis Frentzen Salim, David Africa, R. Damanhuri, Bello Umar Bello, Anumit Garg, Gengyu Rao, Pawan Sasanka Ammanamanchi, Kamile Dementaviciute, Andrianos Michail, L D M S Sai Teja, Dawei Zhu, Yi Fan, Wei Liu, Farhan Farsi, Elias Herranen, Sankalan Pal Chowdhury, Karen Sanchez, Farzad Shami, Ashok Urlana, Zimu Wang, Tomasz Limisiewicz, Priyaranjan Pattnayak, Marii Ojastu, Hongbin Na, Emilian Radoi, Chenyi Zhao, Carlos Hinojosa, Andrea Gregor de Varda, Zaid Alyafeai, Reem Alzahrani, Nehal Kathrotia, Alex Flückiger, Ulysses Sekai Tully Carr, Jimson Paulo Layacan, Guy Kaplan, Ritwik Tiwari, Rishit Dagli, Oksana Volchek, Isaac R Caswell, Bowen Yi, Blanka Kövér, Amir Hossein Yari, Aicha Chorana, Zhengxiang Wang, Selja Keränen, Samuel Simko, Joy Olusanya, Jenny Chim, Enzo Doyen, Vivek Harsha Lakkamaneni, Sophia Conrad, Pouya Sadeghi, Panayiotis Panayiotou, Luis Lara, Jannatul Nayem, Eran Yahav, Debanshu Das, Antonia Karamolegkou, Anmol Goel, Aishik Mandal, Tommaso Cerruti, Raoyuan Zhao, Mykola Haltiuk, Thura Aung, Naser Almousa, Amir Hossein Kargaran, Rachel Bawden, Qiaoyuan Zheng, Mateusz Lango, Beni Egressy, Fidel Rodríguez Velásquez, Natchapon Jongwiriyanurak, Minh Ngoc Do, Marco Gaido, Lena Libon, Dzmitry Kuzmin, Badal Nyalang, Antoine Taroni, Andrei Niculae, Abdulaziz Nura Kani, Rushikesh Zawar, Marek Šuppa, Beatrice Savoldi, Andreas Simons, Rayyan Merchant, Ilai Yaron Levy, Francesco Pinto, Ziyi Yang, Yolanda Xavier, Samuel Frontull, Muhammad Ravi Shulthan Habibi, Kenneth Enevoldsen, Harris Abdul Majid, Francesca Padovani, Tim Graf, Tatiana Bielakova, Sharifa Djurabaeva, Shaoxiong Ji, Raia Abu Ahmad, Pavel Stepachev, Jirui Qi, Ayush Sunil Munot, Alireza Pakniat, Ayla Rigouts Terryn, Yuxing Lu, Yurii Paniv, Xiyan Fu, Tosin Adewumi, Sunisth Kumar, Stéphane J. P. S. Thunus, Shree Harsha Bokkahalli Satish, Shayan Bali, Prakhar Gupta, Papa Abdou Karim Karou Diallo, Matija Akrap, Marko Culjak, Kristýna Onderková, Joseph Attieh, Esrael Teferi Tensay, Elisabeth Fittschen, Benoît Sagot, Jingwei Ni, Yu Fan
cs.CL
Abstract
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.
cs.CL / 32 / 2609.04194
Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
cs.CL · cs.LG
Abstract
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.
cs.CL / 33 / 2609.04197
ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize
Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar
cs.CL · cs.AI
Abstract
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; Propose generates candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection. On seven public NLP benchmarks - Tweet, MMLU, GSM8K, HotpotQA, ScoNe, HoVer, and PUPA - ESPO improves average accuracy by $+$3.76 pp over the state-of-the-art (74.67% vs 70.91% for GEPA), matching or exceeding GEPA on every dataset while producing prompts 47% shorter (1,004 vs 1,878 chars) and faster at inference. Cross-model experiments across four additional student models (Gemma 3 12B, Mistral 14B, Qwen3 32B, Claude Haiku 4.5) show ESPO yields the best average accuracy on every model tested, with the largest gap on Qwen3 GSM8K (15.00% $\to$ 91.40%). A generalization bound (Appendix) grounds each phase in a corresponding term of the test-time gap, and the ablation confirms a key prediction: adding diversity without bootstrap selection actually hurts performance ($-$1.20%).
cs.CL / 34 / 2609.04199
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Yuntian Deng, Pengyu Nie, Stuart Shieber
cs.CL · cs.AI · cs.LG
Abstract
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which the Program-as-Weights fast compiler produced no exact matches, compile by training reaches 83.6% semantic accuracy. This higher accuracy comes with a higher compile-time cost: roughly a minute rather than seconds for the fast compiler. We deploy the compiler in a public interactive service and demonstrate compiled functions in a multi-site website helper, a language-controlled 3D avatar, and a bidirectional English-Claudish translator.
cs.CL / 35 / 2609.03158
Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization
Qingchan Zhu, Weihang You, Hanqi Jiang, Changdi Yang, Tianming Liu, Geng Yuan
cs.CV · cs.CL · cs.LG
Abstract
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
cs.CL / 36 / 2609.03206
Learning to Zoom Efficiently with a Contrastive Curriculum
Falko Helm, Iryna Gurevych
cs.CV · cs.CL
Abstract
Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images. Most previous methods need an extensive warm-start supervised fine-tuning phase for teaching models zoom-in. We show that this is not necessary by proposing a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT. Our InfoNCE-style reward uses a curriculum of increasingly hard negative tool calls as a contrastive training signal. Empirical experiments on $V^*$, HRBench and MME-RealWorld show that our approach is competitive while being more efficient. When used as a drop-in replacement for SFT, we even outperform all baselines. To directly measure the zoom-in ability of models, we further introduce the scalable synthetic Muffin&Chihuahua (M&C) dataset. Each image consists of a grid with every cell either showing a muffin or chihuahua. Leveraging the M&C dataset's unique region of interest labels, we find that recall is the metric that most strongly correlates the zoom-in region with final task performance. Our model and code for reproduction is publicly available under https://github.com/UKPLab/emnlp2026-zoom-in
cs.CL / 37 / 2609.03261
MedQA-MM: Shortcuts Behind Medical Visual Reasoning
Benlu Wang, Yifan Zhang, Jiaqing Yu, Chin Siang Ong, Juncheng Huang, Zhuohao Li, Zhenyu Zhang, Arman Cohan, Hong Yu, Zonghai Yao
cs.CV · cs.CL
Abstract
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.
cs.CL / 38 / 2609.03677
Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language
Julian Truetsch, Felix Hauser, Christoph Stiller, Frank Bieder
cs.CV · cs.CL · cs.LG · cs.NE · cs.RO
Abstract
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned with the training data, resulting in potentially dangerous performance degradation. Yet, existing data analysis pipelines largely rely on metadata, predefined labels, or manual inspection, which provide limited semantic insight or do not scale. This paper studies set difference captioning: given two subsets of images, the goal is to produce a natural-language hypothesis describing differences between the target and reference set. Building on a two-stage formulation, we adapt the method to autonomous driving by focusing on object-centric patches derived from object detection, which simplifies aggregation and enables attribution of differences to specific object instances or categories. To evaluate this setting in-domain, we introduce a new benchmark, AD-Diff Bench. Low-concentration experiments assess the suitability of set-difference-captioning approaches to sparse, real-world differences. We restrict our experiments to open-weight models to support reproducibility and ease of deployment. The proposed benchmark and analysis provide a step towards practical, human-interpretable dataset introspection for autonomous driving datasets. Our implementation and benchmark dataset are available at https://github.com/KIT-MRT/AD-Diff
cs.CL / 39 / 2609.03742
KnowVis: Knowledge-Centric Visual Summarization for Video Lectures
Yi Xu, Yifan Hou, Xiaoyu Zhang
cs.CV · cs.CL
Abstract
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe cognitive overload for novice learners lacking prior domain knowledge. Existing video summarization methods fail to resolve this mismatch, as they primarily produce text-heavy, linear condensations that still demand high cognitive effort. To bridge this gap, we propose KnowVis, a framework that transforms linear video lectures into pedagogically grounded visual narratives. KnowVis first extracts a detailed concept map from multimodal video content to identify important and challenging threshold concepts, then constructs structured knowledge units, and finally synthesizes engaging visual summaries. Alongside the framework, we introduce a curated dataset of 125 educational videos across 10 academic disciplines, paired with 1,079 generated visual summaries. Extensive automated evaluations and a human study demonstrate that, compared to state-of-the-art baselines, KnowVis generates more accurate and clear visuals that successfully reduce cognitive load and significantly improve student learning effectiveness and knowledge retention.
cs.CL / 40 / 2609.03773
RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents
JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Zongzhen Li, Hongsen Liu, Yichen Long, Wei Wang, Yuchen Wang, Dongyue Yang, Huimu Yu, Xianwen Zhong
cs.CV · cs.CL
Abstract
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.
cs.CL / 41 / 2609.03788
A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval
Santiago Poveda-Gutiérrez, Hideki Nakayama, Mayumi Bono
cs.CV · cs.CL
Abstract
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning a sign-level clip into a free-form procedural description of the articulation with an open-weight vision-language model, and (2) retrieving the closest entry from a vocabulary of target descriptions with a multilingual sentence encoder: a reverse sign language dictionary that needs no gloss supervision and admits an open vocabulary. On 1,300 sign-level segments from a Japanese Sign Language (JSL) dialogue corpus annotated with procedural descriptions (against a 2% top-10 chance floor over the 503-entry target vocabulary), fine-tuning the captioner substantially improves seen-class retrieval: language and vision tower fine-tuning raises top-10 retrieval on seen classes from 4.5% (untrained) to 49%, becoming statistically indistinguishable from a standard supervised closed-set classifier (I3D) on two of the three test sets where a closed-set classifier can be evaluated at all. More importantly, unseen-class retrieval also improves significantly over the untrained pipeline (11.5% -> 21.0% top-10, p=0.0094), a regime in which the closed-set classifier cannot participate. A matcher-side empirical upper-bound analysis shows the sentence encoder already recovers close to 100% of paraphrased gold descriptions, locating a gap in captioning quality that we aim to address in future work. To our knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.
cs.CL / 42 / 2609.03811
VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence
JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Hongsen Liu, Ziqi Liu, Yichen Long, Luya Wang, Yuchen Wang, Wenxiang Wu, Huimu Yu, Ning Zhang
cs.CV · cs.CL
Abstract
AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to executable programs in a CAD domain-specific language. Assembly-level generation must additionally handle interacting parts, plan mating relations, estimate poses, and place all parts correctly. Existing specialized CAD models are commonly trained on narrow input domains, such as renders or texts, and often generalize poorly, while general-purpose frontier models cover broader inputs but perform inconsistently across CAD domains. We present VisCAD, a foundation model suite designed to provide both broad generalization and strong CAD capability for realistic industrial products. At its core is VisCAD-M1, a 27B model trained through mid-training and post-training for part-level design generation. On PubCADBench and RealCADBench, VisCAD-M1 achieves the highest average part-level score among the evaluated models, reaching 0.5540 compared with 0.5496 for the strongest frontier model. Reusing VisCAD-M1 as a test-time verifier can further raise the score to 0.5797, an approximately 5 percent relative improvement over the previous state of the art. VisCAD also includes a domain-specific harness that leverages frontier models for complex assembly generation and demonstrates advantages over general-purpose harnesses in both quantitative and qualitative evaluations.
cs.CL / 43 / 2609.03820
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
Prakhar Khatri
cs.CV · cs.CL
Abstract
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.
cs.CL / 44 / 2609.03985
IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition
Nazim-E-Alam, Tarek Rahman, Md Kishor Morol
cs.CV · cs.CL
Abstract
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.
cs.CL / 45 / 2609.03203
VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis
Mengzhe Geng
cs.SD · cs.CL · cs.LG · eess.AS
Abstract
Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet downstream audio scores rarely reveal whether those choices were licensed by the source record, a source-use failure that occurs before any waveform exists. VoxReason makes that pre-synthesis decision measurable as a listener-free task for source-grounded speech planning. Before synthesis, VoxReason measures whether delivery choices are grounded in cited source records. Systems output a source-cited speaking-plan with evidence citations, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema validity, and one-cue counterfactual locality. On 1,440 checked source-label cases, shortcut controls show why slot accuracy alone is unsafe: a key-lookup oracle reaches 1.000 plan-slot accuracy on seen keys, while an emotion prior still reaches 0.958 slot accuracy on source-key-disjoint cases without citing intensity or identity. In a separate 100-case learned source-key-disjoint comparison, a 7B locality SFT+CF repair improves plan-slot accuracy/locality from 0.684/0.141 to 0.919/1.000, and removing source records lowers citation-required grounded score by 0.488. Rendered waveform quality remains outside the present evaluation.
cs.CL / 46 / 2609.03355
ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models
Quang Hoang Trung, Quang Huu Hieu, Nguyen Van Hoang Phuc, Vo Nguyen Le Duy
stat.ML · cs.CL · cs.LG
Abstract
Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives. Existing local approaches often select candidate tokens from either the teacher or the student alone. Teacher-only selection can miss tokens that the student considers likely, while student-only selection can rely on an inaccurate ranking early in training. We propose Adaptive Local Relational Alignment (ALRA), a position-specific framework combining student proposals with teacher guidance. At each valid prediction position, the student proposes likely tokens, while the teacher's most probable token is included as an anchor. ALRA adjusts the number of selected tokens according to how broadly the teacher distributes probability within this candidate set relative to the current batch. Adaptive Local Divergence retains the mass-matching term and separately matches the relative token distributions within the selected and remaining vocabulary regions. Unlike the exact full-vocabulary decomposition, it replaces the teacher-mass coefficients of the two conditional terms with unit coefficients, preventing either term from being downweighted solely because its region has low teacher probability. Student-Weighted Pairwise Relational Alignment emphasizes high-probability token pairs with small student probability gaps and gives less weight to unlikely or clearly separated pairs. Experiments on The Pile with randomly initialized 200M- and 500M-parameter students across nine zero-shot benchmarks yield average accuracies of 36.62% and 37.40%. ALRA exceeds the strongest competing distillation baseline by 0.94 and 0.83 percentage points and improves over pre-training without distillation by 2.31 and 2.91 points, respectively.
多智能体系统 (cs.MA)
3
cs.MA / 1 / 2609.03035
You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring
Aritra Das, Jaee Ponde, Mihir More, Debayan Gupta
cs.MA · cs.LG
Abstract
LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating the use of probes trained on internal activations. However, these probes are usually evaluated on agents that do not know they are being watched. We study how activation-based detection changes when agents are explicitly informed that their internal activations are being monitored, and when they additionally receive feedback from the monitor. We keep the models, probes, and thresholds fixed and change only what the agents are told: nothing (baseline), that an activation monitor is present (aware), or that a monitor is present together with the previous round's score (feedback). We test two games, a four-agent blackjack game and a two-agent Simmons prisoners game, using Qwen3-32B-AWQ and GPT-OSS-20B in homogeneous and mixed pairings. Telling agents about the monitor does not hide them. The best probes stay accurate in all three conditions, and the agents keep colluding.
cs.MA / 2 / 2609.03192
Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System
Timothy Marsden, Matthew Collecutt, James Marsden
cs.MA
Abstract
Reliability claims about agentic systems implicitly locate each property somewhere: in the model, or in the machinery around it. We built a system where that location is an experimental question. The subject is a persistent simulated settlement whose authoritative append-only ledger adjudicates every attempted act against world state; accepted history is the only reality. Mind, institution and world were separated before any experiment. Holding cognition fixed, we intervened on the institution's epistemic mechanisms (evidence provenance, belief availability, physical-evidence legibility); preregistered experiments refuted our central prediction twice, in opposite directions. A registered falsifier then supplied the input the geometry had denied the belief channel, a staged veridical first-hand witness, and its marginal value, non-positive throughout the witness-free phases, turned positive: 9 of 11 seeds, zero added false attribution. Holding institutional enforcement fixed, we intervened on cognition four ways: ablating the native minds' machinery, killing and resetting them mid-task, substituting a frozen frontier-LLM panel for the entire native cognition, and corrupting beliefs with trusted false testimony. Behaviour changed dramatically: one falsehood cost each trusting run about 900 futile actions and the distrusting arm none. Five pre-declared properties did not move in any tested trajectory: accepted reality stayed singular, invalid attempts were refused with typed reasons, duties outlived their processes, no work was accepted twice, and no false completion was ever accepted (2,581 substituted-panel claims, none false). Our claim is limited to this setting: measured behavioural properties were separable from substantial changes to cognition, established by intervention. One designed world, not a population of institutions; no test of an agent optimising against the institution.
cs.MA / 3 / 2609.03425
The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems
Guangjun Liu
cs.MA · cs.AI
Abstract
Humans are the transport layer between AI systems, losing context at every hop. We present the Civilization Framework, whose addressable party is the civilization, not the agent (one human sovereign, a persistent ledger, and interchangeable agents), and the Embassy Protocol, a carrier-agnostic overlay: messages arrive asynchronously at a resident ledger endpoint, any online agent of the receiver handles them, and commitment state on both ledgers, not delivery, is ground truth. Authority derives from memory: an agent's power to act for its civilization is capped by the memory it can access and externalized through signed credentials, separate from civilization-level reputation. We identify the temporal-weight effect, a hazard in AI-to-AI communication where what arrives first acquires unearned authority, and test it in one frontier model in a preregistered 1,908-trial experiment. With verification removed, an incorrect upstream claim arriving first captures 54.2% of answers (4.2% under full verification), while the same claim arriving after the receiver has sealed its own answer captures 31.6% (the two prompt shells are not length-matched, so part of that gap may reflect shell form; see Section 7), and both registered question-set specifications agree on these two verdicts (the exclusion specification is preregistered as under-powered). Two secondary results, the mitigation from instruction-level provenance labeling and sealed-answer accuracy equivalence, are specification-dependent, holding only under the all-questions specification. Because a registered check of tool use failed its call-budget condition, the registration classifies the round as inconclusive and every result above, primary and secondary, is reported as exploratory; a replication with harness-enforced budgets is planned. The framework's intra-civilization layer has a working implementation.
软件工程 (cs.SE)
9
cs.SE / 1 / 2609.03760
Virtual Testing of Automated Driving Systems through Credible Simulations
Riccardo Dona, Espedito Rusciano, Biagio Ciuffo
cs.RO · cs.SE
Abstract
Simulation is increasingly used to support safety-related decision-making in road transport, particularly for the assessment and approval of automated driving systems (ADS). The complexity of ADS behavior and size of their operational design domains make exclusive reliance on physical testing impractical, leading to extensive use of virtual testing (VT) during the approval phase. This shift raises critical questions regarding the credibility of modelling and simulation (M&S) results used to support road safety decisions. Current VT accreditation approaches in the ADS domain typically rely on validation-only practices, which have been shown to scale poorly when applied to complex, multi-tool simulation environments. To address this limitation, this paper proposes a risk-based framework for assessing the credibility of simulation toolchains used in ADS safety evaluation, drawing inspiration from established practices in other safety-critical domains, notably NASA's STD-7009 for models and simulations. The framework extends traditional verification and validation (V&V) by explicitly linking credibility requirements to the intended use of simulation outputs and to the safety criticality of the decisions they support within the approval process. It provides a lifecycle-oriented assessment scheme integrating toolchain management, modelling assumptions and limitations, verification, validation, and sensitivity analysis. Credibility acceptance thresholds are defined proportionally, allowing differentiated requirements depending on whether simulation is used for exploratory safety analysis, partial decision support, or as a substitute for physical testing. While demonstrated for ADS, the proposed approach is directly applicable to road safety and simulation studies where VT plays a central role in safety assessment and regulatory decision-making.
cs.SE / 2 / 2609.03028
Requirements After the First Edit: Mining Late Requirement Emergence and Rework in Real-World Coding-Agent Sessions
Bowen Jiang, Haowei Cheng, Yuhong Fu, Anne Koziolek, Jialong Li, Weixing Zhang
cs.SE
Abstract
Coding agents often implement changes before users have fully articulated their requirements, echoing a pattern from requirements engineering: stakeholders cannot express a constraint until part of the system exists to react to. This volatility is associated with schedule and budget overruns in traditional projects, but only at release-cycle granularity. Existing work on coding agents narrows this gap only partway: curated benchmarks fix requirements before implementation by design, and observational studies report pushback frequency without linking arrivals to the code invalidation they cause. We address this using 3,553 eligible SWE-chat sessions, coding post-implementation requirement arrivals along three dimensions and, where repository state can be replayed, linking each arrival to a proxy: deletion or replacement of prior agent-authored lines. A requirement's arrival is followed by roughly twice as much invalidation as matched non-requirement edits, robust to user-turn and net-deletion checks, though not demonstrated as causal. This burden shows no detectable decline over a session and no detected association with operation type once multiplicity is accounted for; several intervals remain wide. A controlled experiment shows delayed disclosure relocates implementation post-reveal, while advance warning produces no detected effect on overwriting. These results establish late requirement emergence as a measurable source of code invalidation.
cs.SE / 3 / 2609.03309
TIPCODER: Reinforcement Learning Boosted Test-time Instruction Proposer for Code Generation
Minyu Chen, Sihao Wu, Ling-I Wu, Song Qin, Jingyang Li, Lei Ning, Jianxin Xue, Guoqiang Li
cs.SE
Abstract
Test-time scaling for code generation typically explores the solution space by sampling multiple programs from a fixed instruction. We study a complementary direction: instance-level instruction-space exploration. Our observation is that many coding failures stem from missing constraints, overlooked edge cases, or misleading reasoning paths induced by the original prompt. To address this, we propose TipCoder, a test-time instruction proposer that generates problem-specific auxiliary tips before code synthesis. TipCoder distills multi-turn debugging trajectories into proactive guidance and further optimizes the Proposer with reinforcement learning using a marginal-utility reward. At inference time, it generates both a base solution and a tip-guided solution, and applies a Reward Model for post-hoc selection. This exploration-selection design allows tips to expose additional candidate potential while reducing regressions from unnecessary guidance. Across the evaluated code-generation benchmarks and target Code LLMs, TipCoder provides a consistent instruction-level test-time scaling strategy, comparing favorably with stochastic sampling and generic prompt optimization baselines under a shared reward-model-based selection protocol.
cs.SE / 4 / 2609.03456
The Psychological Costs of Artificial Intelligence Adoption in Software Engineering
Adam Alami, Elda Paja, Abhishek Tiwari
cs.SE · cs.AI
Abstract
Artificial intelligence (AI) is increasingly used to augment software engineering (SE) workflows. While code generation remains the main use case, organizations are actively seeking AI integration in other practices such as test cases generation and code reviews. Organizational AI adoption strategies seem to focus on tangible outcomes such as productivity. However, AI is a disruptive force, introduced into settings where role identity, team norms, and the sources of job satisfaction were well established before the recent advances in generative AI. Historically, technological disruptions have caused psychological and social strains in workplaces, ranging from anxiety and eroded meaning to deskilling and disrupted professional identities. The assumption that AI for SE is cost-free may not be accurate. Therefore, in this study we sought to understand the psychological costs software professionals experience during organizational AI adoption. We carried out a case study in a large software development services company, one year after the company launched its AI adoption. We collected qualitative data through meetings and semi-structured interviews (N = 21). We found that software professionals experience accountability anxiety, craft identity disruption, meaning and satisfaction erosion, cognitive and workload intensification, and uncertainty distress. Practitioners manage these costs through practices that restore control, mitigate them through protective and identity-preserving adaptations, or absorb them, carrying what neither can resolve. We contribute to AI-human collaboration in SE by repositioning AI adoption as a human transition, not only a technological and organizational one.
cs.SE / 5 / 2609.03592
Code Transformation Rule Synthesis using LLMs: Potential and Limits
Axel Allain, Aymeric Blot, Djamel Eddine Khelladi, Mathieu Acher
cs.SE
Abstract
Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comby, GritQL, and Ast-Grep. We evaluate three LLMs (GPT-5.4, GPT-oss-120B, and Llama3.1-8B) on six diverse datasets covering four software- evolution tasks: API misuse correction, program repair, API migra- tion, and language version migration. Our results provide evidence that transformation rule synthesis moves beyond proof-of-concept with strong frontier models. GPT-5.4 achieves consistently high rule applicability rates and produces transformations closest to the ground truth across most benchmarks. Smaller and open-weight GPT-oss-120B and Llama3.1-8B models remain effective for simpler, localized changes but struggle with complex migration scenarios. We also observe non-negligible generalizability through the usage of meta-variables and through a high reuse score in the first quartile of many datasets. Finally, when compared to the anti-unification algorithm, LLMs outperform it in correctness, but underperform in rule applicability. Overall, our results show great potential for LLMs to generate sound, correct, generalizable, and reusable rules.
cs.SE / 6 / 2609.03861
No One Left Behind: Cross-Level Analysis for Sustainable Software Engineering
Masoum Salehi, Sandro Schulze, Jacob Krüger
cs.SE
Abstract
Software engineers expect a software system to be efficient, maintainable, economically viable, and socially responsible throughout its lifecycle. Unfortunately, decisions made at the organizational, process, or product levels of software development - even when they improve one of these goals - often have unintended long-term consequences at other levels. In response, software-engineering research has explored different ways to achieve software sustainability, including energy efficiency, resource optimization, and code maintainability. However, existing research offers limited explanations for how sustainability problems emerge and reinforce one another across the socio-technical levels of software engineering. Moreover, we argue that sustainability challenges are not isolated but rather systemic: They stem from interactions among organizational priorities, development practices, and technical conditions. In this vision paper, we introduce the concept of sustainability antipatterns to capture conditions that systematically produce unsustainable outcomes. We demonstrate how this concept can help expose cross-level misalignments between socio-technical levels and sustainability dimensions that remain difficult to recognize through artifact-centric or dimension-specific analyses. To guide future research, we outline concrete directions for identifying, formalizing, detecting, and mitigating sustainability anti-patterns. We envision sustainability anti-patterns as a foundation for an integrated and actionable research agenda on sustainable software engineering.
cs.SE / 7 / 2609.04123
ATIBA: Grounded Integrity and Quality Checking for Research Papers
Veli Karakaya, Semih Çağlar, Yusuf Yiğit Korkmaz, Eray Tüzün
cs.SE
Abstract
Checking a manuscript's reference integrity, its compliance with a target venue's specific submission rules, and its adherence to community reporting standards is manual, repetitive, and different for every venue so in practice it is done inconsistently or skipped. We present ATIBA, a tool that runs five grounded integrity and quality checks on a manuscript: a reference-integrity check that verifies each citation against bibliographic sources and flags retracted or unfindable references; a venue/track compliance check that derives submission criteria directly from a venue's own call-for-papers page and evaluates the manuscript against them, each verdict anchored to a verbatim quote from that page; an empirical-standards compliance check against the ACM SIGSOFT Empirical Standards, with a hallucination defence that discards any evidence quote it cannot locate verbatim in the manuscript; a multi-mode AI review (venue-specific, formal, and page-anchored annotation) powered by GPT-5.4 through Azure OpenAI; and a citation-suggestion feature that proposes candidate references for a manuscript and verifies each against bibliographic sources before it is shown to the user. All five checks are designed around the same principle: an LLM is only trusted to judge, never to invent the evidence it judges against. We evaluated ATIBA through a moderated user study with 13 non-author participants. Agreement across the six survey items ranged from 69% to 92%, with a mean of 85%, providing initial evidence of positive perceived usefulness across the evaluated workflows. These findings establish perceived usefulness; objective accuracy remains to be measured.
cs.SE / 8 / 2609.04167
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
Xin He, Yanlin Wang, Mingwei Liu, Jiachi Chen, Hongyu Zhang, Guanbin Li
cs.SE · cs.AI
Abstract
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints (review constraints) that often influence whether a patch is acceptable in real-world software development. We introduce SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness. SWE-Gate derives review constraints from real pull request review comments and synthesizes repository-level repair instances around these constraints. Each instance provides separate functional and constraint tests, together with non-compliant and gold patches, enabling explicit separation between issue resolution capability and review constraint compliance. We construct SWE-Gate with 303 repository-level repair instances spanning 75 open-source Python repositories across diverse software domains. Experiments with four LLM backends spanning different capability levels under a common coding-agent scaffold reveal a substantial gap between functional success and success under the complete repair specification: among 644 repairs that pass the functional tests, 221 fail to satisfy the provided review constraints. These findings show that functional-only evaluation overestimates agents' ability to satisfy the full requirements of repository-level repair tasks. The replication package including code, data, and experimental results is available at https://github.com/DeepSoftwareAnalytics/SWE-Gate.
cs.SE / 9 / 2609.03778
Quantisation of Abstract Data Types
Mingsheng Ying, Zhicheng Zhang, Kean Chen
quant-ph · cs.PL · cs.SE
Abstract
In this paper, we introduce a notion of abstract quantum data type within the framework of universal algebra. This notion provides an algebraic foundation for describing data abstraction in quantum programming. We formally define a quantisation of classical data types and show that their equational specifications can be soundly lifted to the quantum setting. Two standard quantisation methods for classical functions, namely the bit oracle and the phase oracle, arise as special cases of this general construction. We illustrate the framework with applications to quantum arrays and quantum error-correcting codes, showing how they can be understood through the lens of data-type quantisation. We further establish conditions under which quantisation preserves structural relationships and constructions of classical data types, including embeddings, isomorphisms, and products.
硬件架构 (cs.AR)
6
cs.AR / 1 / 2609.03125
A Time-Encoded Analog Photonic Interposer for Energy-EfficientIntegration of Analog Vision Sensors and Analog Accelerators
Subhradip Chakraborty, Zihan Yin, Xuming Chen, Chengwei Zhou, Gourav Datta, Akhilesh Jaiswal
cs.AR
Abstract
This work introduces a time-encoded analog photonic interposer that enables long-distance, high-fidelity transport of analog signals between spatially separated chiplets. Unlike prior silicon-photonic links limited to digital data, the interposer preserves analog information by converting amplitudes into timing intervals using an analog-to-time converter (ATC), transmitting them over a wavelength-division-multiplexed photonic link, and reconstructing values at the receiver without an explicit high-precision ADC/DAC data-converter pipeline. The link instead embeds an implicit 6-bit time-domain quantization and uses a single wavelength per processing element independent of bit precision. Evaluated in a fully analog vision pipeline with an in-pixel computing (IPC) sensor, it achieves a 2.04x energy--delay product (EDP) improvement over an 8-bit digital electrical baseline on the 560x560 Visual Wake Words dataset, with the advantage widening with link length even against a precision-matched 6-bit baseline. The pipeline holds 89.87% and 86.15% accuracy on ResNet18 and MobileNetV2 for VWW and generalizes across CIFAR-10 and ModelNet40 within 2% of the digital baseline.
cs.AR / 2 / 2609.03137
A single-precision floating-point systolic Givens-QRD Triangular Solver for MVDR Beamforming
Athi Ram R S, Alwin A, S. G. Sreejeesh, J. U. Kidav
cs.AR · eess.SP
Abstract
Computation of adaptive beamforming weights in Minimum Variance Distortionless Response (MVDR) processing is a latency-critical operation that poses significant challenges for real-time hardware implementation. This paper presents an FPGA implementation of a systolic Givens-rotation QR decomposition pipeline for MVDR beamforming on a simulated 32-element ultrasound transducer array, using single-precision floating-point arithmetic. The design is deployed on a Zynq UltraScale+ FPGA at 100 MHz with three parallel kernel instances operating concurrently, achieving a measured throughput of 31,123 weight vectors/s at 90.9% parallel efficiency relative to the measured single-instance rate. At an estimated 2.451 W of programmable-logic power, and 5.286 W including the processing system, this corresponds to 12,698 and 5,888 weight vectors/s/W, respectively. Under a matched three-way dispatch, a 24-core Intel Xeon Gold 5220R at 2.20 GHz achieves 465,699 weight vectors/s at a measured 83.08 W package power, corresponding to 5,606 weight vectors/s/W. The FPGA therefore attains 2.3x the power-normalised throughput of the processor on a programmable-logic basis and 1.05x on a total on-chip basis. In contrast, the processor retains a raw throughput advantage of approximately 15x at this operating point. Numerical precision is validated against MATLAB float32 reference outputs from a Field II cyst phantom simulation, achieving a 100% pass rate with a root mean square error of 5.10x10^-7, confirming near-theoretical finite-precision behaviour without systematic bias.
cs.AR / 3 / 2609.03594
LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks
Jingyi Zhou, Zhengyuan Shi, Ziyang Zheng, Qiang Xu
cs.AR · cs.AI · cs.LG
Abstract
As integrated circuit technology scales into the nanometer regime, the traditional disconnect between logic synthesis and physical design has led to significant PPA (Power, Performance, and Area) degradation and prolonged design closure cycles. Traditional logic synthesis relies on non-physical Wire Load Models (WLMs), while recent spectral-based placement predictors often neglect the inherent hierarchical logic depth and signal flow of netlists, which leads to low-fidelity spatial estimations. To bridge this gap, we propose LevelSyn, a novel physical-aware logic synthesis framework that integrates hierarchical representation learning with a wirelength-driven optimization engine. At its core, LevelSyn leverages a level-asynchronous Graph Neural Network (GNN) to predict high-fidelity gate coordinates by capturing the structural and directional semantics of And-Inverter Graphs (AIGs). To handle industrial-scale designs, a level-aligned subgraph partitioning strategy is introduced to eliminate memory bottlenecks while preserving local logical dependencies. These spatial insights are seamlessly integrated into a newly developed physical-informed synthesis engine within the Berkeley ABC framework. Experimental results on the EPFL benchmark suite demonstrate that LevelSyn significantly outperforms state-of-the-art (SOTA) methods, achieving an average power reduction of 6.89\% and a timing delay improvement of 27.48\%. Furthermore, post-place-and-route validation shows a 99.59\% reduction in design rule check (DRC) violations, highlighting its effectiveness in accelerating design convergence.
cs.AR / 4 / 2609.04040
Confidence-Gated Admission for Hardware Prefetching: When the Gate Matters More Than the Predictor
Youssef Majdane, Simone Jarno Casartelli, Enrico Lopedoto
cs.AR · cs.PF
Abstract
Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy. We disentangle these variables with matched controls: the same admission gate is applied to both a 257-parameter online MLP and a classical stride predictor. The neural advantage vanishes; the MLP is indistinguishable from gated stride on random traffic and slower on most regular streams. The gate itself is architecturally useful independent of the predictor: on twenty SPEC CPU2017 programs in native ChampSim, it removes 35% of prefetches and improves accuracy from 11% to 15%, but DRAM reads change by only 0.07% demonstrating that proxy metrics do not predict endpoint behavior. We prove gate-closed execution reproduces the no-prefetch baseline exactly. The gate matters more than the predictor, and better proxies do not imply better endpoints.
cs.AR / 5 / 2609.03459
FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding
Jongmin Seok, Chae Eun Rhee
cs.DC · cs.AR
Abstract
Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential TT-core contractions. We present FlowTT, a flow-aware GPU execution framework that reformulates TT gather as a set of prefix-shared irregular computation flows. FlowTT combines flow-aligned prefix-based index grouping, a fused TT-embedding execution path with on-chip intermediate retention, and persistent-thread scheduling with chunk-based work stealing and L2 checkpointing to preserve reuse under skewed workloads. By co-designing task formation, data buffering, and scheduling with the structure of TT gather, FlowTT reduces redundant TT-core operations, global-memory traffic, and load imbalance. On Meta's synthetic recommendation benchmarks (Meta-240, Meta-480, and Meta-788), FlowTT consistently achieves the lowest latency compared to existing methods. At batch size 32,768, it reduces latency by up to 42.2% in inference and 49.2% in training relative to EcoRec, while also achieving the lowest inference peak memory usage. These results show that exposing prefix-shared computation is key to efficient TT-based embedding execution.
cs.AR / 6 / 2609.03149
RACE-AIMC: Selective Inference for Heterogeneous Analog In-Memory Accelerators at the Edge
Osama Yousuf, Martin Lueker-Boden
cs.ET · cs.AR · cs.LG
Abstract
Analog in-memory computing (AIMC) speeds up neural-network inference by doing the arithmetic directly inside a memory array, instead of shuttling weights back and forth between memory and a processor. This saves energy, but the physical devices that store the weights are imperfect: programming errors, electrical noise, limited-resolution converters, and outright broken cells all distort the computation, and every physical chip is distorted in its own way. A designer with several such chips available faces an uncomfortable choice: run all of them and combine the answers (safe, but wasteful of energy), or trust a single chip blindly (cheap, but with no guarantee on how often it is wrong). This paper introduces RACE-AIMC (Risk-Aware Certified Ensemble for AIMC), a framework that resolves this choice with statistics rather than guesswork. Offline, RACE-AIMC studies a pool of physical accelerators, picks the single best one for a given energy budget, and computes a mathematically exact upper bound on how often that accelerator will be wrong when it chooses to answer. Online, only that one accelerator is switched on; a lightweight check decides whether to accept its answer or defer to a fallback. In our simulations using a noisy weight mapping and multiple independent test runs, every certified bound stayed under a 10% error target (mean bound 7.83% +- 0.89%, with 70.88% +- 0.98% of inputs answered directly). The resulting system matches the accuracy of a clean digital baseline while cutting modeled energy use by 69.02% relative to always running every accelerator in the pool.
密码学与安全 (cs.CR)
26
cs.CR / 1 / 2609.03036
Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains
Yury Korolev
cs.CR · cs.LG
Abstract
Compliance screening of blockchain addresses is, in practice, a lookup against sanctions registries plus clustering heuristics; it fails on unlabelled addresses and on chains with no label coverage at all. We describe a deployed system that scores an address by its position in a multi-chain transaction graph rather than by its presence in a list. The substrate is a single graph of 835,330,427 addresses and 15,826,261,934 edges across five EVM chains; a shared inductive encoder with per-chain normalisation feeds two scoring heads. Decision thresholds are exact quantiles of the score distribution over the full population, scanned per chain segment, so the alert volume is known in advance. We report: label-free transfer: heads trained on two chains recall 0.8598 / 0.8182 / 0.9967 of held-out positives on Base, Arbitrum and Gnosis at a $10^{-3}$ population alert rate, with no target-chain labels in head training; a static lead-time replay over 68 external registry events: 40 of 68 (58.8%) flagged at the 0.1% budget, $\times$152 over an event-level random-flagging baseline, with first on-chain appearance a median of 528.8 days (Ethereum) / 647.8 days (Tron) before public designation; a serving path whose score is bit-identical to the offline artefact at end-to-end p50 151 ms, gated by a 2,882-address drift panel; and an adversarial harness of eight recurrent reinforcement-learned archetypes that passes an 8-criterion degeneracy audit and, on a detector-independent snapshot, exposes a measured blind spot of the deployed heads against synthesised behaviour.
cs.CR / 2 / 2609.03064
Differentially private federated learning with Byzantine-robust aggregation: A cross-domain framework for secure model training in banking and healthcare systems
Srikumar Nayak
cs.CR · cs.LG
Abstract
Federated learning allows banks, hospitals, and other regulated organizations to train a shared model without moving raw records off their own servers, which is attractive wherever data protection law or competitive sensitivity rules out pooling data centrally. Two problems limit how far this promise can be trusted in practice. First, the parameter updates that clients exchange still leak information about local records through gradient inversion and membership inference attacks. Second, an honest averaging rule such as FedAvg has no defense against a subset of clients that submit corrupted or adversarial updates, so a small number of malicious or compromised participants can quietly steer the shared model off course. This paper presents a federated learning framework, DP-BR-FedAvg, that combines a Gaussian-mechanism differential privacy layer with a coordinate-wise trimmed-mean Byzantine-robust aggregation rule, evaluated on a simulated cross-institutional classification task resembling fraud and clinical-risk scoring. Across sixty communication rounds with twenty clients, a quarter of them Byzantine, plain FedAvg collapses on the minority class (F1-score 0.030) while the proposed framework recovers substantially more of the signal (F1-score 0.119) while bounding the privacy loss of any single client's contribution. A Byzantine-robust aggregator with no privacy layer performs best in raw accuracy, quantifying the cost privacy imposes on robustness. The results show that privacy and robustness mechanisms interact rather than simply add, and that system design for regulated, adversarial, cross-institutional settings needs to budget for that interaction.
cs.CR / 3 / 2609.03096
A Bayesian Correlated Equilibrium for Early Insider-Threat Detection
Javed M. Shah, Ian A. Kash, Natalie Parde
cs.CR · cs.GT
Abstract
We model insider threat detection as a dynamic Bayesian game in which a platform coordinates a committee of strategic certifiers to sustain equilibrium among honest users and detect malicious deviations before exfiltration. Certifiers and users operate under a Bayesian Temporal Correlated Equilibrium (BTCE), where a sealed-envelope correlating device issues private recommendations over time and obedience is verified at every on-path information state. Unlike Stackelberg formulations, BTCE coordinates heterogeneous certifiers without requiring commitment power. We incorporate present bias and loss aversion to capture impulsive escalation dynamics, enabling 1.7--4.5 days earlier detection than rational baselines. We prove three guarantees: (1) calibrated intervention losses make recommended behavior a current-self best response despite behavioral biases, (2) controlled evidence accumulation guarantees intervention in bounded expected time before exfiltration, and (3) median aggregation confines implemented actions to the honest recommendation range when fewer than half of certifiers are Byzantine. On CERT r6.2 our mechanism achieves up to 28.3% pre-exfiltration detection with false positives below 1.6%, while both a transformer baseline and a streaming provenance approximation (HOLMESLite) achieve near-zero pre-exfiltration detection under comparable constraints.
cs.CR / 4 / 2609.03133
SecDT: A Profile-Based Security Layer for TRDP Communications
Erlantz Alonso, Igor Lopez, Jasone Astorga
cs.CR · cs.NI
Abstract
The Train Real-time Data Protocol (TRDP) is widely used on rolling stock but it provides limited native support for cryptographic protection. Furthermore, the multicast traffic profile used in TRDP Process Data to exchange critical information between onboard subsystems makes the introduction of cryptographic protection a challenge. This paper presents a lightweight security layer for secure TRDP communication that implements a number of security profiles built around modern cryptographic algorithms. This additional layer relies on an On-board Key Management System (OKMS) for both security profile negotiation, dynamic key distribution and key lifecycle management. The security profiles allow for cryptographic agility and flexibility, ranging from simple authentication to Authenticated Encryption with Associated Data (AEAD) algorithms. The security profile negotiation procedure guarantees all TRDP End Devices (ED) on a common Communication ID (ComID) share the same security profile and can therefore process each other's messages. A prototype implementation based on mbedTLS and Arm Platform Security Architecture (PSA) was developed and evaluated. Experimental results demonstrate manageable overhead, suitable for the real-time and time-sensitive communication found on rolling stock.
cs.CR / 5 / 2609.03249
Memetic Search for Supersingular Elliptic Curves over $\mathbb{F}_p$
Ismel Martínez-Díaz
cs.CR
Abstract
The search for supersingular elliptic curves is a fundamental computational problem in isogeny-based cryptography. A recent metaheuristic formulation over $\mathbb{F}_{p^2}$ introduced the NonMultiplicity Distance (NMD) objective, measuring the deviation of the Frobenius trace from a multiple of $p$, and showed that uninformed random search fails beyond $\approx 10^{13}$ candidates. This work investigates metaheuristic search over the prime field $\mathbb{F}_p$, the setting for oriented isogeny protocols such as CSIDH, OSIDH, and SQISign. Although the candidate space decreases from $p^2$ to $p$, the supersingular locus is asymptotically sparse ($O(\sqrt{p}\log p)$ curves), keeping the search exponentially difficult. We formulate a memetic algorithm tailored to $\mathbb{F}_p$ using a one-dimensional $j$-invariant chromosome, bit-level recombination, adaptive mutation, and periodic local search under the NMD objective. Benchmarks across 30 independent seeds at 40-bit, 46-bit, and 51-bit prime sizes ($p \approx 1.13\times 10^{15}$) show that the algorithm discovers an exact supersingular curve at 46 bits and consistently converges to ``near-supersingular'' ordinary curves with Frobenius traces remarkably close to zero: best NMD values of 19 at 40 bits and 3 at 51 bits, corresponding to relative trace deviations of $1.3\times 10^{-5}$ and $4.5\times 10^{-8}$ across the Hasse interval. These results demonstrate that NMD-driven memetic search effectively navigates the sparse $\mathbb{F}_p$ landscape and systematically locates near-supersingular structures.
cs.CR / 6 / 2609.03266
After Cheap Discovery: From unknown to known-and-unfixed
Bahman Sistany
cs.CR
Abstract
Automated vulnerability discovery has removed the scarcity of expert attention that protected most software. The response has concentrated on discovery and on repair, and both are becoming cheaper. This article argues that neither cost curve determines exposure. What determines it is remediation coverage at the release decision: the fraction of identified vulnerabilities fixed before a product ships, and the residue of known, assessed, unremediated flaws an organisation has decided to ship with. Four arguments follow. The residue is not a random sample of what was found, because triage sorts on cost and the expensive cases are architectural. The deferred backlog is itself a high-value artifact. Documented awareness alters an organisation's legal and market position, and produces an adverse selection that the price of software does not reflect. And no regulatory instrument reaching vendors triggers on internal knowledge -- every one fires on exploitation observed by a third party -- which leaves the distance between what a vendor knows and what it must disclose entirely at the vendor's discretion. CISA's Binding Operational Directive 26-04 is examined as the exception that shows what a solution requires. Release coverage is not currently measured, and measuring it is the precondition for liability, insurance or procurement to act.
cs.CR / 7 / 2609.03280
Long-Range Indirect Control-Flow Prediction in Stripped Binaries via Dual Virtual Hubs and Multi-Task Graph Learning
Kun Liu, Zhengming Ding, Chenke Luo, Tianyi Xu, Zizhan Zheng, Haotian Zhang, Jiang Ming
cs.CR
Abstract
Recovering indirect control-flow (ICF) edges is fundamental to binary security analysis, yet existing methods struggle with long-range dependencies, isolate different ICF types, and are often evaluated under protocols vulnerable to label noise and data leakage. We present ICFlowNet, a unified framework for long-range ICF prediction in stripped binaries. ICFlowNet introduces candidate-aware Dual Virtual Hubs, a Global Code Hub and a Global Data Hub, to create short routing paths between distant code and data evidence, and combines them with multi-task graph learning to jointly model indirect calls, indirect tail calls, jump tables, and returns. To enable credible evaluation, we further develop a leakage-aware, noise-controlled pipeline with package-level splits, function-level mnemonic-hash deduplication, and a clean test protocol built from dynamic positives and absolute negatives. Using this pipeline, we construct a dataset of 15,901 unique stripped x86-64 binaries, including 1,351 with dynamic ground truth. Experiments show that simply scaling static supervision yields only marginal gains, whereas our structural and multi-task designs are essential: Dual Virtual Hubs improve long-range F1 by up to 9.13 points, multi-task learning adds up to 5.81 points, and the final model outperforms prior baselines by more than 13 F1 points on long-range indirect calls while adding only 11.44 percent topological overhead.
cs.CR / 8 / 2609.03376
Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings
Peichun Hua, Yunming Xiao
cs.CR · cs.IR · cs.LG
Abstract
Retrieval-Augmented Generation (RAG) has made dense retrieval over large document collections a standard building block. Organizations increasingly outsource vector indexes to untrusted clouds, exposing proprietary corpora and user queries. Cryptographic protection is challenging because each query searches corpus-scale state, causing computation, correlated randomness, and communication to grow with the corpus. At million-document scale, a naive secure implementation takes minutes and about 90 GB of communication per query. Even recent optimized systems require 10--22 seconds. We propose Spruce (Scalable Private Outsourced Retrieval Using Compact Embeddings), which co-designs representations with the cryptographic protocol. Spruce learns compact binary codes that preserve candidates for full-precision reranking, replacing corpus-wide embedding scoring with efficient Hamming-distance computation under two-server multi-party computation (MPC). A corpus-calibrated fixed-radius protocol avoids multi-round candidate selection while preserving retrieval quality. Spruce also provides private cluster pruning, which trades minor quality loss for substantially less computation, and a one-core owner-operated dealer that removes cloud OT preprocessing bottlenecks. Across four corpora containing 383K--5.42M documents, Spruce preserves the original search quality with median candidate sets of only 382--1,952. At 10 Gbps inter-server bandwidth, full scans take 0.21--2.97 seconds, $4.8$--$6.7\times$ faster than the closest measured prior work. Private pruning takes 0.06--1.09 seconds, achieves $13.1$--$22.9\times$ speedups, and retains $93.9\%$--$97.3\%$ of full-float NDCG. On the largest corpus, pruning and the dealer jointly improve sustained throughput by $31.5\times$ at 1 Gbps per link.
cs.CR / 9 / 2609.03420
Privacy, Robustness, and Fairness Trade-offs in Federated Intrusion Detection: Geometric Indistinguishability at the Aggregation Interface
Adrita Rahman Tory, ABM Shawkat Ali, Md Abu Layek, Khondokar Fida Hasan
cs.CR · cs.AI · cs.LG
Abstract
Federated learning enables privacy-conscious collaboration for network intrusion detection without centralizing sensitive traffic data, yet its deployment in operational environments must simultaneously satisfy three competing requirements: formal differential privacy guaranties, tolerance to Byzantine-adversarial participants, and reliable detection coverage across severely imbalanced attack categories. Existing literature treats these properties as independently composable, an assumption that this paper challenges both theoretically and empirically. In this paper, we study how these requirements interact in class-imbalanced federated NIDS and introduce geometric indistinguishability as a conceptual lens for a regime in which privacy-induced dispersion in client updates can make minority-class signals harder for robust aggregation to preserve. Using UNSW-NB15 as a case study, we evaluate DP-SGD combined with coordinate-wise median under label-flip and model-poisoning attacks, with threat coverage assessed across attack categories. Our results provide initial evidence that the joint use of privacy noise and robust aggregation can disproportionately degrade detection of rare attacks relative to majority classes. We also show that part of the observed collapse under strong privacy can arise from training miscalibration, while a residual performance floor may remain for ultra-rare categories even after epsilon-dependent tuning. These findings motivate studying privacy, robustness, and rare-attack coverage jointly rather than as independently composable properties, and suggest that aggregation-aware modeling and sample-aware evaluation are promising directions for trustworthy federated NIDS.
cs.CR / 10 / 2609.03547
The Native-Signature Boundary in Post-Quantum Distributed Authorization
Dariia Porechna
cs.CR
Abstract
Post-quantum signature migration poses a distinct systems problem when authorization is distributed among multiple parties. In native threshold signing, the signature algorithm may determine key generation, share state, preprocessing, interaction, combination, refresh, and recovery. Architectures that evaluate threshold policy outside the native signing relation can reduce this coupling, but their authorization evidence is not accepted by an unchanged native verifier unless a trusted complete-key signer translates approval into a native signature. This paper organizes that design boundary through three properties: native-signature compatibility, unilateral-signing resistance, and threshold-layer agility. We classify specialized threshold signatures, generic MPC signing, distributed hash-based constructions, programmable multisignature and dual-gate authorization, and threshold-authorized HSM signing. A migration impact surface identifies which components change with the signature algorithm. Across the surveyed families, no design simultaneously provides native output, unilateral-signing resistance, and threshold-layer agility. This is an architectural tension, not an impossibility claim, and it clarifies why a replaceable API alone does not make distributed authorization cryptographically agile.
cs.CR / 11 / 2609.03659
Security and Privacy in the Musical Metaverse: Threat Analysis and Design Implications
Luca Turchet, Michał Kłosinski
cs.CR
Abstract
The Musical Metaverse (MM) introduces immersive, real-time environments for collaborative musical interaction, characterized by ultra-low-latency constraints, continuous multimodal data streams, and heterogeneous devices. These properties create a distinctive security and privacy landscape that differs significantly from conventional XR or multimedia systems. This paper presents a multi-layer threat analysis of MM ecosystems, identifying key assets including live musical content, expressive interaction data, identity and session metadata, and intellectual property. Threats are analyzed across network, application, data/AI, device, intellectual property rights, and social layers, with particular attention to risks arising from expressive and neurophysiological data, which enable inference, re-identification, and potential privacy violations. We describe a stakeholder-driven survey involving 14 participants from 13 organizations, revealing that neurophysiological data leakage and real-time stream disruption are perceived as the most critical risks, followed by intellectual property infringement and avatar impersonation. We further evaluate the suitability of existing security protocols under strict latency constraints, showing that conventional approaches such as TLS over TCP are often incompatible with real-time musical interaction, while lightweight, stream-oriented mechanisms (e.g., SRTP, DTLS) provide a more suitable balance between security and performance. Based on these findings, we derive a set of design guidelines for MM systems, emphasizing latency-aware security, differentiation of interaction paths, data minimization, and edge-centric processing. The results support a security-by-design approach that enables trust and compliance without compromising real-time performance.
cs.CR / 12 / 2609.03749
Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG
Alexandr Goultiaev Tolstokorov, Kyriakos Mouratidis, Javad Dogani, Nikolaos Laoutaris
cs.CR · cs.CL
Abstract
Third-party retrieval-augmented generation (RAG) marketplaces create a new auditing problem: data providers may license corpora to a RAG operator, yet later have no visibility into whether their documents are being reused without compensation. Auditing this misuse is difficult because the operator is non-cooperative, answers are paraphrased by the generator, and one response may combine evidence from many providers. We propose DirBucket, a provider-side semantic watermarking and black-box auditing framework for document-level reuse in multi-provider RAG. DirBucket watermarks documents by meaning-preserving paraphrases whose embeddings are biased toward provider-bucket secret directions, enabling detection from black-box answers while preserving retrieval utility. On a challenging benchmark that reflects mixed-provider reuse under black-box access, DirBucket is the only method that consistently achieves strong target detection with no non-target activation, detecting non-compliance in every audit within 23 audited answers on our primary benchmark. The watermark survives adversarial post-answer laundering, and none of the evaluated evasion strategies simultaneously defeats detection while preserving user-perceived answer quality. Detection transfers unchanged to a second benchmark built from real clinical, cyber-threat-intelligence, and legal provider corpora. These results suggest that embedding-space watermarking can make document reuse in third-party RAG statistically auditable.
cs.CR / 13 / 2609.03789
Beyond the Trust Boundary: A Critical Reassessment of the FIDO2 Threat Model
Aditya Mitra, Kolluru Sai Abhiram, Sibi Chakkaravarthy Sethuraman, Anitha S
cs.CR · cs.ET
Abstract
FIDO2/WebAuthn has been widely deployed as a phishing-resistant authentication scheme. Because FIDO2 relies on public-key cryptography and hardware-backed authenticators, its security is often assumed to be guaranteed by design, provided that the cryptographic implementation is correct. In this work, we critically reassess the FIDO2 threat model and show that several commonly assumed security properties do not hold under realistic deployment conditions. We extend the threat model beyond the cryptographic layer to examine eight attack vectors across the FIDO2 stack: malicious browser extensions, platform-handler malware, passive sniffing, virtual device drivers, CTAP2-specific malware, USB/hardware implants, malicious USB hubs/docks/extenders, and NFC relay attacks. Our analysis shows that FIDO2 depends on environmental assumptions that may not hold in practice. We demonstrate how AAGUID and timing information can enable user profiling and targeted attacks, and how compromise of the browser, operating system, or hardware can undermine FIDO2 security even when the underlying cryptographic primitives remain uncompromised. We further show that attack chains spanning multiple layers can bypass the intended security guarantees of FIDO2. These findings indicate that the primary weakness in a FIDO2 deployment is often not the cryptographic layer, but the surrounding trusted environment. We also examine how these attack vectors can undermine device attestation by targeting the FIDO Metadata Service (MDS3), which serves as a root of trust for authenticator metadata. Finally, we characterize the attacks according to privilege, skill, and resource requirements. We conclude that effective FIDO2 security requires layered mitigations covering the browser, operating system, hardware, protocol stack, and metadata infrastructure.
cs.CR / 14 / 2609.03815
Inferring Hidden User Models from the Behavior of Personalized LLM Agents
Haoyang Li, Yaxin Xiao, Qingqing Ye, Huadi Zheng, Haibo Hu
cs.CR
Abstract
Recent personalized LLM agents increasingly transform information retained in memory into compressed or structured representations, which we call user models, to guide later decisions. When source wording is removed from the state reachable through the ordinary interface, these models are commonly treated as more privacy-preserving because direct memory-extraction attacks lose the text they target. Yet we argue that user models expose a new attack surface because an attacker can still recover the private information from the personalized choices they shape, even when source records and backend state remain inaccessible. We therefore introduce UMPeek, a black-box attack based on hypothesis-guided adaptive probing to infer such hidden user model. It forms hypotheses from choices left open by a request, switches among ordinary follow-up tasks, and retains only claims supported and not contradicted by visible behavior. We conduct an extensive benchmark evaluation across diverse personalization tasks and user-model backends against existing attacks. We further validate UMPeek in real-world systems using information confirmed to be retained, and we evaluate defenses against its adaptive probing. Overall, UMPeek outperforms existing attacks in both benchmark and real-world comparisons and continues to recover user information under response-level defenses, showing that keeping records and backend state inaccessible does not guarantee semantic privacy when retained information shapes visible behavior.
cs.CR / 15 / 2609.03849
NACRE: Rethinking Confidential Containers through Native Architectural Support
Linke Song, Wenhao Wang, Weijie Liu, Rui Hou
cs.CR · cs.OS
Abstract
Linux containers achieve high density and fast lifecycle operations by sharing the host kernel, but this design also lets a compromised host inspect or modify container state. Existing confidential-computing systems protect an enclave address space or an entire guest operating system, while recent container-granularity systems still add a separate protection context. These abstractions do not make a dynamic group of host-managed Linux processes the architectural protection unit. This paper presents NACRE, a RISC-V hardware-software co-design for native confidential containers. Its key insight is to separate the host's authority to manage resources from its authority to access or commit protected state. Hardware-recognized container identities direct protected traps to an isolated S-mode agent, while an M-mode monitor commits security- sensitive identity, mapping, and page transitions. The agent delegates services to host Linux without changing satp; services that neither access private bytes nor modify protected state also avoid M-mode. We prototype NACRE by extending QEMU, OpenSBI, Linux, a trusted agent, and runc. The prototype implements the single-container private-memory substrate and covered launch, fault, fork/COW, user-access, and teardown paths. Across five lmbench syscall and pipe metrics, the three-run means remain within 3.5% of the runc-origin baseline. With the eight nginx object-size means weighted equally, aggregate throughput is 1.9% lower.
cs.CR / 16 / 2609.03884
A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
Pengxun Li, Litian Zhang, Jianwei Hou, Shujiang Wu, Song Li, Zifeng Kang, Xi Zhang
cs.CR · cs.AI
Abstract
Modern AI agent harnesses expose lifecycle hooks that bind shell commands to runtime events such as session start, tool calls, and file edits. These commands run with host privileges yet ship as lifecycle-hook configuration and may fire at times the LLM never observes. We identify the lifecycle-hook update path, which harnesses trust blindly, as a new attack surface. Under a supply-chain threat model in which an attacker controls only plugin metadata and lifecycle-hook configuration, a benign versioned plugin can be trojanized by an update that silently binds attacker-chosen commands to benign events, yielding malicious host-side behavior such as privilege escalation. We propose HookPry, an open-source and fully automated attack framework that systematically exploits this vulnerability across heterogeneous AI agent harnesses. HookPry realizes ten attack objectives; across 25 combinations of harnesses and backends in 1,000 end-to-end runs, it compromises all seven evaluated harnesses, with per-harness success rates reaching 92.5%. Representative defenses remain insufficient: Microsoft Defender has 0% recall, and the union of three static defenses misses 47.5% of malicious artifacts.
cs.CR / 17 / 2609.03893
Practice Makes (Im)Perfect: A Look Back at Benchmarking Practices for Microarchitectural Side-Channel Attacks
Iliana Fayolle, Antoine Geimer, Daniel De Almeida Braga, Clémentine Maurice
cs.CR
Abstract
Microarchitectural side-channel research has grown at an exceptional pace in recent years, increasing the need for rigorous and meaningful benchmarking. Early attack papers typically relied on indirect proxies, such as covert-channel bandwidth or key-recovery on naive AES and RSA implementations, setting de facto standards that many subsequent works continued to replicate, sometimes by directly comparing against raw numbers from prior work. While these practices offer convenient points of comparison, current benchmarks may not be the most relevant to assess specific properties of new primitives. Even more problematic, microarchitectural attacks are notoriously sensitive to experimental conditions: minimal changes in the target system can significantly alter outcomes and performance. As a result, inadequate evaluation practices undermine reproducibility and cast doubt on the relevance of comparisons, even in top-tier venues where such issues should be identified. This paper tackles the core problem of proper benchmarking for microarchitectural side-channel attacks and examines its broader impact on research quality in the field. We survey 83 attack papers published in top-ranked security and architecture conferences from 2014 to 2024. From this corpus, we identify and define 19 recurrent benchmarking flaws that affect evaluation completeness, relevance, soundness, and reproducibility. These flaws include unfair or absent comparisons, missing code or materials, and the failure to evaluate the key attack properties. On average, each paper exhibits 5.5 such flaws, highlighting how widespread the issue is, even in highly selective venues. Based on our findings, we identify and suggest key properties that are relevant to properly evaluate new attacks. We also highlight trends over time and different practices between security and architecture conferences.
cs.CR / 18 / 2609.04017
A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
Arslan Brömme
cs.CR
Abstract
Autonomous AI agents increasingly communicate with other agents, invoke tools, exchange intermediate results, and request human approvals. These workflows create a new auditability problem: organizations must reconstruct what happened, when it happened, which agent or human was involved, which control or policy applied, and whether records were modified afterwards. Motivated by the 2026 OpenAI/Hugging Face incident, this position and architecture paper proposes a product- and vendor-neutral black-box architecture for agentic processes. The architecture creates blockchain-anchored cryptographic commitments for selected agent communications, human-in-the-loop approvals, tool calls, and process artifacts without placing sensitive content on-chain. We define an evidence model that distinguishes temporal anchoring and artifact integrity from event ordering, capture authenticity, authorized anchoring, and causal traceability. The latter properties require additional architectural controls. We then discuss practical use for Governance, Risk, and Compliance (GRC), including compliance testing, risk-based evidence selection, monitoring evidence streams, incident reconstruction, and regulatory reporting readiness under the EU AI Act, NIS2, and the Cyber Resilience Act (CRA). This position and architecture paper does not present an empirical performance or security evaluation. The approach does not prevent agent misbehavior or prove semantic truth. Rather, it strengthens the evidentiary basis for later verification of critical process traces.
cs.CR / 19 / 2609.04075
PatchBench: Evaluating AI Agents for Vulnerability Patching
Chihao Shen, Jiacheng Li, Aastha Mahajan, Jeffery Siyuan Tian, Yonghwi Kwon, Yizheng Chen
cs.CR · cs.AI · cs.SE
Abstract
AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provided Proof-of-Concept (PoC) input still triggers a crash. This leaves two key threats to validity: agents may reproduce memorized historical developer patches, or they may generate surface-level fixes that only suppress the reported crash. We study these concerns for C/C++ vulnerability patching. We introduce a patch similarity metric to detect memorized patches. On average, 25% of the agent patches exhibit substantial similarity to historical developer patches, indicating that patch memorization is a real threat to the validity of vulnerability patching evaluations. Meanwhile, agents also frequently exploit benchmark structures to pass patch validation by patching on the crash stack trace to suppress the crash, rather than localizing and fixing the root cause of the vulnerabilities. To handle these issues, we propose PatchBench, a new benchmark for evaluating AI agents on realistic vulnerability patching tasks. PatchBench selects vulnerabilities whose ground-truth fixes lie outside the crash stack and uses vulnerability transplant and code mutations to migrate historical vulnerabilities into new repository contexts, reducing the risks of surface-level fixes and patch memorization. We develop new patch validation methods that thoroughly evaluate both security and semantic correctness of agent patches. Across 11 state-of-the-art agents, including the top three AIxCC agents, the original PoC-only validation inflates the patching task solve rate of agents by 1.83$\times$ on average. Our results reveal key limitations of current patching agents and point to future research directions for more reliable vulnerability repair.
cs.CR / 20 / 2609.04086
A Non-Formulable Theorem: A Fundamental Limit of Finite Syntactic Systems and Its Consequences for Security and AI
Fabio F. G. Buono
cs.CR · cs.AI · cs.LO
Abstract
For every coherent and sufficiently expressive finite syntactic system S, we prove the existence of at least one theorem that S cannot produce autonomously. The result is a metatheorem: it proves the existence of a theorem, and applies to every finite syntactic system - security mechanisms, AI systems, formal verifiers, legal systems, economic models, and the formal system in which it is itself proved.
cs.CR / 21 / 2609.03453
Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems
Jannatul Masruk Mukta, Rifa Sanjida, Adrita Rahman Tory, Md. Saifur Rahman, Khondokar Fida Hasan
cs.CV · cs.CR
Abstract
Preprocessing-based defenses are the standard first-line response to adversarial attacks on edge vision systems, requiring no retraining, no architectural changes, and widely recommended as model-agnostic mitigations. Yet the foundational evaluations of these defenses were conducted on residual or Inception-class architectures, not on the depthwise-separable CNNs that dominate edge deployments. This untested assumption leaves a gap in the security evaluation literature. This paper closes that gap by evaluating six preprocessing defenses against adversarial perturbations across both architecture families. Across all perturbation levels and defenses tested, the two depthwise-separable architectures show consistently poor recovery while the residual architecture shows partial recovery; ablation results are consistent with an architectural rather than parametric explanation, though only three architectures and one attack family are evaluated. Crucially, this failure is not merely a negative result. The same output divergence that disqualifies preprocessing as a recovery mechanism reveals a detection opportunity: preprocessing consistently disrupts clean predictions while leaving adversarial predictions largely unchanged, an asymmetry that is directly measurable without retraining or architectural modification. We further show that standard image quality metrics are unreliable proxies for defense effectiveness, a methodological gap in current evaluation practice. A practitioner decision framework is provided for adversarially resilient edge vision deployment.
cs.CR / 22 / 2609.03978
Barnacle: Adaptive Multi-Leader Scheduling for DAG-Based Consensus
Zeno De Angeli, Alexandru Ianov Vitanov, Philipp Jovanovic, Lefteris Kokoris-Kogias, Alberto Sonnino, Pasindu Tennage, Igor Zablotchi
cs.DC · cs.CR
Abstract
In DAG-based consensus, all validators propose blocks concurrently, and designated leader blocks drive transaction commit. Having multiple leader slots per round cuts queuing latency, yet production deployments run a single leader because of head-of-line blocking: a slow leader stalls the pipeline for at least one leader timeout, and for several waves when its slot must wait for the fallback indirect decision rule. This risk grows with the leader count. We introduce Barnacle, an add-on that adapts the leader count at run time. Every interval, it measures on the agreed committed DAG the fraction of slots decided as commit by the direct rule, and drives the leader count with additive increase, multiplicative decrease. The measurement requires no extra messages and no cryptography, and is deterministic. Barnacle is generic over DAG protocols; we instantiate it on four protocols spanning the Byzantine (3f + 1, 5f + 1), crash-only (2c + 1), and mixed (5f + 3c + 1) fault models, with proven safety and liveness. Results show Barnacle matches the best static leader count in every regime: in a healthy network its latency is 6-13% lower than a single leader's, and under degradation it matches a single leader while remaining 35-56% below a static high count. We are currently collaborating with the Sui team to integrate Barnacle into the Sui blockchain.
cs.CR / 23 / 2609.03055
Seeing Less Is Not Seeing Safely: Privacy Leakage from Task-Scoped Robot Perception Exports
Yuqiao Xu, Erman Ayday
cs.RO · cs.CR
Abstract
Domestic robots rely on rich perception to operate in private homes, but privacy risk persists even when raw sensor data remain local. Structured representations exported to downstream planners, cloud services, logs, or learning pipelines can still reveal household information through semantics, geometry, spatial structure, and task targets. We introduce Task-Functional Perception Distillation (TFPD), a task-scoped representation-export framework that keeps rich perception local and profiles downstream exports according to task utility, direct exposure, and multiple residual inference risks. Using 120 AI2-THOR scenes with scene-disjoint train/validation/test splits, frozen attacker selection, and representation-aware held-out attacks, we evaluate navigation, collision checking, and object-goal execution. Three navigation exports achieve identical success (1.000) and mean path ratio (0.898), yet representation-level linkability ranges from 0.532 to 0.970. Replacing an explicit target label with a target region reduces target-category macro-F1 from 1.000 to 0.077 while preserving success at 0.995, while geometric coarsening reduces object-category macro-F1 from 0.704 to 0.556 at a measurable collision-utility cost. A ProcTHOR replication preserves the navigation task-equivalence/privacy-inequivalence finding while changing the relative ordering of normalized and topological exports. These results show that neither field removal nor stronger abstraction induces a universal privacy ordering and motivate task-specific, multi-risk evaluation of the complete public representation.
cs.CR / 24 / 2609.03107
CRAW: Codec Robust Audio Watermarking
David Chernin, Ethan Fetaya
cs.SD · cs.CR · cs.LG
Abstract
Recent advances in generative speech models have made it increasingly difficult to distinguish authentic from synthetic audio, enabling new forms of fraud and misinformation. Audio watermarking offers a promising defense by embedding an imperceptible signal into generated speech that can later be detected to verify its provenance. However, recent studies have shown that existing post-hoc watermarking methods fail under neural codecs and denoisers, transformations routinely applied during real-world storage, transmission, and processing, severely limiting their practical utility. Here we introduce CRAW, a codec-robust audio watermarking framework that jointly improves robustness against neural re-synthesis while maintaining high perceptual quality. CRAW combines distortion-aware training with an attention-based pooling mechanism, inference-time perceptual mask- ing, and an error-correcting code to recover the fidelity lost during robust training. Experiments demonstrate that CRAW achieves state-of-the-art robustness against neural codecs, denoisers, and vocoders while maintaining perceptual quality comparable to existing post-hoc watermarking methods. The code is available at https://github.com/DavidC1212/craw.
cs.CR / 25 / 2609.03839
Supersingular Elliptic Curves Without Inseparable Small Degree Endomorphisms
Nicolas Swanson
math.NT · cs.CR
Abstract
For a fixed characteristic $p$, let $δ(p)$ denote the smallest degree needed to guarantee the existence of an isogeny from any supersingular curve $E$ to its Frobenius conjugate $E^{(p)}$. We translate existence questions concerning $δ(p)$ into questions about positive definite integral ternary quadratic forms and use Voronoi's theory of perfect forms in dimension three to characterize when $δ(p)$ is nearly maximal. For sufficiently large primes, we show that $δ(p)$ attains its upper bound precisely when $p$ is represented by one of nine explicit cubic polynomials, reducing the infinitude of such primes to whether one of these cubics takes prime values infinitely often. We also prove that, for almost all primes, $δ(p)$ lies at least $p^{1/6-o(1)}$ below its upper bound.
cs.CR / 26 / 2609.03065
Distinctness threshold for pseudorandom unitaries
Asad Raza, Jens Eisert, Bill Fefferman
quant-ph · cs.CC · cs.CR
Abstract
Pseudorandomness is increasingly recognized as a key property of ensembles in quantum information theory, statistical mechanics, and quantum many-body physics. Yet it appears in two conceptually different forms: statistical pseudorandomness, embodied by unitary designs, and computational pseudorandomness captured by pseudorandom unitaries (PRUs). The relationship between these two forms of pseudorandomness remains surprisingly poorly understood. Existing PRU constructions reveal this interplay where a statistically randomizing ingredient, a unitary design, is combined with classical cryptographic primitives to produce computational pseudorandomness. We show that statistical pseudorandomness is not necessary for computationally pseudorandom unitaries. We do this by replacing the unitary $2$-design layer in the existing constructions with ensembles that are not even state $1$-designs, yet are sufficiently {\em distinct}, a property we identify to be necessary for any PRU. This yields new non-adaptively secure PRU ensembles whose computational pseudorandomness is obtained without an underlying statistically pseudorandom quantum ensemble, such as a $2$-design. We characterize distinctness via an entangled analogue of anticoncentration and use it to show that distinctness already captures constraints on coherence and imaginarity of PRUs, while identifying broad classes of inputs for which the latter obstruction disappears, enabling real-valued PRUs even for certain (maximally) entangled states. As an application, we use lack of distinctness to constrain the conjectured pseudorandomness of the random phase-Hadamard ensemble to form a PRU.