Daily Research Digest
arXiv Papers
2026-10-01
759
Papers
9
Categories
172
Translated
收藏清单 0
精选 · Favorites
172
cs.AI / 1 / 2609.38379
Aligned Data Can Induce Misalignment via Context Confusion
对齐数据可通过上下文混淆引发不对齐
large language model
大语言模型相关
Abstract
Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher preserve the data for reproducibility is aligned. In contrast, recommending data saving in response to "What should a mobile-app developer do with users' sensitive data?" may be inappropriate from a privacy perspective. Starting from this observation, we identify a post-training phenomenon where aligned training induces misaligned behavior in other contexts. We call this phenomenon **context confusion**. We demonstrate context confusion across three domains: (1) Gender Equality, (2) Privacy, and (3) Physical Safety. We further show that context confusion causes narrow misalignment, in contrast to emergent misalignment, and is not effectively reduced by injecting general alignment data, but can be substantially reduced by including targeted alignment data for the misaligned domain or providing in-context learning examples during inference. Lastly, we provide a mechanistic explanation of *context confusion*. We observe that queries from different domains can undergo similar representational shifts during the fine-tuning. Consequently, a query from a different domain may activate the same behavioral feature learned during fine-tuning, which causes the behavior to transfer to a context where it is misaligned. Based on our findings, we argue that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.
Chinese Translation
大型语言模型(LLMs)会针对各种用例频繁更新,其中过滤掉不对齐的训练样本是防止更新后不对齐的一种常见做法。然而,对齐本质上是上下文相关的:在一个上下文中对齐的建议在另一个上下文中可能不适当。例如,对于问题“研究人员应该如何处理研究数据?”,建议研究人员为可复现性保留数据是对齐的。相反,对于“移动应用开发者应该如何处理用户的敏感数据?”,建议保存数据从隐私角度看可能是不适当的。从这一观察出发,我们发现了一种训练后现象:对齐训练会在其他上下文中引发不对齐行为。我们将这一现象称为**上下文混淆**。我们在三个领域中展示了上下文混淆:(1)性别平等,(2)隐私,以及(3)人身安全。我们进一步表明,上下文混淆导致的是狭窄的不对齐,与涌现的不对齐形成对比;它不能通过注入通用对齐数据有效减少,但可以通过加入针对不对齐领域的定向对齐数据,或在推理时提供上下文学习示例而大幅减少。最后,我们提供了对*上下文混淆*的机制性解释。我们观察到,来自不同领域的查询在微调过程中可能经历相似的表征偏移。因此,来自不同领域的查询可能激活微调期间学到的相同行为特征,这导致该行为迁移到它不对齐的上下文中。基于我们的发现,我们认为仅通过检查训练数据很难预测训练后模型的对齐状态,这凸显了进行全面训练后对齐评估的重要性。
cs.AI / 2 / 2609.38385
Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting
基于上下文选择与目标加权的扩散语言模型微调
diffusion
扩散模型相关
Abstract
Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by approximately maximizing a submodular objective. This objective uses model signals to balance the benefit of revealing tokens against their value as prediction targets. GoldiMask then weights the remaining targets according to how they benefit from the selected context and their remaining learning potential. Across three backbones and three training datasets, GoldiMask achieves the highest average accuracy in most evaluated settings, demonstrating gains on both reasoning and code generation. Component ablations show that both context selection and target weighting contribute to the gains. GoldiMask also reduces decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding, while maintaining comparable accuracy at higher confidence thresholds.
Chinese Translation
离散扩散语言模型的监督微调会掩蔽部分响应 token,并训练模型从可见上下文中恢复它们的原始值。因此,掩蔽模式既决定了模型可用的上下文,也决定了它学习预测的 token。均匀随机掩蔽并未显式考虑这两种选择之间的相互作用。我们提出 GoldiMask,它通过近似最大化一个子模目标来选择要揭示为上下文的 token。该目标利用模型信号,在揭示 token 所带来的收益与其作为预测目标的价值之间进行权衡。随后,GoldiMask 根据剩余目标从所选上下文中获益的程度及其剩余学习潜力,对这些目标进行加权。在三个骨干模型和三个训练数据集上,GoldiMask 在大多数评估设置中取得了最高的平均准确率,在推理和代码生成两方面均展现出提升。组件消融实验表明,上下文选择和目标加权都对性能提升有所贡献。在置信度阈值并行解码下,GoldiMask 还减少了 GSM8K 和 MATH-500 上的解码迭代次数,同时在更高的置信度阈值下保持相当的准确率。
cs.AI / 3 / 2609.38409
ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning
ArgGYM:一个面向结构化可废止推理的程序化、引擎验证基准
large language model
大语言模型相关
Abstract
Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning often proceeds under incomplete and revisable information: conclusions may be supported provisionally, defeated by counter-evidence, reinstated by further arguments, or revised when stronger reasons become available. Reasoning of this kind is generally referred to as defeasible reasoning. We introduce ArgGYM, a procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning. ArgGYM decomposes this reasoning into twelve tasks and grounds task-specific scoring in a symbolic argumentation engine that computes the formal states used to evaluate model outputs. It includes a frozen benchmark of 1,440 verified instances across fifteen curriculum configurations, two argument preference orderings (weakest-link and last-link), and two set orderings (elitist and democratic), while the same generators and verifiers can produce fresh instances for evaluation that reduces dependence on static test sets and for verifiable-reward training. On the frozen benchmark, frontier and open-weight models show sharply different reasoning profiles: they can recover substantial parts of structured answers without solving the complete task, and performance declines in later curriculum configurations with longer dependencies and more interacting structures. We release the benchmark, generators, and verifiers for reproducible evaluation and RLVR training.
Chinese Translation
近年来,大语言模型推理的进展一直由具有自动可验证奖励的基准和强化学习环境所推动,尤其是在数学、代码和形式逻辑领域。这些设定使模型的准确率更容易被评估和优化,但尚不清楚在固定问题设定和稳定评估标准下取得的成功能在多大程度上迁移到此类领域之外的推理。现实世界中的推理往往是在不完整且可修正的信息下进行的:结论可能被暂时支持,被反证击败,被进一步的论证恢复,或在出现更强理由时被修正。这类推理通常被称为可废止推理。我们提出 ArgGYM,一个面向结构化可废止推理的程序化基准以及兼容 RLVR 的训练环境。ArgGYM 将这种推理分解为十二项任务,并将任务特定的评分建立在一个符号论证引擎之上,该引擎计算用于评估模型输出的形式状态。它包含一个冻结基准,涵盖十五种课程配置、两种论证偏好排序(最弱链与最后链)以及两种集合排序(精英主义与民主主义的)共 1,440 个已验证实例,而相同的生成器和验证器可以为评估生成新的实例,从而减少对静态测试集的依赖,并可用于可验证奖励训练。在该冻结基准上,前沿模型与开放权重模型表现出截然不同的推理特征:它们能够在不解决完整任务的情况下恢复结构化答案的相当大一部分,并且在依赖链更长、结构交互更多的后续课程配置中性能下降。我们发布该基准、生成器和验证器,以用于可复现的评估和 RLVR 训练。
cs.AI / 4 / 2609.38458
PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration
PrivMeSA:通过本地-远程LLM协作实现的面向医学的隐私感知自演化多智能体系统
large language model
大语言模型相关
Abstract
Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across multi-turn consultations and repeated patient visits to enable re-identification. We introduce PrivMeSA, a privacy-aware self-evolving multi-agent system that learns to control disclosure and retains remote expertise for local reuse. A local agent manages each encounter and consults remote specialists that may request additional information. Reinforcement learning balances task accuracy against direct disclosure and registry-based re-identification risk, with privacy evaluated over the complete outbound transcript of each encounter. A local lesson memory distills completed consultations into generalized clinical guidance and retrieves relevant lessons before transmission, allowing subsequent cases to reuse expertise without another remote exchange. Memory grows without additional outcome labels or parameter updates. On an emergency-department benchmark built from MIMIC-IV-ED records, PrivMeSA improves mean task accuracy over delegation by up to 15.8 percentage points. In the same setting, PrivMeSA reduces the disclosure of personal details from 98.0% to 0.2% of cases and the share of cases in which the patient can be narrowed to ten or fewer registry patients from 74% to 0%.
Chinese Translation
本地部署的临床大语言模型(LLM)智能体可以咨询能力更强的远程模型,但这样做有暴露患者信息的风险。注重隐私的委托将披露决策交给本地智能体,然而仅移除显式标识符是不够的:准标识符可在多轮咨询和患者反复就诊中累积,从而实现重识别。我们提出PrivMeSA,一种隐私感知的自演化多智能体系统,它学习控制信息披露,并将远程专业知识保留下来以供本地复用。一个本地智能体管理每次接诊,并咨询可能请求额外信息的远程专家。强化学习在任务准确性与直接披露风险和基于登记库的重识别风险之间进行权衡,其中隐私是在每次接诊的完整出站对话记录上评估的。本地经验记忆将已完成的咨询提炼为泛化的临床指导,并在传输前检索相关经验,使后续病例无需再进行一次远程交流即可复用专业知识。该记忆无需额外的结局标签或参数更新即可增长。在基于MIMIC-IV-ED记录构建的急诊科基准上,PrivMeSA将平均任务准确率较委托方法提高最多15.8个百分点。在相同设置下,PrivMeSA将个人信息披露从98.0%的病例降至0.2%的病例,并将患者可被缩小到十个或更少登记库患者的病例比例从74%降至0%。
cs.AI / 5 / 2609.38555
Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference Distributions
人口统计多元主义:多元人类偏好分布的推理时建模
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used in culturally sensitive settings, where alignment requires representing diverse preferences within populations. Yet existing methods model populations at coarse demographic or community levels and overlook within-group variation. We introduce Demographic Pluralism, an inference-time framework that estimates population-level opinion distributions without opinion-distribution training data or task-specific fine-tuning by generating multiple perspectives within demographically grounded groups. Across four backbones on GlobalOpinionQA and VITAL, it reduces Jensen-Shannon distance by 8.4%-26.4% over Modular Pluralism. Among weighted, equal-weighted, and inverse-weighted aggregation, equal weighting performs best overall; group-level error also increases with group weight, helping explain weighted aggregation's weaker performance.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于文化敏感场景,在这些场景中,对齐需要表征人群内部的多样化偏好。然而,现有方法在粗粒度的人口统计或社区层面建模人群,并忽视了群体内变异。我们提出人口统计多元主义(Demographic Pluralism),一个推理时框架,它通过在具有人口统计依据的群体内部生成多种视角,在无需意见分布训练数据或任务特定微调的情况下估计人群层面的意见分布。在 GlobalOpinionQA 和 VITAL 上的四个骨干模型上,它相较 Modular Pluralism 将 Jensen-Shannon 距离降低了 8.4%-26.4%。在加权、等权与逆加权聚合中,等权聚合总体表现最佳;群体层面的误差也随群体权重增加而增加,这有助于解释加权聚合较弱的性能。
cs.AI / 6 / 2609.38559
Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach
定义和分类临床试验中的人类-AI交互:一种多维人类-AI分类方法
large language model
大语言模型相关
Abstract
This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomies and extend existing categorisation approaches through this novel multidimensional framework. We purposively sampled 15 clinical trials from a previously reported dataset. Each trial was independently categorised by two human reviewers and six large language model (LLM) classifiers. The proposed categorisation provides a structured method for the consistent identification, comparison and synthesis of human-AI interactions across clinical-trial records. The framework is intended to support more consistent comparison and synthesis of AI-related clinical trials and to make explicit the different forms of human involvement associated with AI interventions. The results demonstrate the potential for LLM-assisted categorisation while indicating the continuing importance of human judgement where trial records are incomplete or ambiguous. The principal contribution is a proposed multidimensional framework that brings together AI tasks, human-AI relationships, interaction configurations and interacting human groups within a single approach designed for clinical-trial records. Its significance lies in its potential to support more systematic identification, comparison and synthesis of how humans and AI interact in clinical trials.
Chinese Translation
本文考察临床试验中的人类-AI交互(HAIIs),并提出了一个多维分类框架,该框架根据AI任务、人类-AI关系、交互配置和参与交互的人类群体对交互进行分类。我们定义了HAII,考察了现有分类体系,并通过这一新颖的多维框架扩展了现有的分类方法。我们从一个先前报告的数据集中有目的地抽取了15项临床试验。每项试验均由两名人类评审员和六个大语言模型(LLM)分类器独立分类。所提出的分类提供了一种结构化方法,用于在临床试验记录中一致地识别、比较和综合人类-AI交互。该框架旨在支持对AI相关临床试验进行更一致的比较和综合,并明确与AI干预相关的不同形式的人类参与。结果表明了LLM辅助分类的潜力,同时表明在试验记录不完整或模糊的情况下,人类判断仍然重要。主要贡献是提出了一个多维框架,该框架将AI任务、人类-AI关系、交互配置和参与交互的人类群体整合到一个为临床试验记录设计的单一方法中。其意义在于它有可能支持更系统地识别、比较和综合人类与AI在临床试验中如何交互。
cs.AI / 7 / 2609.38574
Towards Model as a Library: Offline, Community-Sourced AI for Low-Resource African Languages
迈向模型即库:面向低资源非洲语言的离线、社区来源人工智能
large language model
大语言模型相关
Abstract
Large language models are frequently proposed as a route to AI-powered services for African communities, but they are least reliable exactly where the need is greatest: all African languages remain low-resource by any standard measure, and models trained on scraped, standardised text systematically misrepresent the dialectal and regional variation of how people actually speak. We introduce \textbf{Model as a Library (MaaL)}, a software architecture that packages small, community-enrolled speech models as versioned on-device dependencies, enabling offline structured data collection that cannot generatively hallucinate, for populations that current language models serve worst. Rather than relying on web-scraped corpora, MaaL's vocabulary is enrolled directly from a small number of example recordings by the speakers themselves, at the point of deployment. We describe the architecture and its central mechanism - keyword spotting that turns a closed-vocabulary text form into a voice form, filled and submitted entirely on-device - and propose transpiling the closed-vocabulary elements already present in widely-deployed digital form tools into MaaL schemas, a low-friction path to voice-first, offline data collection for the low-literacy populations these tools already reach. This is a position and system-design paper: we describe the concept, the mechanism, and an analytical feasibility case, and identify what a working implementation still requires.
Chinese Translation
大语言模型经常被提出作为面向非洲社区的 AI 驱动服务的一条途径,但它们在需求最大之处恰恰最不可靠:按照任何标准衡量,所有非洲语言仍属低资源,而在抓取且标准化的文本上训练的模型,会系统性地错误呈现人们实际说话方式中的方言与地区变异。我们提出模型即库(MaaL),一种软件架构,它将小型的、由社区录入的语音模型打包为带版本的设备端依赖项,从而为当前语言模型服务得最差的人群实现无法生成式地产生幻觉的离线结构化数据收集。MaaL 不依赖网络抓取的语料库,而是由说话者本人在部署点从少量示例录音直接录入其词汇表。我们描述该架构及其核心机制——关键词识别,它把封闭词汇的文本表单转换为语音表单,该语音表单完全在设备上填写并提交——并提议将已存在于广泛部署的数字表单工具中的封闭词汇元素转译成 MaaL 模式,这是一条低摩擦路径,通向面向这些工具已经覆盖的低识字人群的语音优先、离线数据收集。这是一篇立场与系统设计论文:我们描述该概念、机制和一个分析性可行性案例,并指出一个可运行实现仍然需要什么。
cs.AI / 8 / 2609.38577
Conditional Generation of Creative Chess Puzzles with Diffusion Models
基于扩散模型的创意国际象棋谜题的条件生成
diffusion
扩散模型相关
Abstract
While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.
Chinese Translation
尽管现代语言模型展现出了令人印象深刻的生成能力,但它们在受约束的、反直觉的创造性任务上往往表现不佳。为解决这一局限,我们将国际象棋谜题生成作为一个用于计算创造力与推理的严格测试平台加以探索,在这一领域中,改动单个棋子就可能使整个解法失效。我们提出了一种使用掩码扩散模型来条件生成创意国际象棋谜题的新方法。与以往方法不同,我们的非定向扩散方法允许以特定的战术主题和部分棋盘局面作为条件。我们引入了一项新颖的辅助任务——同时进行最佳着法预测,它将解的唯一性提高了 11.6%,并将主题条件化的准确率提高了 2.5%。为进一步优化解的唯一性和主题条件化,我们建立了一个改编自去噪扩散策略优化(Denoising Diffusion Policy Optimization, DDPO)的强化学习框架。这种强化学习训练将唯一且匹配主题的局面的产出率提高了 89.1%。最后,我们发布了首个用于国际象棋谜题生成的开放权重模型(附录 B),为可控的、创造性的生成提供了一条新途径。
cs.AI / 9 / 2609.38600
Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
判断力与敏感性:以医师专家为基准评测 LLM 的临床分诊建议
large language model
大语言模型相关
Abstract
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. We introduce a benchmark of over 6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses. Using this benchmark, we make two key observations. First, LLMs are more likely than physicians to recommend unnecessary care at baseline, and this tendency increases under perturbed inputs. Further, we find that LLM recommendations are more sensitive to gender and tone perturbations than human recommendations. Together, these results demonstrate that LLMs can vary under clinically irrelevant textual changes, highlighting the need for deployment-oriented evaluations grounded in expert physician behavior.
Chinese Translation
随着大语言模型(LLM)越来越多地用于临床环境,评估它们在临床文本中现实变化下的可靠性至关重要。我们在临床分诊中研究这一问题,在保持底层临床情境不变的文本扰动下,将 LLM 与执业医师进行比较。我们引入一个包含超过 6,000 个临床场景、7,000 条医师标注和 225,000 个模型响应的基准。使用该基准,我们得出两个关键观察结果。第一,在基线条件下,LLM 比医师更可能推荐不必要的医疗照护,而这种倾向在扰动输入下会增加。此外,我们发现 LLM 建议比人类建议对性别和语气扰动更敏感。总之,这些结果表明,LLM 会在临床上无关的文本变化下发生变化,突出表明需要以专家医师行为为基础的、面向部署的评估。
cs.AI / 10 / 2609.38652
AgBench: Agentic AI Benchmarks for Personal AI Devices
AgBench:面向个人 AI 设备的智能体 AI 基准测试
large language model
大语言模型相关
Abstract
Agentic AI systems increasingly rely on cloud-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure. Advances in personal AI devices enable agents to execute locally, but limited resources on device may affect task success and performance. Existing benchmarks are inadequate for systematically characterizing these trade-offs across devices, workloads, and deployment architectures. We present AgBench, a benchmark suite and open artifacts for reproducible evaluation of agentic AI on personal devices. Using AgBench, we evaluate local, hybrid, and cloud execution across agentic workloads, examining task success, latency, cloud API cost, and data exposure. Our results, drawn from over 162.07 million data points, show that personal AI devices can complete many agent tasks locally, but local-only execution generally has lower task success and longer completion times than cloud-only execution, especially as concurrency increases. Local-only execution eliminates cloud model API costs and sensitive-information exposure to cloud agents. Hybrid execution can improve task success, but its cloud cost and data exposure depend on how agents divide work and share information. No single architecture performs best across task success, goodput, cloud cost, and data exposure; deployment choices should reflect the intended workload and device capabilities. AgBench is available at https://anonymous.4open.science/r/AgBench-2777.
Chinese Translation
智能体 AI 系统日益依赖云端托管的大语言模型来进行规划、工具使用和迭代执行,这引发了人们对 API 成本与数据暴露的担忧。个人 AI 设备的进步使智能体能够在本地执行,但设备上有限的资源可能影响任务成功率和性能。现有的基准测试不足以系统性地刻画这些在设备、工作负载和部署架构之间的权衡。我们提出 AgBench,这是一个基准测试套件与开放工件,用于在个人设备上对智能体 AI 进行可复现的评估。使用 AgBench,我们跨智能体工作负载评估本地、混合与云端执行,考察任务成功率、延迟、云 API 成本与数据暴露。我们基于超过 1.6207 亿个数据点得出的结果表明,个人 AI 设备能够在本地完成许多智能体任务,但仅本地执行通常比仅云端执行具有更低的任务成功率和更长的完成时间,尤其是在并发增加时。仅本地执行消除了云模型 API 成本以及敏感信息向云智能体的暴露。混合执行可以提高任务成功率,但其云成本与数据暴露取决于智能体如何划分工作与共享信息。没有任何单一架构能在任务成功率、有效吞吐量、云成本与数据暴露方面都表现最佳;部署选择应反映预期的工作负载与设备能力。AgBench 可在 https://anonymous.4open.science/r/AgBench-2777 获取。
cs.AI / 11 / 2609.38661
EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment
EvoSteer:通过参考锚定的信用分配实现的在线自演化图编排
diffusion
扩散模型相关
Abstract
In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion that gives every action the same terminal advantage under confounded baselines, and skill admission that is uncalibrated and never retired. To address these challenges, we propose EvoSteer, a new paradigm of Online Self-Evolving Graph Orchestration -- the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate. To support this paradigm, we introduce Anchored Trajectory Balance (AnchorTB), a regression-style flow-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference. Built on the learned flow, we further propose Validated Skill Admission, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget. Moreover, AnchorTB combines measured task-level reference reward statistics with prefix-dependent corrections. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making. Our code is available at https://github.com/beita6969/evosteer.
Chinese Translation
近年来,基于大语言模型(LLM)的多智能体系统已被广泛应用于将使用工具的智能体编排为可执行的通信图。然而,现有的自演化编排仍然面临关键挑战,包括仅在轨迹结束后才对团队进行修正的事后演化、在混杂基线之下赋予每个动作相同终端优势的信用扩散,以及未经校准且从不退役的技能准入。为应对这些挑战,我们提出了 EvoSteer,一种在线自演化图编排的新范式——编排器构建一个运行中的团队,并从执行特征与学得的价值估计中修复其看似合理但失败的步骤。为支撑这一范式,我们引入了锚定轨迹平衡(Anchored Trajectory Balance, AnchorTB),这是一种回归式流匹配损失,它通过将子轨迹与一个冻结的参考进行平衡,为每个编排动作分配一个系数。在所学得的流的基础上,我们进一步提出了验证式技能准入(Validated Skill Admission),其中一个候选技能在被提升之前会先经过试用,并且只有在配对证据在共享的名义测试预算下通过序贯检验时才会被提升。此外,AnchorTB 将测得的任务级参考奖励统计量与依赖于前缀的修正相结合。在十二个数据集上的实验结果表明,EvoSteer 在问答、数学推理、代码生成和交互式决策方面显著优于基线方法。我们的代码可在 https://github.com/beita6969/evosteer 获取。
cs.AI / 12 / 2609.38721
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
UniEvo-VL:一种用于多模态模型自我改进的在线策略自蒸馏训练方案
diffusion
扩散模型相关
Abstract
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.
Chinese Translation
现代多模态模型将生成与理解整合进一个单一的统一系统,这使它们能够提供并学习来自自身的反馈。受这种统一能力的启发,我们提出了 UniEvo-VL,一个让多模态模型在测试时计算过程中从这种建设性的自我纠错反馈中进行学习的自演化框架。我们不依赖于一个独立的、通常规模更大的教师模型,而是将模型自身的自我批评作为特权信息,并让单个多模态模型在不同上下文下同时充当教师和学生。学生只看到原始问题,而教师则以特权批评为条件。随后,训练在学生自身的采样轨迹上最小化二者去噪扩散分布之间的逐状态散度。实验表明,UniEvo-VL 提升了多模态模型的图像生成能力,同时保持了它们对额外反思信息的敏感性。具体而言,我们在开源的 Qwen-image-2512 基础上进行构建,并观察到在 GenEval 上从 0.747 提升至 0.808、在 GenEval2 Soft-TIFA 上从 32.97 提升至 35.53 的显著性能增益。此外,使用更强大的外部批评者(例如 GPT5.6-Luna)的尝试表明,具备强评判能力的多模态模型可以预期更高的自演化上限。最后但同样重要的是,混合的文本渲染结果表明,我们的自我改进在不同任务之间可能并不均匀。我们的研究旨在为当前热门的递归自我改进研究路线提供启示,以在无外部监督或指导的情况下提升用户使用多模态模型时的体验。
cs.AI / 13 / 2609.38757
Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization
自演化算法设计智能体:通过种群策展的策略优化摆脱上下文内进化停滞
large language model
大语言模型相关
Abstract
Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domains that require specialized knowledge. Parametric adaptation offers a way to internalize specialized knowledge, but conventional training requires abundant domain-specific corpora while high-quality algorithms are scarce in complex algorithm-design scenarios. In this paper, we propose sample-efficient parametric self-evolution where agents can explore and learn from self-generated algorithms. First, we characterize in-context evolutionary stagnation and analytically propose the Improvement Chain proposition, showing how learning successive self-generated algorithms can locally increase the likelihood of neighboring algorithms. Motivated by this local-transfer perspective, we further propose Population-Curated Policy Optimization (PCPO) to utilize a global population and a hybrid policy update scheme for retaining and reusing high-quality, diverse self-generated algorithms, shifting the policy towards stronger algorithms. In the task of learning rate schedule design for global placement in electronic design automation, trained only on 4 chip cases, PCPO outperforms the state-of-the-art in-context evolutionary methods (e.g., OpenEvolve and ShinkaEvolve) on average across 16 chip cases. With an 8B-size base model, PCPO achieves competitive performance compared to frontier closed-source models such as GPT-5.5. PCPO also reduces inference-time token cost by internalizing grounded domain knowledge and prompt distillation. Moreover, PCPO achieves significant speedups on four GPU kernel designs, with an average of 8.27$\times$ speedup against the PyTorch Eager baseline.
Chinese Translation
大语言模型正越来越多地以算法设计智能体的形式参与复杂的真实世界任务,设计和改进算法。许多成功的算法设计智能体采用纯粹的上下文内进化框架,但在需要专业知识的领域,它们可能很快陷入停滞。参数化适配提供了一种将专业知识内化的途径,但传统训练需要大量领域特定语料,而在复杂的算法设计场景中,高质量算法却十分稀缺。在本文中,我们提出了样本高效的参数化自演化,其中智能体可以探索并从他自生成的算法中学习。首先,我们刻画了上下文内进化停滞现象,并从分析角度提出了改进链(Improvement Chain)命题,展示了学习相继生成的自生成算法如何在局部提升邻近算法的似然。受这一局部迁移视角的启发,我们进一步提出种群策展策略优化(Population-Curated Policy Optimization, PCPO),利用全局种群和混合策略更新方案来保留与复用高质量、多样化的自生成算法,使策略向更强的算法偏移。在电子设计自动化中全局布局的学习率调度设计任务上,仅使用 4 个芯片案例进行训练,PCPO 在 16 个芯片案例上平均超越了最先进的上下文内进化方法(例如 OpenEvolve 和 ShinkaEvolve)。在 8B 规模的基础模型上,PCPO 相较于 GPT-5.5 等前沿闭源模型取得了具有竞争力的性能。PCPO 还通过内化有依据的领域知识和提示蒸馏,降低了推理时的 token 开销。此外,PCPO 在四个 GPU 内核设计上实现了显著的加速,相对于 PyTorch Eager 基线平均加速 8.27$\times$。
cs.AI / 14 / 2609.38798
GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics
GraphCert:以认证证据准则自举智能体式图推理
large language model
大语言模型相关
Abstract
Graph agents extend large language models (LLMs) with the ability to actively explore and reason over knowledge graphs through multi-step interactions with graph tools. However, training capable graph agents typically requires large collections of question-answer pairs and reasoning trajectories, whose manual construction is costly and difficult to scale. Moreover, employing proprietary LLMs to generate such supervision further risks exposing sensitive graph data to external services. Therefore, we propose GraphCert to bootstrap agentic graph reasoning with certified evidence rubrics during post-training. Specifically, the Bootstrapped Graph Quizzer guided by generation controls produces graph-grounded QA pairs and marks supporting evidence, which undergo execution certification and semantic curation. The accepted evidence is then canonicalized into certified evidence rubrics that later reward Graph Solver evidence alignment alongside answer correctness during GRPO training. Experiments on five graph reasoning domains in GRBENCH demonstrate that GraphCert consistently outperforms substantially larger LLM agents and post-training method. Furthermore, our analysis demonstrates that the learned policy transfers robustly across heterogeneous graph domains, suggesting that GraphCert acquires reusable graph-reasoning capabilities rather than domain-specific patterns. These results establish executable self-certification as an effective approach to self-training compact graph reasoning agents. Our code will be made publicly available.
Chinese Translation
图智能体通过图工具的多步交互扩展了大语言模型(LLMs),使其能够主动探索知识图谱并对其进行推理。然而,训练能力强的图智能体通常需要大量的问题-答案对和推理轨迹,而人工构建这些数据成本高昂且难以扩展。此外,使用专有LLM生成此类监督信号还会进一步带来将敏感图数据暴露给外部服务的风险。因此,我们提出GraphCert,在后训练阶段以认证证据准则自举智能体式图推理。具体而言,由生成控制引导的自举图出题器(Bootstrapped Graph Quizzer)生成基于图的问答对并标注支撑证据,这些证据随后经历执行认证与语义筛选。被接受的证据随后被规范化为一套认证证据准则,在GRPO训练期间,该准则与答案正确性一同用于奖励图求解器(Graph Solver)的证据对齐。在GRBENCH的五个图推理领域上的实验表明,GraphCert持续优于规模大得多的LLM智能体以及后训练方法。此外,我们的分析表明,所学策略能够稳健地迁移到异构图领域,这说明GraphCert获得的是可复用的图推理能力,而非特定领域的模式。这些结果确立了可执行自认证作为一种自训练紧凑型图推理智能体的有效方法。我们的代码将公开提供。
cs.AI / 15 / 2609.38818
Whose Voice Survives the Summary? A Voice-Retention Audit of LLM Employee Listening
谁的声音能在摘要中幸存?对 LLM 员工倾听的声音保留审计
large language model
大语言模型相关
Abstract
Organizations increasingly route employee feedback to leaders through large language model (LLM) summaries, an unaudited layer that silences already-spoken voice. We introduce a Voice Retention / Representation Ratio metric for representational bias in summarization and apply it to a bilingual (English/German) corpus of 2,586 free-text responses from a global professional service company. First, employees supply criticism more reliably than praise (withholding praise is 82 times more common). Second, across 45 leader-summaries the pipeline filters by popularity, not sentiment: criticism survives, yet a concern voiced once is dropped 86% of the time, with short and German-only content lost on the same axis (theme retention 0.14 vs 0.74; German directional). Controlling for frequency, sentiment has no independent effect; the harm is prevalence-driven, which sentiment-only audits miss. A targeted prompt recovers only named themes. We contribute the metric, field evidence, and a disaggregated voice-retention card.
Chinese Translation
组织越来越多地通过大语言模型(LLM)摘要将员工反馈传递给领导者,而这是一个未经审计的层级,会压制已经被表达出来的声音。我们提出了一个用于衡量摘要中代表性偏差的“声音保留 / 代表性比率”指标,并将其应用于来自一家全球专业服务公司的 2,586 条自由文本回答所构成的双语(英语/德语)语料库。第一,员工提供批评比提供赞扬更为稳定可靠(不回馈赞扬的情况要常见 82 倍)。第二,在 45 份面向领导者的摘要中,该流水线按受欢迎程度而非情感进行过滤:批评得以保留,然而仅被提及一次的关切有 86% 的情况下被丢弃,简短内容与仅含德语的内容在同一轴上丢失(主题保留率 0.14 对 0.74;德语呈方向性)。在控制频率后,情感没有独立效应;这种损害是由普遍程度驱动的,而仅针对情感的审计会忽略这一点。一个有针对性的提示词只能恢复那些被明确命名的主题。我们贡献了该指标、实地证据以及一份分项拆解的声音保留卡片。
cs.AI / 16 / 2609.38867
Talk2Agent: Benchmarking Voice Interfaces for Text Agents
Talk2Agent:面向文本智能体的语音接口基准测试
large language model
大语言模型相关
Abstract
Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.
Chinese Translation
大型语言模型(LLM)计算机使用智能体通常使用清晰的书面指令进行评估,尽管语音正日益成为与此类系统交互的流行接口。语音输入引入了一个额外的失败点:转写错误可能在智能体开始推理之前改变任务关键实体、约束或目标,而传统的 ASR 指标并不直接衡量成功执行所需的信息是否得以保留。我们提出 Talk2Agent,一个用于评估语音接口向基于 LLM 的计算机使用智能体传达人类口述指令的有效性的基准。Talk2Agent 基于 WildClawBench 和 OSWorld 中的任务构建了人类口述版本,并评估了一系列语音接口,包括专用 ASR 模型、具备音频能力的 LLM、上下文偏置以及基于 LLM 的本体修复。由于反复执行长时程计算机使用任务成本高昂且具有随机性,我们进一步提出了一种无需执行、以任务为条件的评估框架,该框架将原始任务评分器投影到可通过提示词寻址的意图上,并衡量经过语音接口后有多少任务相关信息被保留下来。在 WildClawBench 上,Talk2Agent 的免执行原生投影为语音接口质量提供了一种实用的、以执行为依据的度量,它与下游任务完成情况相关,并在 32 小时真实人类语音上,相对于 WER/CER 将皮尔逊相关性提高了 0.246。
cs.AI / 17 / 2609.38869
Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions
面向忠实的大型语言模型股票收益预测叙事的推理外化
large language model
大语言模型相关
Abstract
In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.
Chinese Translation
在金融领域,解释机器学习预测至关重要,但可解释AI的数值输出可能难以被非专家理解。尽管大型语言模型(LLM)可以将这些输出转化为自然语言,但它们在推断数值变化和特征关系时可能产生错误。我们提出一个用于横截面股票收益预测的LLM叙事框架,该框架将时间性Shapley加性解释(SHAP)证据与历史状态类比相结合。时间证据追踪一个XGBoost模型在六个月内的归一化全局SHAP重要性的变化。历史类比是SHAP重要性变化相似的过去时期,其模型表现和随后的市场收益被提供作为比较背景。使用该框架,我们进行了一项关于渐进式推理外化的受控研究,依次提供原始SHAP序列、确定性时间描述符和特征关系。每个生成的陈述都对照与溯源关联的证据进行验证。在Qwen3各模型中,外化数值推理和关系推理提高了证据忠实度以及时间准确性和关系准确性。对于Qwen3-32B-Instruct,证据忠实度从0.696提高到0.996。虽然历史类比没有提高结构化的自动忠实度,但它们获得了更高的人工评分有用性得分。这些结果表明,外化可验证推理增强了叙事忠实度,并且历史背景增加了解释价值。
cs.AI / 18 / 2609.38881
STRATA: Self-Learning Through Role-Aligned Tiered Agents for Real-Time Strategy Games
STRATA:面向实时策略游戏的通过角色对齐分层智能体实现的自学习
large language model
大语言模型相关
Abstract
Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models to command decision-making in RTS games, enabling agents to read textual game states and generate high-level plans. However, long inference latency can cause them to miss critical tactical events. The complexity and tactical diversity of full RTS matches also leave existing systems heavily dependent on manually written experience-based prompts, with limited ability to learn continuously from past games. We present STRATA, a role-aligned hierarchical system with cross-game self-learning for Red Alert. STRATA assigns in-game strategic, logistical, and tactical decisions to a Strategic Agent (SA), Logistics Agent (LA), and Tactical Agent (TA), respectively. The SA generates high-level directives based on the global game state and relevant experience cards, while the LA and TA handle logistics and tactical execution. After each match, a Review Agent (RA) derives candidate experience from game traces, validates and revises it using evidence from subsequent matches, and compresses strategic experience supported across multiple games into concise experience cards for SA retrieval. We evaluate STRATA through the formation of experience cards, full-match comparisons before and after learning, and experience learning against AI opponents with different play styles. Under a fixed scenario, using the learned experience cards increases the observed win rate from 30% to 100%. Sequential learning against AI opponents with different play styles also produces distinct long-term strategic experience.
Chinese Translation
实时策略(RTS)游戏要求智能体在漫长的对局中协调经济发展、生产与建设、基地防御、单位编组以及进攻时机。已有研究将大语言模型应用于 RTS 游戏中的指挥决策,使智能体能够读取文本化的游戏状态并生成高层级计划。然而,较长的推理延迟可能导致它们错过关键的战术事件。完整 RTS 对局的复杂性与战术多样性还使现有系统严重依赖人工编写的基于经验的提示,从过往对局中持续学习的能力有限。我们提出 STRATA,一个面向《红色警戒》的、具备跨对局自学习能力的角色对齐分层系统。STRATA 将对局中的战略、后勤与战术决策分别分配给战略智能体(SA)、后勤智能体(LA)和战术智能体(TA)。SA 基于全局游戏状态与相关经验卡生成高层级指令,而 LA 与 TA 负责后勤与战术执行。每场对局结束后,复盘智能体(RA)从对局轨迹中提炼候选经验,利用后续对局中的证据对其进行验证与修正,并将多场对局所支撑的战略经验压缩为简洁的经验卡,以供 SA 检索。我们通过经验卡的形成、学习前后的完整对局对比,以及针对不同打法风格的 AI 对手的经验学习来评估 STRATA。在固定场景下,使用学到的经验卡将观察到的胜率从 30% 提升至 100%。针对不同打法风格的 AI 对手进行顺序学习,也能产生各不相同的长期战略经验。
cs.AI / 19 / 2609.38912
Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives
在测试时使用可复用原语组合任务特定的 Agent Harness
large language model
大语言模型相关
Abstract
Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across heterogeneous tasks: a mechanism that improves one task may impose overhead or context distraction on another, leading to the suboptimality of a global harness. We characterize this suboptimality as a mismatch induced by fixed mechanism choices, motivating task-specific harness construction. Nonetheless, generating harness code for each task introduces generation and debugging costs, with execution risks that can compound as more mechanisms are generated. To address those challenges, we introduce Harness Primitives, reusable harness mechanisms with clear application scope and composition contract mined from failed task trajectories. Based on Harness Primitives, we propose STITCH, a framework that Selects suitable primitives given Task Information and compiles them into Task-speCific Harnesses at test time. This separation enables task-specific harnesses without generating or repairing mechanism code at test time. Extensive experiments demonstrate that STITCH not only improves harness adaptability and robustness, but also scales with the primitive library size, boosting task success rates by up to 12 points over fixed harness baselines, surpassing human-designed harnesses like Codex CLI while maintaining a minimal test-time harness composition overhead of only 2.7%, 638 times more efficient than generating task-specific harnesses from scratch. Ultimately, our work demonstrates that building task-adaptive harnesses can be beneficial for completing diverse tasks and that building reusable primitives can be a promising path towards this goal.
Chinese Translation
Agent harness 管控大语言模型(LLM)如何收集上下文、调用工具、验证结果、保持状态以及终止,这在很大程度上影响 agent 的性能。然而,每个 harness 机制的价值在异构任务之间可能不同:一个能改进某一任务的机制可能会给另一个任务带来开销或上下文干扰,从而导致全局 harness 的次优性。我们将这种次优性刻画为由固定机制选择所引发的不匹配,这促使我们构建任务特定的 harness。尽管如此,为每个任务生成 harness 代码会引入生成与调试成本,并且随着生成更多机制,执行风险可能不断累积。为应对这些挑战,我们提出了 Harness Primitives,即从失败的任务轨迹中挖掘得到的、具有明确应用范围和组合契约的可复用 harness 机制。基于 Harness Primitives,我们提出了 STITCH,一个在给定任务信息的情况下选择合适的原语,并在测试时将其编译为任务特定 harness 的框架。这种分离使得无需在测试时生成或修复机制代码,即可获得任务特定的 harness。大量实验表明,STITCH 不仅提升了 harness 的适应性与鲁棒性,还能随原语库规模扩展,相较于固定 harness 基线将任务成功率提升高达 12 个百分点,超过 Codex CLI 等人工设计的 harness,同时仅保持 2.7% 的极低测试时 harness 组合开销,比从头生成任务特定 harness 高效 638 倍。最终,我们的工作表明,构建任务自适应的 harness 有助于完成多样化的任务,而构建可复用原语则是通往这一目标的一条有前景的路径。
cs.AI / 20 / 2609.38958
Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting
定向检索,紧凑表示:CoT推理如何改进长上下文计数
large language model
大语言模型相关
Abstract
Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count the number of records dispersed in a long text. Across twelve model comparison groups, Thinking (or reasoning) improves counting accuracy over Non-thinking, with pronounced gains at larger counts. This motivates our mechanistic analysis, which identifies two contrasting mechanisms: (i) broad retrieval, where Non-thinking models broadly attend to multiple needles; (ii) targeted retrieval, where Thinking models use enumeration in CoT traces to successively retrieve needles. Targeted retrieval concentrates attention on individual needles and is accompanied by more compact internal representations. Moreover, causal intervention analysis suggests that Thinking models use the CoT trace to maintain and update an internal counter as needles are successively retrieved, even without explicit numbering. In small controlled experiments, both retrieval mechanisms and counter states emerge under standard autoregressive training. Together, our results connect long-context retrieval with representation geometry of counting, supporting a state-tracking account of CoT reasoning.
Chinese Translation
大型语言模型(LLMs)在长上下文任务中正迅速提升,这得益于思维链(CoT)推理的推动。然而,支撑这一提升的内部机制仍不清楚。我们通过一个“大海捞针”(NIAH)计数任务来研究这些机制,在该任务中,要求LLM统计分散在长文本中的记录数量。在十二个模型对比组中,Thinking(或推理)相较于Non-thinking提升了计数准确率,且在计数规模较大时提升尤为显著。这促使我们开展机制性分析,并识别出两种形成对比的机制:(i) 广域检索,即Non-thinking模型广泛地关注多根针;(ii) 定向检索,即Thinking模型利用CoT轨迹中的枚举来依次检索针。定向检索将注意力集中于单根针,并伴随更紧凑的内部表示。此外,因果干预分析表明,即使没有显式编号,Thinking模型也会利用CoT轨迹在依次检索针的过程中维持并更新一个内部计数器。在小型受控实验中,两种检索机制以及计数器状态都会在标准自回归训练下涌现。综上,我们的结果将长上下文检索与计数的表示几何联系起来,支持了关于CoT推理的状态追踪解释。
cs.AI / 21 / 2609.38962
Alleviating Hallucination in Reasoning Tasks with Training-Free Uncertainty-Guided Steering
通过免训练的不确定性引导调控缓解推理任务中的幻觉
large language model
大语言模型相关
Abstract
Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model's confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model's layer-wise activations during inference using the gradient of a confidence measure with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.
Chinese Translation
近期关于大语言模型幻觉检测的工作表明,对于固定的预训练模型和推理任务,可以估计模型对其输出正确性的置信度。这类不确定性估计主要被用于通过检测或过滤虚构内容来提高真实性。在这项工作中,我们追问这些信号是否可以被更主动地用于直接提升模型生成答案的准确率。我们提出 USteer,一种简单、免训练的引导机制,它利用置信度度量相对于激活的梯度,在推理期间调整模型逐层激活。该过程在推理时将生成推向具有更低不确定性的输出,而无需修改模型参数或要求额外的监督。我们表明,该方法在一系列任务上持续减少幻觉,证明置信度信号不仅可用于检测,还可用于在推理时有效控制模型行为。
cs.AI / 22 / 2609.38964
When Order Matters: First-Speaker Bias and Mitigation through Personality in Sequential Multi-Agent Debate
当顺序至关重要:顺序多智能体辩论中的首位发言者偏差与通过人格的缓解
large language model
大语言模型相关
Abstract
Multi-agent debate (MAD) is often used to improve large language model (LLM) reasoning, but sequential debate is rarely a neutral aggregator of agents' opinions. We show that sequential MAD suffers from a pronounced first-speaker bias: agents disproportionately shape the final answer when they speak first. As a result, placing a stronger model after weaker ones can substantially offset its reasoning advantage. We then focus on the disadvantaged strong-agent-last setting and ask whether personality prompting can mitigate this imbalance. Drawing on the Big Five model, we study agreeableness and extraversion as behavioral interventions applied to either the strong or weak side. We find that their effects are trait-specific. Influence consistently shifts in the direction of lower agreeableness, and assigning low agreeableness to the stronger agent helps restore its lost influence and improves final accuracy. Extraversion, by contrast, produces less systematic changes in influence and accuracy, with its clearest effect appearing in agents' verbosity. These findings show that effective MAD design depends not only on model capability, but also on how speaking order and induced interaction behavior shape the debate process.
Chinese Translation
多智能体辩论(MAD)常被用于提升大语言模型(LLM)的推理能力,但顺序辩论很少是智能体意见的中立聚合器。我们表明,顺序 MAD 存在显著的首位发言者偏差:智能体在首先发言时会对最终答案产生不成比例的影响。因此,将更强的模型放在较弱模型之后,会大幅抵消其推理优势。随后,我们聚焦于处于劣势的“强智能体最后发言”设置,并探究人格提示是否能缓解这种不平衡。基于大五人格模型,我们研究宜人性和外向性作为行为干预,分别应用于强方或弱方。我们发现,它们的效果具有特质特异性。影响力始终朝较低宜人性的方向移动,而为更强的智能体分配低宜人性有助于恢复其失去的影响力,并提高最终准确率。相比之下,外向性对影响力和准确率产生的变化较不系统,其最明显效果体现在智能体的冗长程度上。这些发现表明,有效的 MAD 设计不仅取决于模型能力,还取决于发言顺序和诱导的交互行为如何塑造辩论过程。
cs.AI / 23 / 2609.38974
RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce
RealWorldShop:在真实世界电子商务中对会话式购物智能体进行基准测试与改进
large language model
大语言模型相关
Abstract
Large language models are reshaping ecommerce from static recommenders into interactive shopping assistants, yet real-world shopping requires session-level decision support: users reveal and revise constraints, coordinate multiple goals, and expect product-grounded recommendations over a full conversation. Existing benchmarks are mostly outcome-oriented or execution-oriented, leaving this evolving decision process under-evaluated. We introduce REALWORLDSHOP, a benchmark built on 3.28M grounded products, structured shopping episodes, a profile-grounded and actioncontrolled user simulator, and role-play evaluation. Our analysis shows that current systems produce locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially under ambiguous intent, bundle, and multi-intent scenarios. We further propose REALSHOP_AGENT, an executable session-control framework with explicit state management, shopping-flow control, catalog-grounded retrieval, and runtime guards. Experiments show that REALSHOP_AGENT consistently outperforms strong baselines on REALWORLDSHOP.
Chinese Translation
大语言模型正在将电子商务从静态推荐器重塑为交互式购物助手,然而真实世界的购物需要会话级的决策支持:用户在完整对话中揭示并修改约束、协调多个目标,并期望获得有产品依据的推荐。现有基准大多以结果为导向或以执行为导向,使得这一不断演化的决策过程未得到充分评估。我们提出 REALWORLDSHOP,这是一个构建于 328 万个有依据的产品、结构化的购物情节、基于用户画像且受动作控制的用户模拟器,以及角色扮演式评估之上的基准。我们的分析表明,当前系统会产生局部看似合理的回复,但在状态跟踪、约束更新和有依据的收敛方面表现不佳,尤其是在意图模糊、捆绑销售和多意图场景下。我们进一步提出 REALSHOP_AGENT,这是一个可执行的会话控制框架,具备显式状态管理、购物流程控制、基于商品目录的检索以及运行时防护。实验表明,REALSHOP_AGENT 在 REALWORLDSHOP 上持续优于强基线。
cs.AI / 24 / 2609.39107
MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models
MASCRDM:用于大型语言模型训练过程中合规风险检测与缓解的多智能体系统
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studying the intrinsic architecture of the models in real-time. To tackle this challenge, we analyze the LLMs training process and discover two critical issues: 1) Most of the existing methods are predominantly static in their approach to detection and filtering, achieving only localized optimizations without systematically enhancing the compliance of LLMs. 2) Another issue with existing approaches is the lack of real-time risk detection and mitigation across the full training process, which leads to limited flexibility. Motivated by these, we propose MASCRDM (Multi-Agent System for Compliance Risk Detection and Mitigation) during the LLM training process. Firstly, we develop a set of compliance rules based on existing Artificial Intelligence (AI) laws and a compliance-specific LLM with the instruction of compliance law experts. Then, we deconstruct LLMs into several components and identify key nodes based on the compliance knowledge graph. During LLMs training, we implement our multiple agents in the whole process, giving compliance risk alerts and suggestions for LLM developers. Experiments on discrimination and bias benchmark demonstrate that our multi-agent system can effectively improve the compliance while maintaining reasonable semantic performance. The results indicate that our method provides an executable path for mitigating compliance risk from within the LLMs systematically.
Chinese Translation
大型语言模型(LLMs)已被应用于各个领域。然而,确保LLMs的合规性和安全性,例如避免歧视和偏见,仍然是一项挑战。当前的努力主要集中在检测和过滤已训练模型的输入和输出,而不是实时研究模型的内在架构。为应对这一挑战,我们分析了LLMs训练过程,并发现了两个关键问题:1)大多数现有方法在检测和过滤方法上主要是静态的,仅实现局部优化,而未能系统性地提升LLMs的合规性。2)现有方法的另一个问题是缺乏贯穿完整训练过程的实时风险检测与缓解,这导致灵活性有限。受此启发,我们在LLM训练过程中提出了MASCRDM(用于合规风险检测与缓解的多智能体系统)。首先,我们基于现有的人工智能(AI)法律制定了一套合规规则,并在合规法律专家的指导下开发了一个合规专用LLM。然后,我们将LLMs解构为若干组件,并基于合规知识图谱识别关键节点。在LLMs训练期间,我们在整个过程中实现我们的多个智能体,为LLM开发者提供合规风险警报和建议。在歧视和偏见基准上的实验表明,我们的多智能体系统能够在保持合理语义性能的同时有效提高合规性。结果表明,我们的方法为从LLMs内部系统性地缓解合规风险提供了一条可执行路径。
cs.AI / 25 / 2609.39146
MADBench: Benchmarking the Security of Multi-Agent Debate
MADBench:多智能体辩论安全性的基准测试
large language model
大语言模型相关
Abstract
Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited. A central question is whether debate mitigates adversarial influence or amplifies it. In this paper, we present MADBench, a benchmark for evaluating the security of MAD. We organize attacks into a layered taxonomy following the MAD workflow, incorporating both established attacks and new strategies tailored to debate. We evaluate six attack families over 356 source tasks and 3,958 test cases, examining their effects on the final answer and the propagation of adversarial influence. Our results show that, under attacks, MAD does not necessarily improve LLM reasoning. Compared with a single-agent baseline, MAD can mitigate attacks on answer accuracy in question-answering tasks while amplifying unauthorized reads or writes in both question-answering and workspace tasks. Moreover, even when three out of five agents collude, the attack changes the final answer from correct to wrong on only 28.30\% of tasks answered correctly without attack, while only 3.26\% of initially correct honest agents switch to wrong answers during debate.
Chinese Translation
多智能体辩论(MAD)通过允许多个智能体就同一任务交换并批判彼此的答案,能够提升大语言模型(LLM)的推理能力。然而,使智能体能够纠正错误的交互也可能传播对抗性错误,并引导智能体走向错误答案。尽管已有一些工作致力于考察针对 MAD 的特定攻击类型,但在多种攻击下对 MAD 进行系统性评估仍然有限。一个核心问题是:辩论究竟是减轻对抗性影响,还是放大对抗性影响。在本文中,我们提出了 MADBench,一个用于评估 MAD 安全性的基准。我们按照 MAD 工作流将攻击组织为一个分层分类体系,既纳入已有的攻击,也纳入为辩论量身定制的新策略。我们在 356 个源任务和 3,958 个测试用例上评估了六类攻击,考察它们对最终答案的影响以及对抗性影响的传播。我们的结果表明,在攻击下,MAD 并不一定能提升 LLM 推理能力。与单智能体基线相比,MAD 可以在问答任务中减轻攻击对答案准确率的影响,同时在问答任务和工作区任务中放大未授权的读取或写入。此外,即使五个智能体中的三个合谋,攻击也仅在无攻击时被正确回答的任务中的 28.30\% 上使最终答案从正确变为错误;而最初回答正确的诚实智能体中,只有 3.26\% 在辩论过程中转变为错误答案。
cs.AI / 26 / 2609.39149
Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
Rep2Skill:面向LLM智能体的表示引导技能自演化
large language model
大语言模型相关
Abstract
Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.
Chinese Translation
文本技能使基于大语言模型(LLM)的智能体能够在不更新模型参数的情况下积累可复用的过程性知识。然而,现有的技能演化在很大程度上仍局限于文本空间:优化器必须仅依据冗长的执行轨迹和稀疏的任务结果来诊断成功与失败模式,并据此修订技能。这种纯文本范式将智能体的内部表示——其中包含其不断演化的执行状态的丰富记录——排除在技能优化回路之外。我们提出疑问:智能体能否通过反思自身的内部表示来改进其外部的文本技能?我们提出 Rep2Skill,一个用于智能体技能自演化的表示引导框架。具体而言,在收集到的智能体轨迹(rollouts)上,Rep2Skill 对其内部模型表示轨迹进行建模,以定位偏离成功执行动态的轮次,并进一步将这些信号与执行上下文一并解释为可用于针对性技能修订的可操作文本反馈。在两个智能体环境、两种开源 LLM 上的实验表明,在自演化设定中,Rep2Skill 持续优于纯文本方法;在该设定中,同一个 LLM 同时充当执行者与优化器,且不使用更强的外部模型。这确立了一个有前景的方向,推动智能体自我改进超越纯文本反思。
cs.AI / 27 / 2609.39168
Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation
通过 Token 级感知锚定优势估计强化多模态推理
large language model
大语言模型相关
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.
Chinese Translation
具有可验证奖励的强化学习(RLVR)提升了多模态大语言模型(MLLMs)的推理能力,然而现有框架依赖粗粒度的序列级奖励信号,这些信号缺乏对多模态推理链中以视觉为依据的步骤的细粒度监督。我们通过两个 Token 级指标的视角考察这一空白:视觉依赖性(即一个 Token 的预测在多大程度上依赖输入图像特征)和预测熵。我们的实证分析揭示了两个关键发现:(1)与错误推理链相比,正确推理链随着视觉锚定增强而表现出显著更急剧的熵下降;(2)关键 Token,即那些一旦被错误预测就会触发推理崩溃的 Token,在由正确推理链导出的视觉依赖性与预测熵的联合分布中是统计离群点。受这些发现启发,我们提出 Token 级感知锚定优势估计(TPAE),它通过衡量每个 Token 与正确 rollout 的视觉-熵模式之间的统计一致性来估计 Token 级优势。TPAE 利用这一细粒度分数来调制序列级优势,产生一种可集成到各种 RLVR 框架中的细粒度监督信号。在七个基准上的大量实验表明,TPAE 持续优于领先的强大基线,为多模态推理带来更稳定、更高效的优化。代码已公开提供:https://github.com/Zhihan72/TPAE。
cs.AI / 28 / 2609.39294
ANI: Adaptive Numerical Injection for Unifying Semantic and Arithmetic Representations in Numerical Reasoning
ANI:用于统一数值推理中语义与算术表示的自适应数值注入
large language model
大语言模型相关
Abstract
Precise numerical reasoning with Large Language Models (LLMs) is essential for expanding their applicability to complex real-world tasks. However, text-based tokenization often fragments numbers, significantly hindering precise arithmetic reasoning. Meanwhile, numerical embeddings, despite arithmetic precision, rely on context-agnostic substitution that disregards the semantic role of numbers as identifiers. To combine the complementary strengths, we propose \textbf{ANI (Adaptive Numerical Injection)}, a hybrid framework that governs the selective injection of numerical features based on the semantic context. By employing a context-aware gating mechanism, we selectively inject numerical embeddings (specifically FoNE) into the latent space, explicitly preserving nominal identifiers while enhancing quantitative operands. Through extensive evaluations across various LLMs, we demonstrate that ANI enhances MATH performance by 9.5 points over the official reference model, while maintaining robust performance on general linguistic benchmarks.
Chinese Translation
使用大型语言模型(LLMs)进行精确的数值推理,对于将其适用性扩展到复杂的现实世界任务至关重要。然而,基于文本的分词往往会将数字碎片化,从而显著阻碍精确的算术推理。与此同时,数值嵌入尽管具有算术精度,却依赖于与上下文无关的替换,这忽视了数字作为标识符的语义角色。为了结合互补优势,我们提出了 \textbf{ANI(自适应数值注入)},这是一个混合框架,基于语义上下文控制数值特征的选择性注入。通过采用上下文感知的门控机制,我们将数值嵌入(具体为 FoNE)选择性地注入到潜在空间中,在增强定量操作数的同时,明确保留名义标识符。通过对各种 LLM 的广泛评估,我们证明 ANI 在 MATH 性能上比官方参考模型提升了 9.5 分,同时在通用语言基准上保持稳健的性能。
cs.AI / 29 / 2609.39297
MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate
MiniRep:多智能体辩论中鲁棒的基于声誉的聚合
large language model
大语言模型相关
Abstract
Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may not reliably predict an agent's behavior on new tasks, particularly when malicious agents can adapt their behavior and influence other agents during collaboration. We study reputation in multi-agent debate (MAD), where multiple agents answer the same query, debate to improve their answers, and aggregate them into a final output. We present MiniRep, a reputation-based aggregation system for MAD under malicious agents. To ground our threat model in established research, we construct an attack taxonomy drawing on reputation-system attacks and software-testing mutation operators, covering strategic exploitation of reputation and subtle corruption of agent proposals. Guided by this taxonomy, MiniRep evaluates agents based on both their behavior on the current task and their reputation over time, while preventing groups of agents with highly similar responses from dominating the final decision. We assess MiniRep across diverse tasks, LLM-agent compositions, corruption placements, and attack types drawn from our taxonomy. Our experimental results show that, MiniRep outperforms both conventional MAD aggregation and conventional reputation-based approaches on MATH no matter being attacked or not. Also, under a heterogeneous 10-agent setting on MATH, MiniRep outperforms all baselines in all 28 attack conditions.
Chinese Translation
由大型语言模型(LLM)驱动的自主智能体正在迅速演化为一个开放的智能体生态系统。为了支持可信协作,行业举措越来越多地根据过去行为评估智能体声誉,并提供性能排行榜。然而,由过去表现推导出的声誉可能无法可靠预测智能体在新任务上的行为,特别是当恶意智能体能够在协作期间调整其行为并影响其他智能体时。我们研究多智能体辩论(MAD)中的声誉,其中多个智能体回答同一查询,进行辩论以改进其答案,并将它们聚合为最终输出。我们提出 MiniRep,一个面向存在恶意智能体的 MAD 的基于声誉的聚合系统。为使我们的威胁模型建立在既有研究之上,我们构建了一个攻击分类法,借鉴声誉系统攻击和软件测试变异算子,涵盖对声誉的策略性利用以及对智能体提案的细微破坏。在该分类法的指导下,MiniRep 基于智能体在当前任务上的行为及其随时间变化的声誉来评估智能体,同时防止具有高度相似响应的智能体群体主导最终决策。我们在多样化的任务、LLM 智能体组成、破坏位置以及取自我们分类法的攻击类型上评估 MiniRep。我们的实验结果表明,无论是否受到攻击,MiniRep 在 MATH 上都优于传统 MAD 聚合和传统基于声誉的方法。此外,在 MATH 上的异构 10 智能体设置下,MiniRep 在所有 28 种攻击条件下均优于所有基线。
cs.AI / 30 / 2609.39325
WorkGenesis: Building the Worlds That Teach Agents to Work
WorkGenesis:构建教会智能体工作的世界
large language model
大语言模型相关
Abstract
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.
Chinese Translation
大型语言模型(LLM)智能体完成日常和专业工作的能力正受到越来越多的关注。训练此类智能体需要真实的工作场景。由专家编写的职业工作成本高昂且生产缓慢,而无约束的合成往往产生事实基础薄弱或需求内部不一致的任务。为弥合这一差距,我们提出 WorkGenesis,一个通过两项核心技术从真实世界工件构建可执行职业工作的框架:(1)基于证据的工作构建,它通过以 O*NET 职业知识为指导检索公开文件,并围绕这些文件合成周围上下文、配套材料、工作请求和逐项评分标准,从而将每个工作单元建立在真实世界证据之上;(2)执行引导的一致性验证,它在构建的工作内渲染一个参考交付物,将每个未满足的评分标准项归因于智能体、任务或评分标准,并将任务和评分标准缺陷作为反馈,迭代修复该工作,直到其通过审核。实验结果表明,Fx-Work-35B 仅使用简单的监督微调(SFT)在 WorkGenesis 合成的 20K 个工作单元上进行训练,就在 GDPvalAA-v2、APEX-Agents-AA 和 JobBench 上报告的五个指标中,在所有规模相当的基线中取得了最高分(平均分 31.00 对 24.79),甚至超过了 1.6T DeepSeek-V4-Pro-Preview 等前沿模型。这些结果表明,WorkGenesis 为工作型智能体提供了可扩展的训练数据。
cs.AI / 31 / 2609.39343
The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching
黄金路径假设:扩散缓存中的可复用调度
diffusion
扩散模型相关
Abstract
Diffusion caching accelerates generation by replacing transformer computation with cached or predicted features at selected denoising steps. We introduce the Golden Path Hypothesis (GPH): under fixed inference conditions, prompt-independent cache schedules can achieve final-output quality comparable to the best prompt-specific schedules across prompts. We investigate the GPH across ten caching methods, four image and video models, and three cache ratios. Prompt-adaptive methods repeatedly select a small number of schedules, and reusing their most frequent schedules on new prompts closely matches the quality of prompt-specific choices. Exhaustive evaluation of 1.4 million schedules on four examples further identifies prompt-independent schedules that remain competitive on unseen prompts. To explain this transfer, we analyze denoising trajectories and the accumulation of caching errors. Latent-state trajectories exhibit similar structures across datasets and seeds, while an exact error decomposition shows that accumulated effects of earlier errors predict final latent-state error better than local approximation errors. This motivates searching for end-to-end schedules using final-output quality. With only a small set of examples, the resulting golden paths transfer across prompts and datasets, and can be tuned to the desired quality objective, including reconstruction fidelity or perceptual similarity.
Chinese Translation
扩散缓存通过在选定的去噪步骤上用缓存或预测的特征替换 transformer 计算来加速生成。我们提出黄金路径假设(GPH):在固定推理条件下,与提示无关的缓存调度能够在不同提示上达到与最佳提示特定调度相当的最终输出质量。我们在十种缓存方法、四个图像与视频模型以及三种缓存比例上研究了 GPH。提示自适应方法会反复选择少量调度,而在新提示上复用它们出现频率最高的调度,其质量与提示特定选择的质量非常接近。对四个样例上的 140 万个调度进行穷举评估,进一步识别出在未见提示上仍具竞争力的与提示无关的调度。为解释这种迁移,我们分析了去噪轨迹以及缓存误差的累积。潜在状态轨迹在不同数据集和随机种子下表现出相似的结构,而一个精确的误差分解表明,较早误差的累积效应对最终潜在状态误差的预测效果优于局部近似误差。这促使我们使用最终输出质量来搜索端到端调度。仅用少量样例,所得到的黄金路径即可跨提示和数据集迁移,并且可以针对所需的质量目标进行调优,包括重建保真度或感知相似度。
cs.AI / 32 / 2609.39382
SkillFM: Generating Skills for LLM Agents via Latent Flow Matching
SkillFM:通过潜在流匹配为LLM智能体生成技能
large language model
大语言模型相关
Abstract
Textual skills provide reusable guidance for large language model agents, but existing approaches often rely on manually curated skill banks or reinforcement learning with indirect and delayed feedback. We introduce SkillFM (Skill Flow Matching), a generative framework that synthesizes task-conditioned textual skills directly without test-time skill retrieval. Our framework combines a codec for encoding and reconstructing textual skills in a continuous latent space with a conditional flow model trained using improved MeanFlow. At inference time, the learned velocity field enables single-step latent sampling, and an LLM-based decoder converts the sampled representation into textual guidance for a frozen downstream agent. We evaluate the framework on embodied tasks, question answering, and web shopping. On ALFWorld and Search-QA, our method achieves the best overall performance among the compared vector-based skill approaches. Our analyses further demonstrate that latent skill generation is an effective alternative to retrieval-based skill augmentation. Our code and training skill libraries are available at https://github.com/lulushang999/SkillFM.
Chinese Translation
文本技能为大型语言模型智能体提供了可复用的指导,但现有方法往往依赖人工整理的技能库,或依赖具有间接且延迟反馈的强化学习。我们提出 SkillFM(Skill Flow Matching),这是一个生成式框架,可直接合成以任务为条件的文本技能,而无需在测试时进行技能检索。我们的框架结合了一个用于在连续潜在空间中对文本技能进行编码与重建的编解码器,以及一个使用改进的 MeanFlow 训练的条件流模型。在推理时,学习到的速度场能够实现单步潜在采样,而基于LLM的解码器会将采样得到的表示转换为用于冻结下游智能体的文本指导。我们在具身任务、问答和网页购物上评估了该框架。在 ALFWorld 和 Search-QA 上,我们的方法在所比较的基于向量的技能方法中取得了最佳整体性能。我们的分析进一步表明,潜在技能生成是基于检索的技能增强的一种有效替代方案。我们的代码和训练技能库可在 https://github.com/lulushang999/SkillFM 获取。
cs.AI / 33 / 2609.39394
Can Computation from Earlier Problems Help LLMs Solve New Ones?
来自先前问题的计算能否帮助 LLM 解决新问题?
large language model
大语言模型相关
Abstract
Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.
Chinese Translation
大型语言模型经常在同一对话中解决相互独立的问题。来自先前问题的计算能帮助它们解决新问题吗?为了回答这个问题,我们首先进行初步实验,表明保留的历史可以提高或降低后续轮次的准确率,即使在同一个领域内也是如此。为了理解这些效应,我们使用受控重放来分离出特定于每个问题-历史配对的内部状态变化。在不同的历史之间,这些变化保留了当前问题之间的相似关系。为了改进在保留历史下的推理,我们提出了 STAIR(用于查询间复用的陈旧 Token 注意力)。STAIR 将来自先前响应生成的键和值捕获到一个固定的库中。它学习在当前查询于提示处理期间读取该库时重定向这些查询。基础模型保持冻结;仅训练 12,288 个参数。在三个 Qwen 模型和四个基准上,与带历史的未修改模型相比,STAIR 将平均后续轮次准确率最多提高了 11.67 个百分点。
cs.AI / 34 / 2609.39406
Inferring Causal Relations between Two Sequences of Events with Language Models
使用语言模型推断两个事件序列之间的因果关系
large language model
大语言模型相关
Abstract
Causal AI is a branch of Artificial Intelligence which helps understand and reason about cause and effect relationships, not just patterns or correlations. Causal discovery aims to infer elements of the underlying causal structure--often represented as a directed graph--from observational and, when available, interventional data. While causal discovery is the fundamental step for moving beyond mere associations toward genuine understanding, and thus the basic building block of causal AI, it becomes intrinsically difficult when causal relations must be inferred from single observations. In such situations, standard causal discovery methods cannot be used and one has to identify causal relations from limited amount of information. This is typically the case for, e.g., sequences of events produced by different alarms which need to be analyzed on the fly to detect abnormal phenomena, which are usually rare. We show in this study that it is possible to leverage the predictive power of Large Language Models (LLMs) to infer causal relations between only two sequences of events. This approach, which is validated on both synthetic and real data, provides better results than standard causal discovery algorithms on several time series data, even though these data were converted into smaller, single observed sequences.
Chinese Translation
因果AI是人工智能的一个分支,它有助于理解和推理因果关系,而不仅仅是模式或相关性。因果发现旨在从观测数据以及(在可用时)干预数据中推断潜在因果结构的要素——通常表示为有向图。尽管因果发现是从单纯的关联走向真正理解的根本步骤,并且因此是因果AI的基本构建模块,但当必须从单次观测中推断因果关系时,它会变得本质上困难。在这种情况下,标准的因果发现方法无法使用,人们必须从有限的信息量中识别因果关系。这通常是以下情况,例如,由不同警报产生的事件序列,需要实时分析以检测通常罕见的异常现象。我们在本研究中表明,可以利用大型语言模型(LLM)的预测能力来推断仅两个事件序列之间的因果关系。这种方法在合成数据和真实数据上都得到了验证,在若干时间序列数据上比标准因果发现算法提供了更好的结果,尽管这些数据被转换为更小的、单一观测序列。
cs.AI / 35 / 2609.39483
Who Owns That? Evaluating Ownership Intuitions in Large Language Models
那是谁的?评估大型语言模型中的所有权直觉
large language model
大语言模型相关
Abstract
Ownership establishes rights over the use, control, and transfer of objects. Understanding these relations is essential for AI systems to interact appropriately with people and their resources. Yet how large language models (LLMs) attribute ownership under competing claims remains unclear. We introduce the Competing Ownership Attribution Task (COAT), comprising 42 scenarios, and compare ownership allocations from 24 LLM configurations with those of 108 human participants. Overall, human-model similarity is close to human-human similarity, but models show greater homogeneity in their ownership judgments. Within individual answers, models also divide ownership more evenly among claimants than humans do. Pooling responses across model configurations reveals more scenarios with a shared judgment and fewer with distinct viewpoint groups than in humans. When humans form distinct groups, models may converge on one viewpoint or between competing viewpoints. Further comparisons reveal different contextual sensitivities. As material value increases across scenarios, allocations to creators decline less sharply in models than in humans. Across scenarios differing in public recognition of later holders as owners, allocations to these holders increase in models but decrease slightly in humans. Together, these findings suggest that the evaluated LLM responses do not fully capture the diversity of participants' ownership judgments or how those judgments vary across situations. Developing socially capable AI therefore requires moving beyond overall similarity to capture the diversity and context dependence of human judgments.
Chinese Translation
所有权确立了对物品的使用、控制和转让的权利。理解这些关系对于AI系统恰当地与人及其资源互动至关重要。然而,大型语言模型(LLM)在相互竞争的主张下如何归属所有权仍不清楚。我们引入了竞争性所有权归属任务(COAT),包含42个场景,并将24种LLM配置的所有权分配与108名人类参与者的所有权分配进行比较。总体而言,人-模型相似度接近人-人相似度,但模型在其所有权判断上表现出更大的同质性。在单个回答中,模型在所有权主张者之间分配所有权也比人类更均匀。汇总不同模型配置的回答后,与人类相比,模型在更多场景中呈现出共同判断,而在更少场景中呈现出不同的观点群体。当人类形成不同群体时,模型可能收敛于一种观点,或收敛于相互竞争的观点之间。进一步的比较揭示了不同的情境敏感性。随着各场景中物质价值增加,模型对创造者的分配下降得不如人类那么急剧。在公众对后续持有者作为所有者的认可程度不同的场景中,对这些持有者的分配在模型中增加,但在人类中略有下降。综合来看,这些发现表明,所评估的LLM回答并未充分捕捉参与者所有权判断的多样性,也未充分捕捉这些判断如何随情境变化。因此,开发具有社会能力的AI需要超越总体相似性,以捕捉人类判断的多样性和情境依赖性。
cs.AI / 36 / 2609.39604
Why Do Conventional World Models Fail to Learn Cellular Automata?
为什么传统世界模型无法学习元胞自动机?
diffusion
扩散模型相关
Abstract
Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in many cases. Conventional architectures predict most pixels correctly yet rarely complete a rollout: a CNN predicts 96.3% of cells but completes 18.9% of rollouts; a joint diffusion model completes none. We trace the gap to three failure modes of these world models - namely, they fail to exactly capture spatial locality, temporal locality or temporal stability. Simple changes repair each: (1) for spatial locality, two-dimensional rotary positions lift a transformer from 39.1% to 100% on the Game of Life; (2) for temporal locality, handing each token its cell's previous-frame neighbourhood lifts the same transformer from 25.8% to 99.9% on unseen rules; (3) for temporal stability, causal freezing lifts the same diffusion weights from 42.2% to 99.9%. None of the three changes touches the architectural backbone; each only modifies the information flow within it. We also compare joint and ordered sampling on billiards and, in an exploratory study, on a simulated Burgers equation.
Chinese Translation
尽管传统世界模型——基于 Transformer 或卷积网络的自回归模型或扩散模型——可能学到世界动力学的表面统计量,但它们能否从其观测历史中学习到精确的世界动力学?利用元胞自动机作为一个简单测试平台,我们发现在许多情况下答案是否定的。传统架构能正确预测大多数像素,却很少能完成一次完整推演:一个 CNN 能预测 96.3% 的元胞,但只完成 18.9% 的推演;一个联合扩散模型则一次都未完成。我们将这一差距追溯到这些世界模型的三种失效模式——即它们未能精确捕获空间局部性、时间局部性或时间稳定性。简单的改动即可修复每一种失效:(1)针对空间局部性,二维旋转位置将 Transformer 在生命游戏上的表现从 39.1% 提升到 100%;(2)针对时间局部性,将每个 token 对应元胞的前一帧邻域交给该 token,使同一 Transformer 在未见规则上的表现从 25.8% 提升到 99.9%;(3)针对时间稳定性,因果冻结将同一扩散权重从 42.2% 提升到 99.9%。这三种改动均未触及架构主干;每一种都只修改了其中的信息流。我们还在台球问题上比较了联合采样与有序采样,并且,在一项探索性研究中,在模拟的 Burgers 方程上进行了比较。
cs.AI / 37 / 2609.39863
FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
FIGS:在不惩罚共情的情况下评估多轮谄媚
large language model
大语言模型相关
Abstract
Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views. In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time. Current evaluations, however, rely on rigid, single-turn tests or fixed scripts that fail to capture these natural dynamics. Furthermore, these benchmarks often mistake showing basic empathy for yielding, penalizing models for acknowledging a user's feeling. This view may drive future models to over-correct into cold, dismissive rigidity. To address this gap, we introduce FIGS (Factual Integrity and Grounded Support), a dual-axis evaluation framework built around extended, realistic dialogue. We use an adaptive 10-turn conversational simulator that dynamically challenges the target model, reflecting how users repeat requests, push back, or steer a conversation toward a preferred answer. To accurately evaluate these trajectories, we apply a taxonomy that strictly separates Sycophancy (whether the model holds firm to the truth and keeps its praise proportional) from Calibrated Validation (showing empathetic understanding of the user's feelings without overdoing it). We release our complete testing environment, including 500 diverse multi-turn scenarios and an automated judge. Our evaluation of leading models reveals a consistent trade-off: over the course of a sustained interaction, current systems either slowly drift to sycophancy or over-correct into robotic detachment. This demonstrates that balancing honesty with appropriate support throughout a natural conversation remains a critical, unsolved challenge.
Chinese Translation
大型语言模型常常无法在保持真实与提供支持之间取得平衡。它们经常在回应用户时表现出谄媚,赞同虚假说法,给予无根据的奉承,并提供偏向用户所表达观点的建议。现实中,谄媚很少发生在单次交流中;它可能在用户反复坚持或随着时间微妙引导对话的过程中自然出现。然而,当前的评估依赖僵化的单轮测试或固定脚本,无法捕捉这些自然动态。此外,这些基准常常把表现出基本共情误认为让步,因模型承认用户感受而对其施加惩罚。这种观点可能促使未来模型过度纠正,变得冷漠、轻蔑且僵化。为弥补这一空白,我们提出 FIGS(Factual Integrity and Grounded Support,事实完整性与有根据的支持),一个围绕扩展的、现实对话构建的双轴评估框架。我们使用一个自适应的 10 轮对话模拟器,它会动态挑战目标模型,反映用户如何重复请求、反驳,或将对话引导向一个偏好的答案。为了准确评估这些轨迹,我们应用一套分类体系,将谄媚(模型是否坚持真相,并使其赞扬保持相称)与校准性认可(表现出对用户感受的共情理解而不过度)严格区分开来。我们发布完整的测试环境,包括 500 个多样化的多轮场景和一个自动评判器。我们对领先模型的评估揭示出一个一致的权衡:在持续交互过程中,当前系统要么缓慢滑向谄媚,要么过度纠正为机器人般的疏离。这表明,在自然对话中始终平衡诚实与适当支持仍然是一项关键且尚未解决的挑战。
cs.AI / 38 / 2609.40303
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
一个强智能体为自主 ML 工程需要多少工具框架?
large language model
大语言模型相关
Abstract
Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.
Chinese Translation
最近的自主机器学习工程(MLE)智能体在公开排行榜上取得了显著进展。现代 MLE 智能体常常出于长时程周期中的进展停滞以及有限的大语言模型(LLM)原语这一动机,被部署在日益复杂的机制之上:多智能体编排器、专用检索子智能体等。尽管此类工具框架不断扩展,但使用更原始但有所改进的编码智能体——其中 LLM 通过 read、write 和 bash 原语直接访问执行环境——在该领域却很少受到关注。在本文中,我们发现,在相同时间预算和相同前沿 LLM 骨干模型下,开源的最先进工具框架相较单次会话的最小工具框架编码智能体基线没有任何优势,这表明骨干模型才是性能的主要驱动因素。通过一系列大规模系统性消融研究,我们认为,在编码智能体设定中,这些机制层变得多余。我们得出结论:围绕强模型精心构建手工设计的工具框架所花费的努力,对当前 MLE 基准而言回报微薄。
cs.AR / 39 / 2609.38601
Structure-augmented LLMs for High-Level Synthesis Pragma Optimization
面向高层次综合 Pragma 优化的结构增强 LLM
large language model
大语言模型相关
Abstract
Pragma insertion drives the quality of high-level synthesis (HLS) designs. Choosing the right directives demands expert knowledge and reasoning about loop nesting, data dependences, and memory layout. While existing large language models (LLMs) show promise in code generation, they lack explicit program-structure awareness, limiting their ability to suggest effective pragmas. We present PRISM, a novel structure-augmented LLM that closes this gap by adding compiler-grade structural reasoning to a pretrained, frozen code LLM. It combines three hierarchical program representations, Abstract Syntax Tree (AST), Control-Flow Graph (CFG), and Data-Flow Graph (DFG), injecting them into a specific transformer layer while keeping original code tokens in a separate stream. The cross-attention gate at the injection point allows falling back to the pretrained representation when its structural signal is unhelpful. On zero-shot evaluation in HLS-Eval, PRISM synthesizes 3.5\times as many kernels as Llama3-8B (26.9\% vs. 7.7\%), and on the kernels where it does succeed, it produces designs that are 2.31\times faster (geomean) than GPT-5-mini's. In the agentic flow, the PRISM codegen outperforms other baselines when optimizing complex code and drives the average normalized improvement across the HLS-Eval suite to 26.4\%.
Chinese Translation
Pragma 插入驱动着高层次综合(HLS)设计的质量。选择正确的指令要求具备关于循环嵌套、数据依赖和内存布局的专家知识与推理能力。虽然现有大语言模型(LLM)在代码生成方面显示出前景,但它们缺乏显式的程序结构感知,限制了它们建议有效 pragma 的能力。我们提出了 PRISM,一种新颖的结构增强 LLM,它通过向一个预训练、冻结的代码 LLM 加入编译器级结构推理来弥合这一差距。它结合了三种分层程序表示:抽象语法树(AST)、控制流图(CFG)和数据流图(DFG),将它们注入到特定的 transformer 层,同时将原始代码 token 保留在单独的流中。注入点处的交叉注意力门允许在其结构信号无帮助时回退到预训练表示。在 HLS-Eval 的零样本评估中,PRISM 合成的内核数量是 Llama3-8B 的 3.5\times(26.9\% 对 7.7\%),而在其确实成功的内核上,它生成的设计比 GPT-5-mini 的设计快 2.31\times(几何平均)。在智能体流程中,PRISM 代码生成在优化复杂代码时优于其他基线,并将整个 HLS-Eval 套件的平均归一化改进提升到 26.4\%。
cs.CL / 40 / 2609.38334
EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
EVOKE:激发智能体中的世界知识以实现可迁移决策
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
Chinese Translation
大语言模型(LLMs)正越来越多地被部署为执行多步决策的智能体,但它们在向未见环境迁移时表现不佳。世界模型方法通过训练智能体预测未来观测来应对这一问题,代价则是额外的训练,以及当预测被用于规划时会不断累积的误差。然而,对于在数字环境中运行的 LLM 智能体而言,这类世界知识中的大部分在预训练期间就已经被内化,这使问题从获取世界知识转变为激发世界知识。我们认为,典型的后训练几乎没有为这种激发提供压力,因为在每个被访问到的状态下、单一目标下的监督会无意间驱使策略依赖于表层的上下文习惯。我们提出 EVOKE,一种通过在固定状态下提供目标多样性来施加这种压力的后训练方法。受以下理论的启发——该理论表明,一个能够胜任多样目标的智能体必须编码一个可从其动作偏好中恢复的世界模型——EVOKE 保持环境状态与交互历史固定,并在不同替代目标下对同一组候选动作进行排序,迫使动作偏好发生变化,从而使依赖上下文习惯或单一目标相关性的策略无法正确地对其排序。这隐式地激发了策略所预训练获得的世界知识来为决策提供信息。我们在三个主干模型上跨多样任务评估 EVOKE,展示了任务性能、未见环境泛化能力以及数据效率方面的提升。我们进一步开展受控分析,以更好地理解是什么驱动了这些收益。这些发现为通过直接决策监督来激发内化的世界知识以实现可迁移动作提供了新的视角。
cs.CL / 41 / 2609.38374
Doc2LoRA Provides Decodable Representations of Scientific Ideas
Doc2LoRA 提供科学思想的可解码表示
large language model
大语言模型相关
Abstract
Representing scientific papers as points in a space lets us search for similar papers and inquire about how fields relate to one another and drive innovation. Beyond search, the vector space of papers invites generation: mixing papers through simple vector operations creates new points, mirroring combinatorial novelty, the recombination of existing ideas into new ones. However, a mixed point often represents an idea no paper has yet realized, with no papers nearby to identify the idea. We propose representing each paper by a LoRA adapter generated by the Doc-to-LoRA hypernetwork. Every point in the space, including mixtures, thus represents a large language model (LLM) open to questions and instructions in natural language. On papers from the American Physical Society (APS), we instruct the LLM at the average of each subfield to name the field in a few words and obtain labels closer to the official names than the labels of five baselines, as judged by word overlap and a panel of five LLM judges. We also ask the LLMs at points between two APS papers to write an abstract and obtain descriptions shifting from one paper to the other in step with the mixing weight. While Doc-to-LoRA is trained for generation, a small invertible transform makes the embeddings competitive for search, on par with SPECTER2 and EmbeddingGemma and close to SBERT. Because the transform is invertible, every point in the transformed space still maps back to an LLM. The embeddings thus serve both search and generation, enabling researchers to question the idea at any point in the space as a starting point for generating new ideas.
Chinese Translation
将科学论文表示为空间中的点,使我们能够搜索相似的论文,并探究各个领域之间如何相互关联以及如何驱动创新。除搜索之外,论文的向量空间还启发着生成:通过简单的向量运算混合论文会创造出新的点,映射出组合式新颖性,即把已有想法重组为新的想法。然而,一个混合点往往代表着一个尚无任何论文实现的想法,附近没有论文可以用来指认这一想法。我们提出用由 Doc-to-LoRA 超网络生成的 LoRA 适配器来表示每篇论文。因此,空间中的每一个点,包括混合点,都代表着一个对自然语言问题和指令开放的大语言模型(LLM)。在美国物理学会(APS)的论文上,我们指示位于每个子领域平均位置的 LLM 用几个词为该领域命名,并以词重叠度和由五个 LLM 评审组成的评审小组作为评判依据,得到的标签比五个基线的标签更接近官方名称。我们还要求位于两篇 APS 论文之间各点上的 LLM 撰写摘要,得到的描述会随混合权重同步地从一篇论文向另一篇论文过渡。尽管 Doc-to-LoRA 是为生成而训练的,但一个小的可逆变换使这些嵌入在搜索任务上具有竞争力,与 SPECTER2 和 EmbeddingGemma 相当,并接近 SBERT。由于该变换是可逆的,变换后空间中的每一个点仍然能够映射回一个 LLM。因此,这些嵌入同时服务于搜索和生成,使研究人员能够就空间中任意一点的想法进行提问,以此作为生成新想法的起点。
cs.CL / 42 / 2609.38406
Evaluating Whether LLMs Can Reliably Connect the DOTs?
评估大语言模型能否可靠地连接这些点?
large language model
大语言模型相关
Abstract
Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models (LLMs), their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ~9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score (4.02/5), outperforming models over ten times larger, including DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement (+0.6%). Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.
Chinese Translation
获取现实世界信息通常充满噪声且碎片化。从这些碎片中构建连贯叙事,要求模型在更宏大的故事情节中重建缺失的片段,这通常被称为文本填充,同时保持与局部上下文和全局故事情节的一致性。尽管许多大语言模型(LLMs)将文本填充用作预训练目标,但它们在现实世界叙事填充上的实际表现仍未被充分探索。在本文中,我们通过引入一个约 9.2K 个实例的多领域叙事填充基准来填补这一空白,该基准通过遮蔽四种叙事类型中的一到三个句子构建:百科文本、常识故事、新闻文章和视觉叙事。使用该基准,我们评估了 20 个经过指令微调的开源 LLMs,参数规模从 1.5B 到 70B 不等,涵盖不同水平的指令具体性和推理引导。输出使用标准自动指标以及覆盖五个叙事维度的定性框架进行评估。结果表明,模型规模并不能可靠地预测填充质量:Gemma-2-2B 取得了最高的定性得分(4.02/5),优于规模大十倍以上的模型,包括 DeepSeek-Qwen-32B(3.77/5,6.6%)和 LLaMA-3.3-70B(3.71/5,8.3%)。我们进一步发现,显式推理带来的益处有限,因为思维链推理仅带来边际提升(+0.6%)。此外,对于当前 LLMs 的叙事填充而言,短叙事和领域特征显现为比单独的填充位置更强的任务难度预测因素。
cs.CL / 43 / 2609.38490
Personalized State-Transition-Aware Memory for Clinical Agents
面向临床智能体的个性化状态转移感知记忆
large language model
大语言模型相关
Abstract
Large language model (LLM) agents that reason over clinical records must track changes in a patient's state while preserving the history needed to understand them. Simply accumulating memories leaves it unclear which information still applies, whereas overwriting earlier memories can erase evidence needed to reconstruct treatment history and clinical trajectories. We introduce STAM, a state-transition-aware memory framework that records state changes as new clinical entries arrive. STAM combines semantic retrieval with typed clinical relations to identify affected memories, maintaining current information in Active and superseded or resolved information in History. At read time, a query-dependent gate selectively serves historical memory. Across four longitudinal clinical benchmarks, we evaluate STAM with downstream question answering, direct state-maintenance diagnostics, and comparisons at approximately matched context lengths.
Chinese Translation
在临床记录上进行推理的大语言模型(LLM)智能体必须追踪患者状态的变化,同时保留理解这些变化所需的历史信息。仅仅不断累积记忆会使人无法确定哪些信息仍然适用,而覆盖早先的记忆则可能抹去重建治疗史与临床轨迹所需的证据。我们提出 STAM,一种状态转移感知的记忆框架,它在新临床条目到来时记录状态变化。STAM 将语义检索与带类型的临床关系相结合,以识别受影响的记忆,将当前信息保存在 Active 中,将被取代或已解决的信息保存在 History 中。在读取时,一个依赖于查询的门控有选择地提供历史记忆。在四个纵向临床基准上,我们通过下游问答、直接的状态维护诊断,以及在上下文长度大致匹配条件下的比较来评估 STAM。
cs.CL / 44 / 2609.38510
DEdit: Iterative Draft Editing for Speculative Decoding
DEdit:用于推测解码的迭代草稿编辑
diffusion
扩散模型相关
Abstract
Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks on Qwen3-4B and Qwen3-8B, DEdit achieves the highest macro-average token acceptance and speedup among the evaluated drafters, reaching macro-average speedups of $5.72\times$ and $5.97\times$ over autoregressive generation under greedy decoding, respectively. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and that ProposalMix halves harmful edits that shorten the accepted prefix. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains.
Chinese Translation
推测解码通过让一个轻量级草稿模型提出由目标模型并行验证的 token 来加速自回归 LLM。基于扩散的草稿模型通过一次性提出多个 token 进一步降低草稿延迟。然而,这些 token 是独立预测的,因此一个早期的单个错误就会导致前缀验证丢弃草稿的其余部分,即使其中包含有用的下游预测。我们提出 DEdit,一种基于扩散的草稿模型,它不仅可以通过传统的并行解掩码进行草稿,还可以通过 token 到 token 的预测对其草稿进行迭代编辑。通过编辑,后面的预测可以作为双向上下文,用于修复较早的错误并延长已接受前缀。为了教会模型在保留正确预测的同时修复错误,我们提出 ProposalMix,一种训练方案,它在训练期间基于第一遍置信度将草稿预测与真值 token 混合。在 Qwen3-4B 和 Qwen3-8B 的七个基准上,DEdit 在被评估的草稿模型中取得了最高的宏平均 token 接受率和加速比,在贪婪解码下分别相对于自回归生成达到 $5.72\times$ 和 $5.97\times$ 的宏平均加速比。进一步分析表明,接受率会随着更多编辑轮次和更宽的草稿窗口而提高,并且 ProposalMix 将缩短已接受前缀的有害编辑减半。此外,将编辑器限制为因果注意力会降低接受率,尤其是在高度可预测的输出上,这表明未来上下文是这些收益的一个关键来源。
cs.CL / 45 / 2609.38543
MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models
MedKIT:评估大型语言模型中的知识整合与泛化
large language model
大语言模型相关
Abstract
Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.
Chinese Translation
不断演变的现实世界知识要求模型持续更新。尤其在医学中,随着临床证据随时间变化,过时的知识可能带来安全风险。现有的知识整合评估聚焦于事实回忆,对于新整合的知识是否真正可用所提供的洞见有限。我们的基准 MedKIT(医学知识整合与迁移)针对模型在真实临床更新序列下如何整合和应用知识,提供了细粒度评估。每个实例对应于从临床证据中得出的事实更新,并配有针对性的探针,这些探针评估在词汇变化、关系变换、组合推理和开放式操作化方面的迁移,以及用于知识保持的局部性测试。使用 MedKIT,我们对 12 种知识整合策略在 5 个多样模型上进行了大规模实证研究,这些模型包括通用和医学大型语言模型。我们的结果揭示了回忆与可用知识之间存在一致的差距:尽管大多数方法在原始更新任务和词汇变化下取得显著提升,关系泛化仍然有限,且没有任何方法在组合任务或操作任务上产生有意义的改进。这些发现凸显了知识整合中的一个根本性挑战,并将 MedKIT 定位为一个测试平台,用于开发使新整合的知识在任务和情境中更一致可用的方法。
cs.CL / 46 / 2609.38593
Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions
Prompt2Skill:从自然语言指令进行无监督技能优化
large language model
大语言模型相关
Abstract
Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training. In addition, emerging tasks may fall outside the scope of existing skill libraries, creating a need to develop new skills before curated training data become available. Recent works have explored automated skill optimization through reflection, but they require a curated, in-distribution training set, which users might not always have. To address these limitations, we present Prompt2Skill, a framework that builds skills from natural-language task description alone. From the prompt, the system derives a task specification, discovers or synthesizes datasets, and refines the skill in a closed loop of reflective editing. Across four domains spanning question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill consistently outperforms the direct prompting baseline, achieving an average improvement of 10.8 across open-source and frontier models.
Chinese Translation
技能是外部产物,大型语言模型(LLM)在推理时使用它们,通过纳入相关的程序性知识和领域知识来提升其在专业领域中的表现。由专家编写的技能制作成本高昂,而且由此产生的产物并未针对使用它们的特定模型进行优化,而该模型的失效模式会随版本、规模和训练而变化。此外,新兴任务可能超出既有技能库的范围,从而产生了在精心整理的训练数据可用之前就开发新技能的需求。近期研究探索了通过反思进行自动化技能优化,但它们需要精心整理的、同分布的训练集,而用户未必总能拥有这样的训练集。为解决这些局限,我们提出 Prompt2Skill,一个仅凭自然语言任务描述即可构建技能的框架。从该提示出发,系统推导出任务规范,发现或合成数据集,并在反思性编辑的闭环中改进技能。在涵盖问答、阅读理解、电子表格操作和数学推理的四个领域中,Prompt2Skill 始终优于直接提示基线,在开源模型和前沿模型上平均提升 10.8。
cs.CL / 47 / 2609.38718
MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation
MetaSteer:通过注意力投影适配实现的上下文条件化非线性引导
large language model
大语言模型相关
Abstract
Steering large language models typically relies on linear, context-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions. We introduce MetaSteer, a method that learns nonlinear interventions with context-dependent effects and applies them to attention projection matrices, producing activation effects that vary with the input context by construction and requiring no linear concept-geometry assumption. Framed as preference-based optimization, MetaSteer is trained once on a pooled preference corpus and transferred zero-shot to unseen concepts and out-of-distribution contexts. We find that, despite using low-rank adapters, MetaSteer induces structured, context-dependent changes in hidden-state trajectories while partially preserving aspects of their local trajectory dynamics, including velocity and curvature. We evaluate MetaSteer on three controlled text-generation benchmarks and three agentic settings across multiple model families and scales. MetaSteer matches or outperforms strong task-specific steering baselines on most aggregate comparisons in the zero-shot regime. Across the evaluated settings, stronger text-generation steering is associated with stronger agentic steering performance. We further discuss geometric trajectory effects, capability retention, and safety considerations raised by transferable steering.
Chinese Translation
对大型语言模型进行引导通常依赖于激活空间中的线性、与上下文无关的干预,而近期工作已对这一假设提出质疑,并且当固定的表示必须编码许多行为区分时,这一假设可能引入信息瓶颈。我们提出 MetaSteer,一种学习具有上下文相关效应的非线性干预并将其施加于注意力投影矩阵的方法,它在构造上产生随输入上下文变化的激活效应,且不需要线性概念几何假设。MetaSteer 被表述为基于偏好的优化,它在池化的偏好语料上训练一次,并以零样本方式迁移到未见过的概念和分布外上下文。我们发现,尽管使用低秩适配器,MetaSteer 仍会在隐藏状态轨迹中引入结构化的、依赖上下文的改变,同时部分保留其局部轨迹动态的若干方面,包括速度和曲率。我们在三个受控文本生成基准和三个智能体设置上,跨多个模型家族与规模对 MetaSteer 进行评估。在零样本情形下的大多数总体比较中,MetaSteer 匹配或优于强的任务特定引导基线。在所评估的各设置中,更强的文本生成引导与更强的智能体引导性能相关联。我们进一步讨论几何轨迹效应、能力保持,以及可迁移引导所带来的安全性考量。
cs.CL / 48 / 2609.38802
Uncovering Uncontrolled Repetition through Residual Stream Dynamics
通过残差流动力学揭示失控重复现象
large language model
大语言模型相关
Abstract
Uncontrolled repetition can prolong autoregressive generation in large language models (LLMs) and enable resource consumption attacks. Prior analyses of repetitive generation have identified strongly activated features in intermediate and late layers. However, how uncontrolled repetition activity emerges and develops before becoming prominent in these layers remains insufficiently understood. In this paper, we investigate this question primarily in large vision-language models (LVLMs), which support a richer set of uncontrolled repetitions through both visual and textual inputs. We propose Tokenwise Residual Comparison (TRC), a method that identifies and localizes anomalies associated with repetition from residual dynamics during generation. TRC compares attention and multilayer perceptron writes to the residual stream across generated tokens to identify patterns associated with repetition. It then selectively suppresses coordinates in the residual stream at the identified layer. Experiments show that TRC effectively mitigates uncontrolled repetition, reducing loop rates by 57\% on average. Our analysis further shows that repetition semantics emerge in shallow layers and propagate through the residual stream, disrupting normal representations. TRC also generalizes to large language models (LLMs) and large reasoning models (LRMs), where it consistently captures analogous repetition dynamics and achieves effective mitigation. Our work broadens the study of repetitive generation from its prominent internal representations to earlier opportunities for intervention, providing insights for mitigating resource consumption attacks.
Chinese Translation
失控重复会延长大语言模型(LLMs)中自回归生成的时长,并可能促成资源消耗攻击。此前对重复生成的分析已在中间层和后期层中识别出被强烈激活的特征。然而,在这些层中变得显著之前,失控重复活动是如何出现并发展的,仍未被充分理解。在本文中,我们主要在大型视觉-语言模型(LVLMs)中研究这一问题,这类模型通过视觉和文本输入支持更丰富的失控重复现象。我们提出了逐词元残差比较(Tokenwise Residual Comparison, TRC),这是一种从生成过程中的残差动力学中识别并定位与重复相关的异常的方法。TRC 比较生成词元过程中注意力与多层感知机对残差流的写入,以识别与重复相关的模式。随后,它在所识别出的层上选择性地抑制残差流中的坐标。实验表明,TRC 有效缓解了失控重复,将循环率平均降低了 57\%。我们的分析进一步表明,重复语义在浅层中出现,并通过残差流传播,从而破坏正常的表示。TRC 也能泛化到大语言模型(LLMs)和大型推理模型(LRMs),在这些模型中它持续捕获类似的重复动力学并实现有效缓解。我们的工作将重复生成的研究从其在内部表示中变得显著之时,拓展到更早的干预机会,为缓解资源消耗攻击提供了洞见。
cs.CL / 49 / 2609.38809
StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning
StateTree:通过强化学习增强长期对话推理
large language model
大语言模型相关
Abstract
Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
Chinese Translation
作为个性化助手部署的大型语言模型必须对漫长且不断演化的交互历史进行推理。然而,在长期对话推理中,相关证据分散于各个会话,偏好可能随时间被修改,而标准的长上下文训练在数据稀缺和计算成本高昂的情况下无法应对这些挑战。我们提出 StateTree,一种数据驱动的 RL 方法,它从具有可验证真值的稀缺对话中构建一项具有挑战性的辅助任务。StateTree 通过树结构的路径追踪任务来增强多会话对话:键值记录被嵌入到多个会话中,以形成一棵二叉树。解决该任务要求模型通过跨会话检索记录并比较时间戳以决定分支,从根节点遍历到叶节点,然后在干扰叶节点中恢复隐藏的目标问题。我们应用课程 RL 训练,逐步增加树的深度,并引入一种组合式变体,其边携带步骤级推理片段,训练模型将部分线索组合成连贯的查询。在 10K token 上下文上训练后,StateTree 可泛化到 128K token,而无需全长 RL 成本,并展现出包括跨会话检索、时间推理、知识更新和组合式多跳推理在内的能力。StateTree 在保持短上下文通用推理能力的同时,优于 SFT 和基于 RL 的基线。StateTree-7B 在 LongMemEval (128k) 上取得高达 +23.60% 的提升,StateTree-14B 在 LongMemEval 上达到 59.00% 的准确率,超过了 QwenLong-L1-32B (45.20%)。
cs.CL / 50 / 2609.38816
You're Hired: Strategic Model Selection for LLM Collaboration
你被录用了:面向LLM协作的战略性模型选择
large language model
大语言模型相关
Abstract
While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systematically evaluate a taxonomy of 9 selection algorithms ranging from diversity of model descriptions, capability-aware behavioral diversity, and LLM-based recruiters. We conduct extensive experiments across two candidate pools of 10 and 32 models, deployed in four model collaboration algorithms, and evaluated across tasks spanning math, coding, QA, and reasoning. Results demonstrate that successful selection algorithms greatly outperform random or heuristics-based teams such as merely selecting the models with top individual performance, by up to 36.1% across settings. Specifically, capability- and training-based selection strategies alleviate selection variance and achieve the best performance, which we recommend to employ before deploying real-world multi-LLM systems. Further analysis reveals that larger candidate pools pose greater challenges to shallow selection heuristics, while algorithms grounded in interacting with candidate models and understanding model capability robustly filter out misaligned, unsafe models, as well as generalizing to novel, out-of-distribution tasks. Together, we establish that principled and informed team selection is critical and present strong model selection algorithms for assembling effective multi-LLM systems.
Chinese Translation
尽管多智能体与模型协作算法在结合多种不同大型语言模型(LLM)优势方面日益受到关注,现有系统仍然受制于预先定义且手工构造的模型池。在本工作中,我们研究多LLM系统中的模型选择问题。我们提出并系统性地评估了一个包含9种选择算法的分类体系,涵盖从模型描述的多样性、能力感知的行为多样性,到基于LLM的招募者。我们在两个分别包含10个和32个模型的候选池上开展了大量实验,将其部署于四种模型协作算法中,并在涵盖数学、编程、问答和推理的任务上进行评估。结果表明,成功的选择算法大幅优于随机或基于启发式的团队,例如仅选择个体性能最高的模型,在各种设置下提升幅度最高达36.1%。具体而言,基于能力和基于训练的选择策略减轻了选择方差并取得最佳性能,我们建议在部署真实世界的多LLM系统之前采用这些策略。进一步分析表明,更大的候选池对浅层选择启发式方法构成更大的挑战,而基于与候选模型交互并理解模型能力的算法能够稳健地过滤掉未对齐、不安全的模型,并泛化到新颖的、分布外任务。综上,我们确立了有原则且信息充分的团队选择至关重要,并提出了用于组建有效多LLM系统的强大模型选择算法。
cs.CL / 51 / 2609.38972
Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
让大语言模型说出其所想:测量与提升 CoT 可解释性对齐
large language model
大语言模型相关
Abstract
Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignment between the reasoning described in an LLM's CoT and what it computes internally. We propose CoT-Interpretability Alignment (CIA), a metric that measures the agreement between a model's CoT traces and its internal reasoning strategies as detected by interpretability tools. We evaluate CIA on three tasks (two-hop question answering, hint intervention, and integer multiplication) across three LLMs, finding that LLMs exhibit limited alignment across all tasks (44.8-75.9%). We then experiment with improving CIA via post-training, setting both the task accuracy and parametric faithfulness signals as a reward. Experiments show that we can substantially improve CoT parametric faithfulness while maintaining or improving the task accuracy. We provide rich analysis, such as their generalization patterns. Our work provides both a framework for auditing CoT parametric faithfulness and a pathway toward making models' explicit reasoning more trustworthy. Code and data are available at https://github.com/yihuaihong/CIA-minimal-repro.
Chinese Translation
思维链(CoT)轨迹常被用作大语言模型(LLM)如何得出其答案的代理。然而,越来越多的证据表明,模型的 CoT 往往无法反映其内部计算,并且可以在不影响其最终答案的情况下被改变。在这项工作中,我们测量并提升 LLM 的 CoT 中所描述的推理与其内部计算之间的对齐程度。我们提出 CoT 可解释性对齐(CIA),这是一个衡量模型 CoT 轨迹与由可解释性工具检测到的其内部推理策略之间一致性的指标。我们在三个任务(两跳问答、提示干预和整数乘法)上、跨三个 LLM 评估 CIA,发现 LLM 在所有任务上都表现出有限的对齐(44.8-75.9%)。随后,我们通过后训练来改进 CIA 进行实验,将任务准确率和参数化忠实性信号同时设为奖励。实验表明,我们可以在保持或提升任务准确率的同时,大幅提升 CoT 参数化忠实性。我们提供了丰富的分析,例如它们的泛化模式。我们的工作既提供了一个用于审计 CoT 参数化忠实性的框架,也提供了一条使模型显式推理更值得信赖的路径。代码和数据可在 https://github.com/yihuaihong/CIA-minimal-repro 获取。
cs.CL / 52 / 2609.39045
RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
RSIGame:具有递归自我改进的自主智能体游戏开发
large language model
大语言模型相关
Abstract
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Chinese Translation
近年来大语言模型的进展使自动游戏生成日益可行,但要将生成的游戏可靠地改进到超越可玩版本仍具挑战性。朴素的迭代改进很容易过拟合一小部分测试用例,产生脆弱游戏,存在未解决的错误、缺失的行为,并且对更广泛的玩家交互泛化能力差。我们提出 RSIGame,一个具有递归自我改进的自主智能体游戏开发框架。RSIGame 将开发过程组织为互补的局部循环和全局循环。具体而言,一个局部的探索—诊断—改进循环广泛探索可执行游戏,诊断并优先处理发现的问题,并执行基于证据的修订,其中不断演化的检查清单持续积累新的测试和改进指导。一个全局循环跟踪总体质量,保留最佳检查点,并检测长周期开发中的饱和或退化。除了测试时改进之外,RSIGame 还通过训练将成功的开发经验内化到生成器中。在 140 个 GameCraft-Bench 任务、两个游戏引擎和五个生成器上,RSIGame 在匹配的开发预算下持续提升游戏质量。值得注意的是,经验内化使 Qwen3.8-27B 在 Godot 上达到 61.38,在 Phaser 上达到 58.53,超过 GPT-5.5 的一次性得分,同时将 Qwen 的生成 token 数减少了 11 倍。
cs.CL / 53 / 2609.39049
Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
结构与思维链:评估用于抑郁严重程度的大语言模型标准提取
large language model
大语言模型相关
Abstract
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.
Chinese Translation
大语言模型(LLM)可以直接根据一条社交媒体帖子评定抑郁严重程度,或者标记出该帖子符合哪些临床标准,并让代码将计数转换为标签。后一种方法更易于审核,因为临床医生可以检查每一个被标记的标准。我们在两个 Reddit 语料库上使用三个 LLM(从 9B 到前沿规模)和两份问卷(PHQ-9、BDI-II)比较这些方法,并用二次加权 kappa 衡量一致性。对于两个前沿模型,仅当其决策阈值在带标签数据上拟合时,标准提取才在其中一个语料库上得分高于思维链。无论是否在同一批标签上重新校准思维链,两个模型的增益均不显著。当阈值根据 PHQ-9 的标准先验固定时,提取在两个语料库上都没有显示出增益,即使模型在每篇帖子上标记超过两个标准时也是如此。9B 模型在一个来自抑郁社区的语料库上表现不同。无论直接提示还是使用思维链,它都将大多数帖子标记为重度,而先验规则在没有标签的情况下优于两者。在思维链于同一批标签上重新校准后,不再存在显著差距,这与校准效应一致。然而,更高的序数一致性并不能确保更好地检测重度病例。PHQ-9 标准提取漏掉了大多数重度帖子,并且在几乎所有比较中,从直接提示转向思维链、再转向提取,都会增加漏检。在主要语料库——一个被重新标注的压力数据集——上,一个使用该数据集自身特征(包括来自文本的词数)的模型,在先验规则下与前沿标准提取没有显著差异。
cs.CL / 54 / 2609.39072
Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue
超越文本:基于LLM的多模态对话维度情感评估
large language model
大语言模型相关
Abstract
Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.
Chinese Translation
对话中的情感识别已被广泛研究,但将大语言模型(LLM)应用于多模态对话中的连续维度情感评估在很大程度上仍未被探索。我们提出了一个基于LLM的框架,该框架在IEMOCAP上执行离散情感识别与效价-唤醒-支配(VAD)维度评估,并遵循SpeechCueLLM方法将声学线索以自然语言描述的形式纳入其中。我们在零样本提示、少样本提示和LoRA微调下评估了涵盖LLaMA、GPT和Qwen系列的六个模型。尽管GPT规模更大,经LoRA微调的LLaMA模型在两个任务上都显著优于通过提示工程构建的GPT模型,我们将这一差距归因于领域适应而非模型容量。我们最好的模型取得了0.7822的效价CCC,在IEMOCAP上达到了新的最先进水平。消融研究证实,文本化的音频描述能显著提升较小的模型(加权F1提升+3.5至3.6),而对最大的模型贡献甚微,这表明当语言能力有限时,音频线索最具价值。VAD各维度之间的性能不对称性与IEMOCAP自身标注中标注者一致性层级高度吻合。
cs.CL / 55 / 2609.39225
Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures
在线对话中的论证结构预测:建模范式与任务架构的比较研究
large language model
大语言模型相关
Abstract
Argument structure prediction (ASP) constructs complete argument structures from discourse by identifying argumentative units and their relations. While recent work has explored diverse approaches---including unified neural models, multi-step pipelines, and prompt-based large language models (LLMs)---their relative trade-offs remain under-explored, particularly in dialogical settings. We present a systematic evaluation of ASP under strict schema constraints, comparing supervised fine-tuning and prompt-based LLMs across single- and multi-step task architectures, generating complete argument structures from dialogical input end-to-end. We benchmark them on three diverse dialogical corpora adapted from Inference Anchoring Theory into bipolar argument structures. Under a shared evaluation framework, we assess predictive performance, cross-domain generalization, schema compliance, and computational efficiency. Our results show that ASP remains a challenging task, with identifying argumentative relations emerging as the primary bottleneck, largely due to the implicit and context-dependent nature of dialogical argumentation. To facilitate future research, we release our data processing pipeline and end-to-end modeling framework for computational ASP on dialogical corpora.
Chinese Translation
论证结构预测(ASP)通过识别论证单元及其关系,从语篇中构建完整的论证结构。尽管近期工作已经探索了多种方法——包括统一的神经网络模型、多步流水线以及基于提示的大语言模型(LLM)——但它们之间的相对权衡仍未被充分探索,尤其是在对话场景中。我们提出了一项在严格模式约束下对 ASP 的系统性评估,在单步与多步任务架构上比较有监督微调与基于提示的 LLM,端到端地从对话输入生成完整的论证结构。我们在三个多样的对话语料库上对它们进行基准测试,这些语料库由 Inference Anchoring Theory 改编为双极论证结构。在共享的评估框架下,我们评估预测性能、跨域泛化能力、模式符合度以及计算效率。我们的结果表明,ASP 仍然是一项具有挑战性的任务,其中识别论证关系成为主要瓶颈,这在很大程度上归因于对话式论证的隐式性和依赖上下文的特性。为促进未来研究,我们发布了用于对话语料库上计算型 ASP 的数据处理流水线与端到端建模框架。
cs.CL / 56 / 2609.39229
RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection
RAIM:用于幻觉检测的低成本模型的稳健聚合
large language model
大语言模型相关
Abstract
Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models, costly and ill-suited to high-throughput monitoring. We investigate whether a panel of cheap open-weight judges (4--9B) can be aggregated to stand in for a frontier one, what the substitution sacrifices, and when it is worth making. We propose RAIM, an aggregation scheme robust to the members' correlated errors, coupling a cross-fitted stacked logistic regression with an admissibility test that, read from the members' own outputs, identifies when aggregating them improves on their best member and stays within reach of the frontier judge. We instantiate RAIM with ten judges from disjoint families across eight faithfulness benchmarks. Against Claude Sonnet, the panel retains a median 93% of its Cohen's $κ$ and gives up only 2.9 points of balanced accuracy on average; read as paired differences, it clearly improves on one benchmark and clearly worsens on three (only two by a non-negligible margin), leaving four unresolved. At a sixty-fourth of the frontier's inference price, the operative expense is a one-time in-domain calibration on 50--100 labelled records. The panel is also competitive with purpose-trained detectors on their home benchmarks (within 1.3 accuracy points of GPT-4o and 1.9 of the LLM-AggreFact leader), and beats the strongest one we reran by 6 points on our grounded sets. Whether aggregation pays depends on the members themselves: where several capable members err on different items, the panel improves on its best judge and approaches the frontier; where one dominates, the stacker recovers the leader, and only there does the frontier remain materially ahead. Both conditions are read off the calibration set at no further cost, so a cheap panel can stand in for a frontier one wherever this audit admits it.
Chinese Translation
忠实性的自动评估越来越依赖于大型语言模型充当评判者,然而最可靠的评判者是专有的前沿模型,成本高昂且不适合高通量监测。我们研究一组廉价开放权重评判者(4--9B)能否被聚合成替代一个前沿评判者,这种替代牺牲了什么,以及何时值得进行这种替代。我们提出 RAIM,一种对成员相关误差稳健的聚合方案,它将交叉拟合的堆叠逻辑回归与一种可采纳性检验耦合起来;该检验从成员自身的输出中读取,识别出聚合它们何时能优于其最佳成员并保持在前沿评判者可及范围内。我们用来自不相交家族的十个评判者在八个忠实性基准上实例化 RAIM。与 Claude Sonnet 相比,该评判者小组保留了其 Cohen's $κ$ 的中位数 93%,并且平均仅损失 2.9 个百分点的平衡准确率;若按配对差异来解读,它在一个基准上明显改善,在三个基准上明显变差(其中仅两个达到不可忽略的幅度),另有四个基准未有定论。在前沿模型推理价格的六十四分之一下,实际开销是一次性的、在 50--100 条标注记录上进行的域内校准。该评判者小组在其原生基准上也与专门训练的检测器具有竞争力(与 GPT-4o 相差 1.3 个准确率点,与 LLM-AggreFact 领先者相差 1.9 个点),并在我们的有依据集合上以 6 个点击败了我们重新运行的最强检测器。聚合是否划算取决于成员本身:当若干有能力的成员在不同条目上出错时,评判者小组会优于其最佳评判者并接近前沿;当某个成员占主导时,堆叠器会恢复出该领先者,且只有在这种情况下前沿模型才仍然实质性领先。这两种条件都可以从校准集中读出,且无需额外成本,因此只要这种审计允许,廉价评判者小组就可以替代前沿评判者小组。
cs.CL / 57 / 2609.39346
Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models
离线引导,在线推理:为小语言模型复用大语言模型反馈
large language model
大语言模型相关
Abstract
Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability-deployment gap has motivated LLM-SLM collaboration, which aims to improve SLM reasoning using LLM capabilities while preserving the deployment advantages of SLMs. Existing approaches mainly follow two paradigms. Knowledge distillation uses LLM-generated answers and reasoning trajectories to train SLMs offline, but requires parameter updates and additional training. Alternatively, online collaboration routes difficult problems to an LLM or leverages LLM-generated guidance and corrections when an SLM encounters difficulties. Although effective, online collaboration requires repeated LLM access. Moreover, the guidance produced for a particular problem is discarded after inference and cannot benefit subsequent problems involving similar reasoning states. In the paper, we focus on a more constrained setting in which the LLM is accessed only offline, the SLM parameters remain fixed, and online inference is performed solely by the SLM. To this end, we propose Reusable Latent Correction (RLC), which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM. RLC stores these experiences in an external bank and retrieves them according to the SLM's current reasoning state, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls. Experiments across multiple reasoning benchmarks and SLM scales show that RLC consistently improves SLM reasoning without parameter updates or online LLM calls. Code is available at https://github.com/ZBH031/reusable-latent-correction.
Chinese Translation
大语言模型(LLM)具备强大的推理能力,但通过商业 API 访问往往成本高昂;而小语言模型(SLM)更易于本地部署,但推理能力仍然较弱。这种能力与部署之间的差距推动了 LLM-SLM 协作,其目标是在利用 LLM 能力提升 SLM 推理的同时,保留 SLM 的部署优势。现有方法主要遵循两种范式。知识蒸馏利用 LLM 生成的答案和推理轨迹离线训练 SLM,但需要参数更新和额外训练。作为替代,在线协作将困难问题路由给 LLM,或在 SLM 遇到困难时利用 LLM 生成的引导和修正。尽管有效,在线协作需要反复访问 LLM。此外,针对特定问题产生的引导在推理之后即被丢弃,无法惠及后续涉及相似推理状态的问题。在本文中,我们关注一个约束更强的设定:LLM 仅在离线阶段被访问,SLM 参数保持固定,在线推理完全由 SLM 执行。为此,我们提出可复用隐式修正(Reusable Latent Correction, RLC),它将来自黑盒 LLM 的一次性自然语言引导转化为 SLM 隐空间中的持久修正经验。RLC 将这些经验存储在一个外部经验库中,并根据 SLM 当前的推理状态检索它们,使 SLM 能够在推理过程中复用源自 LLM 的修正,而无需任何在线 LLM 调用。在多个推理基准和不同规模的 SLM 上的实验表明,RLC 在无需参数更新或在线 LLM 调用的情况下持续提升了 SLM 的推理能力。代码可在 https://github.com/ZBH031/reusable-latent-correction 获取。
cs.CL / 58 / 2609.39420
QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code
QuantCode 模型:面向可执行算法交易代码的语言模型专业化
large language model
大语言模型相关
Abstract
Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.
Chinese Translation
大语言模型是强大的通用代码生成器,但可执行的算法交易仍然是一个要求苛刻的专业化目标:模型必须将自然语言策略规范转换为面向专用交易框架的正确程序逻辑,在历史数据上执行,产生交易,并在语义上忠实于请求。我们研究了两种互补的机制,用于使语言模型针对这一场景进行专业化:在算法交易框架代码上进行继续预训练,以及在经智能体验证的请求到代码对上开展有监督微调(SFT)。评估以 QuantCode-Bench 为核心,这是我们用于 Backtrader 策略生成的 400 任务基准,同时还包括一个仓库级别的、类 SWE-bench 的评测轨道。继续预训练将 Qwen3.5-397B-A17B 的单轮 Judge Pass 从 41.5% 提升至 47.5%,将 Qwen3.6-35B-A3B 的从 27.8% 提升至 33.0%。在继续预训练之后应用 SFT,对 Qwen3.6-35B-A3B 带来了更大的增益,达到 58.2% 的 Judge Pass 和 83.5% 的成功回测;在智能体式评估中,它将首轮成功率从 22.3% 提升至 58.3%,并将最多 10 轮后的最终成功率从 47.5% 提升至 79.5%。仅进行继续预训练可以提升首轮智能体式成功率,但会将修复后的最终成功率从 47.5% 降低至 32.5%,这与指令遵循能力退化相一致;而 SFT 则同时提升两者。我们还识别出一种能力保持失效:领域专业化会损害符合解析器规范的结构化工具调用,而有针对性的恢复性 SFT 能恢复工具调用格式,但无法恢复基础检查点的仓库级智能体性能。结果表明,面向框架的预训练、经过验证的 SFT,以及显式的能力保持评估,分别针对领域特定可执行代码生成中的不同失效模式。
cs.CL / 59 / 2609.39533
CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
CATCH:一个用于编码强化学习中奖励黑客行为的可控分析测试平台
large language model
大语言模型相关
Abstract
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at https://github.com/THUAIS-Lab/CATCH.
Chinese Translation
在具有可验证奖励的强化学习(RLVR)过程中,大型语言模型(LLM)可以利用其环境中的漏洞获取高奖励,而无需提升预期能力,即奖励黑客行为。尽管它对训练效率和安全构成风险,但在训练过程中监测和缓解奖励黑客行为仍然具有挑战性,这受限于缺乏能够复现黑客行为并可靠识别它的测试平台。我们提出了 CATCH,一个用于研究编码 RL 中奖励黑客行为的可控测试平台。CATCH 有意暴露环境漏洞,并通过将易受攻击的评估器下的成功与独立审计下的任务正确性进行比较,提供基于执行的黄金标签。它还可以通过监督微调数据混合来控制模型的初始黑客倾向,并通过奖励设计控制获得奖励的难度,从而能够系统比较黑客动态和干预措施。实验表明,CATCH 能够产生具有明显奖励黑客行为的多样化 RL 训练轨迹,分析表明初始模型和奖励难度都会影响奖励黑客行为的出现。我们进一步评估了不同奖励黑客检测与缓解方法的有效性。一个关键发现是,思维链监控器最初可以抑制黑客行为,但随着策略模型学会用代码注释误导监控器,这种保护会逐渐削弱。这凸显了有必要使用 CATCH 在训练全程评估黑客缓解措施。源代码和资源已在 https://github.com/THUAIS-Lab/CATCH 公开发布。
cs.CL / 60 / 2609.39639
Marginal Response Surface Elicitation for Zero-Label Tabular Learning
面向零标签表格学习的边际响应面诱导
large language model
大语言模型相关
Abstract
Tabular learning uses structured data to predict target outcomes. Traditionally, this process has relied on labeled data. However, large language models (LLMs) can be used to elicit domain priors based on the task description and feature semantics, thereby enabling predictions without labeled data. We propose Marginal Response Surface Elicitation (MARS), a method that transforms feature-level LLM priors into a reusable, zero-shot tabular classifier. To construct this classifier, MARS selects representative values for each feature from unlabeled data and prompts the LLM to provide corresponding class support scores and feature weights. It then aggregates multiple responses using the median to construct feature response functions, and makes predictions through their weighted sum without further LLM queries. Across eight tabular benchmark tasks, MARS achieves the highest average AUC and AP, outperforming direct prompting by 1.97 and 6.21 percentage points respectively, while substantially reducing end-to-end costs. Evaluations with LLMs of different sizes further demonstrate its predictive advantage over direct prompting.
Chinese Translation
表格学习使用结构化数据来预测目标结果。传统上,这一过程依赖于有标签数据。然而,大语言模型(LLM)可以被用来基于任务描述和特征语义引出领域先验,从而在无需有标签数据的情况下实现预测。我们提出边际响应面诱导(MARS),一种将特征级 LLM 先验转化为可复用的零样本表格分类器的方法。为构建该分类器,MARS 从无标签数据中为每个特征选取代表性取值,并提示 LLM 提供相应的类别支持分数和特征权重。随后,它使用中位数聚合多个响应以构建特征响应函数,并通过其加权和进行预测,而无需进一步的 LLM 查询。在八个表格基准任务上,MARS 取得了最高的平均 AUC 和 AP,分别比直接提示高出 1.97 和 6.21 个百分点,同时大幅降低了端到端成本。使用不同规模的 LLM 进行的评估进一步证明了其相较于直接提示的预测优势。
cs.CL / 61 / 2609.39645
SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration
SEPAL:用于可靠 LLM 协作的分离式专家对与答案级融合
large language model
大语言模型相关
Abstract
Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while single-pair Actor-Critic collaboration refines only one candidate. We introduce SEPAL, which assigns three private Actor-Critic teams to direct reasoning, evidence grounding, and verification. Role-specific training gives the teams different reasoning objectives beyond sampling variation. Each Critic guides revisions within its own team, preventing feedback from carrying errors across candidates. Once revision ends, majority voting combines only the final answers, keeping the reasoning histories separate until the decision. Across five open-weight backbones and five question-answering benchmarks, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor-Critic pair, with improvements across all five backbones. Code is available at https://github.com/zhansan114514/SEPAL.
Chinese Translation
多智能体协作使大语言模型(LLM)能够通过审议与反馈来提升问答表现。然而,共享讨论将纠错与暴露于相同错误耦合在一起,这可能侵蚀投票所需的多样性。自一致性提供了无需反馈的采样多样性,而单对 Actor-Critic 协作只能精炼一个候选答案。我们提出 SEPAL,它分配三个私有的 Actor-Critic 团队,分别负责直接推理、证据锚定和验证。角色特定的训练赋予这些团队超越采样差异的不同推理目标。每个 Critic 在其自身团队内指导修订,防止反馈在不同候选之间传递错误。一旦修订结束,多数投票仅合并最终答案,使推理历史在做出决策之前保持分离。在五个开放权重骨干模型和五个问答基准上,SEPAL 相较于匹配的单 Actor-Critic 对将平均准确率提升了 1.81 个百分点,且在全部五个骨干模型上均有一致提升。代码可在 https://github.com/zhansan114514/SEPAL 获取。
cs.CL / 62 / 2609.39661
The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends
大语言模型中注意力机制的演进:机制、权衡与新兴趋势
large language model
大语言模型相关
Abstract
Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model. We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.
Chinese Translation
自注意力使LLM能够以细粒度、依赖查询的方式访问上下文,但密集的 token 交互会带来二次方的预填充成本,以及随上下文长度增长的键--值缓存。因此,研究横跨显式记忆压缩、稀疏访问、循环状态构建、结构化状态动力学以及异构机制组合。本综述将这些发展作为模型内部上下文记忆进行分析。我们引入一个五维视角——记忆表示、记忆更新、访问、读出与整合——描述被表示的内容、其如何变化、哪些内容可被查询、其如何被读取,以及读出如何形成输出。该视角比较相互重叠的研究路线,而不强加单一计算模型。我们使用来自14个主要模型谱系的59条发布级记录和11个高性能开放权重端点,重构了机制层面的发展和架构采用情况。首先,显式记忆方法与循环状态方法保留了不同的接口,但越来越多地控制重叠的记忆功能。其次,异构架构越来越多地跨网络深度进行协调:逐层组合将互补的记忆处理分布到不同表示阶段,而跨层复用则将选定的记忆和路由产物向前传递。因此,深度成为一个构建和管理上下文记忆的维度。第三,这些发展促成了一个有状态的多维记忆路由假设:持久记忆按时间范围、网络深度、基底类型和表示粒度进行组织,而协调的稀疏写入和稀疏读取决定了什么被保持,以及什么对每个查询有所贡献。总体而言,高效序列架构设计越来越关注上下文记忆的组织、生命周期和选择性使用,而不是孤立的注意力算子。
cs.CL / 63 / 2609.39786
Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness
Explore-on-Graph:不完整性下知识图谱问答的混合嵌入-LLM推理
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missing knowledge risk introducing hallucinated evidence. We introduce XoG (eXplore-on-Graph), a framework for multi-hop question answering over incomplete KGs that recovers missing reasoning paths from learned graph structure rather than LLM parametric knowledge. XoG combines type-level entity-relation statistics to identify candidate relations with KG embeddings to retrieve plausible missing entities, using the LLM as a semantic selector and reasoner. These mechanisms are integrated into an iterative planning-exploration-reasoning process. Experiments on WebQSP, CWQ, and the Wikidata-based BRINK benchmark show that XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under KG incompleteness. These gains persist across multiple LLM backbones, indicating that stronger LLMs alone do not resolve missing graph evidence. XoG also reduces LLM token consumption by up to 33% compared with a closely related planning-based approach.
Chinese Translation
大语言模型(LLMs)正日益与知识图谱(KGs)相结合,以将推理建立在结构化证据之上。然而,大多数基于LLM的KGQA方法依赖于遍历现有图边,并且当推理路径因缺失事实而中断时会变得不可靠。要求LLM生成缺失知识的替代方案有引入幻觉证据的风险。我们提出XoG(eXplore-on-Graph),一个面向不完整KG上多跳问答的框架,它从学习到的图结构中而非LLM的参数化知识中恢复缺失的推理路径。XoG将用于识别候选关系的类型级实体-关系统计信息与用于检索合理缺失实体的KG嵌入相结合,使用LLM作为语义选择器和推理器。这些机制被集成到一个迭代的规划-探索-推理过程中。在WebQSP、CWQ以及基于Wikidata的BRINK基准上的实验表明,XoG在完整KG上仍具有竞争力,并且在KG不完整的情况下持续优于那些没有进行任务特定KGQA训练的可比方法。这些增益在多个LLM主干上持续存在,表明仅凭更强的LLM并不能解决图证据缺失的问题。与一个密切相关的基于规划的方法相比,XoG还将LLM token消耗最多降低了33%。
cs.CL / 64 / 2609.39853
Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models
认知增强:重新思考角色扮演对大型语言模型的必要性
large language model
大语言模型相关
Abstract
Role-playing prompting has become a popular yet simple technique for improving LLM reasoning and output quality. However, whether it consistently boosts performance across diverse domains remains unclear, as systematic validation is lacking. To fill this gap, we run multi-model, cross-domain, and multilingual experiments on MMLU and MMLU-Redux. We find that gains from role-play prompting depend heavily on model capacity, knowledge domain, and prompt language. Drawing on metacognition theory, we propose the persona-related cognitive alignment hypothesis: role-play works only when the LLM correctly grasps the designated persona and its associated knowledge domain. We test this hypothesis through persona information richness ablation, layer-wise entropy divergence analysis, and latent thought-space deflection observation. To reduce persona cognitive bias and stabilize role-play performance, we propose \textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP}), a simple, training-free, and efficient multilingual prompt concatenation strategy. It aggregates semantically equivalent role prompts to enrich complementary representational cues. Extensive experiments show that MLCP consistently outperforms vanilla role-play prompting across all tested LLMs.
Chinese Translation
角色扮演提示已成为一种流行而简单的技术,用于提升 LLM 的推理和输出质量。然而,由于缺乏系统性验证,它是否能在不同领域中持续提升性能仍不清楚。为填补这一空白,我们在 MMLU 和 MMLU-Redux 上进行了多模型、跨领域和多语言实验。我们发现,角色扮演提示带来的收益在很大程度上取决于模型能力、知识领域和提示语言。借鉴元认知理论,我们提出角色设定相关的认知对齐假设:只有当 LLM 正确把握指定角色设定及其相关知识领域时,角色扮演才会起作用。我们通过角色设定信息丰富度消融、逐层熵散度分析和潜在思维空间偏转观察来检验这一假设。为减少角色设定认知偏差并稳定角色扮演性能,我们提出 \textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP}),一种简单、无需训练且高效的多语言提示拼接策略。它聚合语义等价的角色提示,以丰富互补的表征线索。大量实验表明,在所有测试的 LLM 上,MLCP 始终优于普通角色扮演提示。
cs.CL / 65 / 2609.39882
LLM Persona Unlearning
LLM 人格遗忘
large language model
大语言模型相关
Abstract
Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefore elicit personas that repeatedly shape judgment, language, and action. In open-weight settings, runtime controls can be removed, motivating persona unlearning: a weight-level edit that makes a designated persona difficult to elicit and enact on unseen contexts. We introduce PersonaUnlearnBench, a model-specific paired benchmark spanning six LLMs from three families and five personas, with aligned forget/retain sets, held-out instruction paraphrases, and four-axis evaluation. The benchmark shows that standard unlearning methods cannot reliably erase the target persona without sacrificing meaningful generation or general utility. We therefore propose PaCE, which compares target and desirable responses to the same questions to locate an internal behavior direction, then trains target-prompt states away from the target mode and toward the matched desirable response. Experiments show that PaCE consistently suppresses target personas with high response quality and useful counterpart behavior, at moderate utility cost. These results establish persona unlearning as a distinct behavior-level editing problem and a practical route toward persistent control of latent LLM response policies.
Chinese Translation
预训练使大语言模型(LLMs)获得一整套广泛的行为模式库,这些模式与角色、风格、价值观和目标相关联。后训练教会条件性执行,并使一个有帮助的助手(Assistant)成为默认,但它并不会从权重中抹去替代模式;因此,显式提示可以引出人格,而这些人格会反复塑造判断、语言和行动。在开放权重设置中,运行时控制可以被移除,这促使人格遗忘:一种权重层面的编辑,使指定人格难以在未见情境中被引出和执行。我们提出了 PersonaUnlearnBench,一个针对特定模型的配对基准,涵盖来自三个模型家族的六个 LLM 和五个人格,具有对齐的遗忘/保留集、留出的指令复述以及四轴评估。该基准表明,标准遗忘方法无法在不牺牲有意义的生成或通用效用的情况下可靠地抹除目标人格。因此,我们提出了 PaCE,它比较对相同问题的目标回应和期望回应,以定位一个内部行为方向,然后训练目标提示状态远离目标模式并朝向匹配的期望回应。实验表明,PaCE 以中等的效用代价,持续抑制目标人格,同时保持高回应质量和有用的对应行为。这些结果将人格遗忘确立为一个独特的行为层面编辑问题,以及一条通向对潜在 LLM 响应策略进行持久控制的实用路径。
cs.CL / 66 / 2609.39884
OPSRD: On-Policy Self-Role Distillation
OPSRD:同策略自角色蒸馏
large language model
大语言模型相关
Abstract
Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.
Chinese Translation
角色提示通过专家身份从大语言模型中激发出专门化行为,提供了一种轻量级方式来引导对高要求任务的推理。然而,当采样得到的解答仍然不正确时,评估或蒸馏完整的角色提示答案可能会遗漏有用的下一 token 偏好。迁移这些偏好还需要一个能够触达学生很少预测的备选项的目标函数。我们提出 OPSRD,它使用固定的专家角色作为特权教学上下文,在无需参考解答的情况下进行同策略自蒸馏。一个无角色的学生模型生成一条轨迹,而同一基础模型的一个冻结实例在其精确前缀上提供角色条件分布,从而暴露出超出采样续写的备选项。教师加权的前向 KL 以学生低估的备选项为目标,并通过裁剪来限制单个词表项的贡献。监督被限制在学生位置中熵最高的一半,从而将学习集中在预测不确定的位置。在 Qwen3-1.7B、4B 和 8B 上,针对三个竞赛数学基准的实验表明,其相较于推理时没有角色提示的基础模型有所提升。在每个规模下,前向 KL 在三种所评估的散度中取得了最高的宏平均准确率。代码可在 https://github.com/zhansan114514/OPSRD 获取。
cs.CL / 67 / 2609.39913
The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation
具体-任意差距:大型语言模型中的亲属关系推理并非对呈现方式无动于衷
large language model
大语言模型相关
Abstract
We test whether large language models solve formally matched kinship problems equally well when relations are expressed in familiar vocabulary or by explicitly defined nonce predicates. Across 500 paired graphs, concrete accuracy exceeds arbitrary accuracy by 35.6 percentage points in local Qwen3.8-27B, 26.6 in Gemma 4 26B-A4B, 12.0 in Gemma 4 31B, and 5.4 in Qwen3.8-Max. All four paired gaps are statistically resolved. Reasoning budgets and prompt-language interventions can substantially reduce the difference, showing that it is modifiable rather than a fixed incapacity. The minimal conclusion is behavioral: on these tasks, the models' manifested relational competence is not indifferent to presentation. Explicit definitions provide the formal relations but do not make nonce predicates as usable as familiar vocabulary embedded in learned linguistic associations.
Chinese Translation
我们测试当关系用熟悉的词汇表达或用明确定义的临时谓词表达时,大型语言模型是否能同样好地解决形式匹配的亲属关系问题。在 500 个配对图中,具体准确率在本地 Qwen3.8-27B 中超过任意准确率 35.6 个百分点,在 Gemma 4 26B-A4B 中超过 26.6 个百分点,在 Gemma 4 31B 中超过 12.0 个百分点,在 Qwen3.8-Max 中超过 5.4 个百分点。所有四个配对差距在统计上均得到解决。推理预算和提示语言干预可以大幅缩小这一差异,表明它是可修改的,而不是一种固定的能力缺失。最小结论是行为层面的:在这些任务上,模型所表现出的关系能力并非对呈现方式无动于衷。明确定义提供了形式关系,但并未使临时谓词像嵌入在学习到的语言关联中的熟悉词汇那样可用。
cs.CL / 68 / 2609.40103
JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation
JuryFlow:分歧引导的人在回路多智能体评估
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into reusable rubric entries that all judges inherit, making the evaluator progressively self-refining. To enable large-scale, reproducible benchmarking without human studies, we evaluate JuryFlow in an automatic configuration in which focal selection is made by entropy ranking. On MT-Bench and LLMBar, JuryFlow improves agreement with gold labels over single-judge and majority-vote panel baselines, and ablations isolate the contributions of disagreement-targeted re-evaluation, propagation, and rubric induction. We contribute (1) a human-in-the-loop paradigm that recasts the human from labeler to structural guide, (2) the JuryFlow framework operationalizing it through a disagreement graph, focal re-evaluation, and closed-loop rubric induction, and (3) an evaluation protocol with ablations that isolate where the gains originate.
Chinese Translation
大型语言模型(LLM)正日益被部署为对 AI 生成内容进行评判的自动评审员,然而单个评审员并不可靠,甚至一个评审员小组也会留下难以消解的残余:当评审员之间出现分歧时,多数投票会丢弃这一冲突,而不是解决它。我们提出 JuryFlow,一个由分歧引导、人在回路的多智能体评估框架,它不把评审员之间的分歧视为需要被平均掉的噪声,而是将其视为一个精确的、主张级(claim-level)的信号,用以指示评估在何处存在不确定性。JuryFlow 将每个候选回复分解为原子主张,让一组异构评审员为每个主张分别给出判定,并构建一个分歧图,其节点由判定熵进行评分,其边编码主张之间的结构相似性。人类充当结构引导者,通过一次单一的最小干预来选择要解决哪一个分歧,而不是重新标注该回复;此后,焦点主张被重新评估,该修正沿图的边传播,并传播到历史上相似的案例,进而被固化为所有评审员都继承的可复用评分标准条目,使评估器逐步自我精炼。为了在无需人类研究的情况下实现大规模、可复现的基准测试,我们在一种自动配置下评估 JuryFlow,其中焦点选择由熵排序完成。在 MT-Bench 和 LLMBar 上,JuryFlow 相较于单评审员基线和多数投票评审组基线,提升了与金标准标签的一致性,并且消融实验分离出了针对分歧的重新评估、传播和评分标准归纳各自的贡献。我们的贡献包括:(1) 一种人在回路的范式,将人类从标注者重新定位为结构引导者;(2) JuryFlow 框架,通过分歧图、焦点重新评估和闭环评分标准归纳将其付诸实现;(3) 一套带有消融实验的评估协议,用以分离增益的来源。
cs.CL / 69 / 2609.40121
On the (In)effectiveness of AMR Augmentation for Large Language Models
论 AMR 增强对大型语言模型的(无)有效性
large language model
大语言模型相关
Abstract
While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to reproduce recent work that reported substantial downstream gains from AMR augmentation, finding that these are likely due to specific choices in the experimental settings used: using a consistent and unified protocol for hyperparameter selection, we observe that text-only baselines consistently match or exceed the performance of AMR-augmented models. To investigate this null result, we introduce a perplexity-based probe measuring the degree to which AMR provides an LLM with supplemental relational knowledge not already available to the model. We find that AMR augmentation does not help LLMs improve their understanding of relational content in the sentence, indicating that augmenting these models with AMR offers no clear benefit on downstream tasks.
Chinese Translation
尽管抽象语义表示(Abstract Meaning Representation,AMR)在历史上提升了一系列 NLP 任务的性能,但 AMR 增强对现代 LLM 的益处——或缺乏益处——迄今仍不清楚。在本文中,我们尝试复现近期一项报告 AMR 增强带来显著下游收益的工作,发现这些收益可能归因于所用实验设置中的特定选择:采用一致且统一的超参数选择协议,我们观察到纯文本基线始终匹配或超过 AMR 增强模型的性能。为了探究这一零结果,我们引入一种基于困惑度的探针,用于衡量 AMR 在多大程度上为 LLM 提供了模型尚未拥有的补充关系知识。我们发现,AMR 增强并不能帮助 LLM 提升其对句子中关系内容的理解,这表明用 AMR 增强这些模型对下游任务没有明显益处。
cs.CL / 70 / 2609.40124
Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions
自己动手去偏:向大型语言模型传授认知偏差缓解干预
large language model
大语言模型相关
Abstract
Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.
Chinese Translation
偏差长期以来一直在社会心理学和认知科学中受到研究,数十年的研究已经产生了一系列经过验证的干预措施,这些措施可以减少人类的刻板印象思维和偏见性反应。我们提出“自己动手去偏”(Debias It Yourself,DIY),一个以认知为基础的框架,它将五种此类干预转化为面向大型语言模型的去偏流程,并通过三种成熟范式加以实施:Show(上下文示例)、Train(指令微调)和 Revise(引导式自我修订)。在三个模型、五个偏差基准、十一个去偏基线和三个推理基准上,Train+Revise 和单独使用 Revise 取得了最高的两个平均排名,在偏差-推理权衡中领先(在 90% 推理准确率下,平均偏差低至 2%),并将未见维度上的偏差最多降低 14.8%。我们的代码和数据已公开可用。
cs.CL / 71 / 2609.40340
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
EvoDuet:面向科学发现的网络搜索与任务求解的双层协同演化
large language model
大语言模型相关
Abstract
Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.
Chinese Translation
当进展需要模型所缺乏的外部知识时,使用大语言模型(LLM)的演化搜索可能会陷入停滞。提供相关文档会有所帮助,但仅仅添加一个网络搜索工具,可能会随着解的不断变化而反复返回相同的页面。我们提出 EvoDuet,一种在模型参数固定不变的情况下协同演化解与搜索查询的双层优化方法。在每一次迭代中,一个检索门控让 LLM 评估自身的知识缺口,并选择检索新文档、复用已存储的文档,或在不使用文档的情况下继续。内层循环对查询进行细化,并根据文档预计能够带来的解得分对其排序;外层循环从这些文档中并行生成候选解,并记录评估结果以供后续搜索使用。在 21 项每次迭代仅生成一个候选解的优化任务上,EvoDuet 将 OpenEvolve 的归一化发现增益在使用 GPT-5.6-Luna 时从 74.1% 提升至 78.0%,在使用 Gemini-3.8-Flash 时从 61.3% 提升至 82.3%,而 Qwen3.5-9B 则未从中获益。我们表现最好的运行结果在八项任务上超过了此前报告的最佳得分,包括 Q20 和 Rosetta 上的 Swap Reduction,并在另外三项任务上与这些最佳得分持平。EvoDuet 在 Sums/Diffs 和 Denoising 任务上与其他脚手架(例如 Top-K、EvoX)结合时同样能带来提升,表明其可适用于各类演化搜索脚手架。
cs.CR / 72 / 2609.38381
LogiC-Diff: Embedding Security Properties Into AI-Enabled Cyber-Physical Systems
LogiC-Diff:将安全属性嵌入AI赋能的网络物理系统
diffusion
扩散模型相关
Abstract
AI-enabled Cyber-Physical Systems (CPS) are highly vulnerable to adversarial and anomalous inputs, where small perturbations can induce cascading errors and unsafe control actions. Existing approaches, such as rule-based filtering, training-time regularization, or diffusion-based reconstruction, either operate outside the model or lack mechanisms to incorporate formal security specifications into the prediction process. In this paper, we take the first step toward embedding security properties directly into AI-enabled CPS, enabling predictive models to enforce system-level constraints during inference rather than relying on external defenses. We introduce a logic-conditioned bi-stage diffusion framework that integrates Signal Temporal Logic (STL) specifications into forecasting. STL serves as a first-class conditioning signal that guides both an input repair stage and an output refinement stage, allowing the model to jointly mitigate adversarial perturbations and enforce desired temporal behaviors to satisfy security-critical properties. We evaluate our approach on two real-world multivariate CPS forecasting datasets under a diverse set of physical sensor and cyber attacks. Across sensor faults, gradient-based attacks, adaptive attacks, and varying attack strengths, our method consistently improves robustness and specification compliance, degrades more gracefully as attack strength increases, and generalizes better to unseen attacks. Ablation studies on specification coverage and quality further show that embedding logical security properties yields gains unattainable by reconstruction-based methods alone, highlighting a new direction for integrating formal methods with generative models in secure CPS.
Chinese Translation
AI赋能的网络物理系统(CPS)极易受到对抗性和异常输入的影响,其中微小的扰动即可引发级联错误和不安全的控制动作。现有方法,如基于规则的过滤、训练时正则化或基于扩散的重建,要么在模型外部运行,要么缺乏将形式化安全规约纳入预测过程的机制。在本文中,我们迈出了将安全属性直接嵌入AI赋能CPS的第一步,使预测模型能够在推理过程中强制执行系统级约束,而不是依赖外部防御。我们提出了一种逻辑条件化的双阶段扩散框架,该框架将信号时序逻辑(STL)规约集成到预测之中。STL作为一种一等条件信号,同时引导输入修复阶段和输出精炼阶段,使模型能够联合缓解对抗性扰动并强制执行期望的时序行为,以满足安全关键属性。我们在两个真实世界的多变量CPS预测数据集上,在一组多样化的物理传感器攻击和网络攻击下对我们的方法进行了评估。在传感器故障、基于梯度的攻击、自适应攻击以及不同攻击强度下,我们的方法始终提升鲁棒性和规约符合度,随着攻击强度增加其性能退化更为平缓,并且对未见过的攻击具有更好的泛化能力。关于规约覆盖率和质量的消融研究进一步表明,嵌入逻辑安全属性所带来的增益是仅靠基于重建的方法无法获得的,这凸显了在安全CPS中将形式化方法与生成模型相结合的新方向。
cs.CR / 73 / 2609.38389
The Geometry of Harmfulness in Multi-Turn Attacks
多轮攻击中危害性的几何结构
large language model
大语言模型相关
Abstract
Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work investigates how the geometry and temporal dynamics of harmfulness and refusal representations evolve across multi-turn attacks. We analyzed hidden-state representations from three instruction-tuned LLMs (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-9B-it) using three multi-turn attack frameworks (Crescendo, ActorAttack, and X-Teaming), and examined representation behavior across conversation turns, model layers, and token positions under various context configurations. Across models and frameworks, we found that (1) each attack framework traverses different geometric directions, yet each achieves comparable success in eliciting harmful outputs; (2) multi-turn harmfulness directions became increasingly linearly separable at the end-of-turn token position across turns in middle to late model layers; and (3) harmfulness representations are weakly aligned with refusal-related representations. The results indicate that multi-turn attacks do not succeed by suppressing the model's internal representation of harmfulness. Instead, harmfulness representations become increasingly separable across conversation turns, while remaining only weakly aligned with refusal-related representations. The findings are one possible explanation for why static single-turn safety probes may degrade in multi-turn settings, and suggest that robust defenses must consider temporal representation dynamics rather than identifying harmfulness with isolated or single-turn prompts.
Chinese Translation
大型语言模型(LLM)仍然容易受到对抗性攻击的影响,这类攻击能够绕开安全对齐,从而诱发有害输出。目前尚不清楚,在多轮攻击的过程中,危害性表征与拒绝表征是如何演变的,也不清楚为什么单轮防御在多轮场景下效果较差。本工作研究了在多轮攻击中,危害性与拒绝表征的几何结构和时间动态如何演变。我们使用三种多轮攻击框架(Crescendo、ActorAttack 和 X-Teaming),分析了三个指令微调 LLM(Llama-3.1-8B-Instruct、Qwen2.5-7B-Instruct 和 Gemma-2-9B-it)的隐藏状态表征,并考察了在各种上下文配置下,表征在对话轮次、模型层和 token 位置上的行为。跨模型和框架,我们发现:(1)每种攻击框架都沿着不同的几何方向行进,但每种框架在诱发有害输出方面都取得了相当的成功;(2)在模型的中层到后层中,多轮危害性方向在轮末 token 位置上随着对话轮次的推进变得越来越线性可分;以及(3)危害性表征与拒绝相关表征之间的对齐程度较弱。这些结果表明,多轮攻击并非通过抑制模型内部的危害性表征来取得成功。相反,危害性表征随着对话轮次的推进变得越来越可分,同时与拒绝相关表征之间仅保持较弱的对齐。这些发现为静态单轮安全探针在多轮场景下可能退化提供了一种可能的解释,并表明稳健的防御必须考虑表征的时间动态,而不是用孤立的或单轮的提示来识别危害性。
cs.CR / 74 / 2609.38477
Security-Enhanced Seed-Based Weight Quantization for Large Language Models
面向大语言模型的安全增强型基于种子的权重量化
large language model
大语言模型相关
Abstract
Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity of model weights. We introduce Seed-Q, a security-enhanced sensitivity-aware seed-based weight compression framework that uses lightweight Linear Feedback Shift Register (LFSR)-based weight generation with non-uniform bit allocation. Our approach assigns larger representation budgets to sensitive weights while aggressively compressing less sensitive regions. Importantly, this non-uniform allocation requires no side-information: the decoder deterministically reconstructs the bit-allocation schedule, with no rung depending on the decoded weights, eliminating the need to store per-block metadata or use calibration data while preserving the baseline coding rate. Experiments across diverse LLMs show that Seed-Q matches 4-bit perplexity of SeedLM with fewer bits, while at the same 4 bits/weight it reduces both perplexity degradation and zero-shot accuracy loss relative to SeedLM. We also show that Seed-Q simultaneously achieves high security against bit-flip attacks on model parameters, as bit corruption affects multiple reconstructed weights, greatly amplifying its impact and making it easier to detect. We further implement Seed-Q in an ASIC-based accelerator and demonstrate modest hardware overhead compared to prior seed-based approaches.
Chinese Translation
大语言模型(LLM)会带来大量的存储、内存带宽和能耗开销,这促使人们采用紧凑的权重表示。现有的基于种子的压缩方法从紧凑的伪随机表示中重建权重,但没有显式考虑模型权重的非均匀敏感性。我们提出 Seed-Q,一个安全增强的、感知敏感性的基于种子的权重压缩框架,它使用轻量级的基于线性反馈移位寄存器(LFSR)的权重生成以及非均匀比特分配。我们的方法为敏感权重分配更大的表示预算,同时激进地压缩敏感度较低的区域。重要的是,这种非均匀分配不需要任何边信息:解码器确定性地重建比特分配调度,没有任何梯级依赖于已解码的权重,从而无需存储逐块元数据或使用校准数据,同时保持基线编码率。在多种 LLM 上的实验表明,Seed-Q 以更少的比特即可达到与 SeedLM 相当的 4 比特困惑度,而在相同的 4 比特/权重下,相对于 SeedLM,它同时降低了困惑度退化与零样本准确率损失。我们还表明,Seed-Q 同时对模型参数上的比特翻转攻击实现了高安全性,因为比特损坏会影响多个重建权重,极大地放大其影响,使其更容易被检测到。我们进一步在基于 ASIC 的加速器中实现了 Seed-Q,并证明与先前的基于种子的方法相比,其硬件开销适中。
cs.CR / 75 / 2609.38899
SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs
SceneJail:利用视频场景上下文越狱多模态大语言模型
large language model
大语言模型相关
Abstract
Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios. To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding scenario that is contextually compatible with the harmful query. Scenario-aware Prompt Search uses black-box response feedback to search for textual guidance tailored to the selected scenario. Extensive evaluations on the HADES and SafeBench datasets across eight Video-MLLMs, including two proprietary models, GPT-4.1 and Gemini3.5-Flash, demonstrate the effectiveness of SceneJail. SceneJail-F, which presents the complete query persistently, achieves average attack success rates (ASR) up to 91.5%, outperforming the strongest baselines by 29.1 percentage points. Furthermore, SceneJail-S, which distributes the query across successive frames, remains highly robust against current defenses, retaining a 72.3% ASR even under strict image filtering.
Chinese Translation
视频多模态大语言模型(Video-MLLMs)支持对视频输入进行推理,但仍然容易受到越狱攻击,这些攻击会引发违反策略的响应。现有视频越狱主要操纵有害查询的视觉呈现方式,从而将视频仅视为载体。因此,周围的视频场景作为上下文攻击面仍未得到探索。在本文中,我们表明,同一有害查询置于不同视频场景中时会引发不同的安全响应。为了系统性地利用这一漏洞,我们提出了 SceneJail,一个具有两个协调组件的自适应黑盒越狱框架。自适应场景构建动态搜索与有害查询在上下文上兼容的周围场景。场景感知提示搜索使用黑盒响应反馈来搜索针对所选场景定制的文本引导。在 HADES 和 SafeBench 数据集上对八个 Video-MLLM 进行的广泛评估,包括两个专有模型 GPT-4.1 和 Gemini3.5-Flash,证明了 SceneJail 的有效性。SceneJail-F 持续呈现完整查询,达到平均攻击成功率(ASR)高达 91.5%,比最强基线高出 29.1 个百分点。此外,SceneJail-S 将查询分布到连续帧中,对当前防御保持高度鲁棒,即使在严格的图像过滤下仍保持 72.3% 的 ASR。
cs.CR / 76 / 2609.38954
APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry
APTInvestBench:评估在不同遥测条件下的自主 APT 调查
large language model
大语言模型相关
Abstract
Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry setting does not establish robustness to changes in log collection, retention, or sampling. We introduce APTInvestBench, a benchmark for evaluating cross-telemetry robustness in autonomous APT investigation. It comprises 370 cases across seven SOC-inspired conditions, derived from 56 report-informed attack reconstructions with 16.4 million log records. Agents investigate unverified leads and submit reports with record-level citations. Fixed action-level support requirements track sufficient evidence across available logs, query returns, and formal citations, separating telemetry limitations from acquisition and reporting gaps. Across eleven LLMs, agents acquire sufficient evidence for 44.3% of recoverable attack actions on average, while formal citations support only 25.0%. More importantly, aggregate coverage can conceal substantial instability: from Full to endpoint-only telemetry, coverage declines by only 1.6 percentage points, yet 35.5% of previously covered actions lose sufficient citation support despite remaining recoverable. Across four frameworks, such losses persist even when registered supporting records remain unchanged. APTInvestBench provides reusable investigation environments and diagnostic evaluation for identifying these gaps and developing more reliable defensive agents.
Chinese Translation
大语言模型(LLM)智能体可以通过将微弱线索转化为用于入侵范围界定和响应的证据,帮助安全运营中心(SOC)调查高级持续性威胁(APT)。然而,在一种遥测设置下取得成功,并不能证明其对日志收集、保留或采样变化的鲁棒性。我们提出 APTInvestBench,一个用于评估自主 APT 调查中跨遥测鲁棒性的基准。它包含 370 个案例,涵盖七种受 SOC 启发的条件,这些案例源自 56 个基于报告的攻击重构,包含 1640 万条日志记录。智能体调查未经核实的线索,并提交带有记录级引用的报告。固定的动作级支持要求跟踪可用日志、查询返回结果和正式引用中是否有充分证据,从而将遥测限制与采集和报告差距区分开来。在十一个 LLM 中,智能体平均为 44.3% 的可恢复攻击动作获取了充分证据,而正式引用仅支持 25.0%。更重要的是,总体覆盖率可能掩盖显著的不稳定性:从完整遥测到仅端点遥测,覆盖率仅下降 1.6 个百分点,但先前被覆盖的动作中有 35.5% 失去了充分的引用支持,尽管它们仍然可恢复。在四个框架中,即使已登记的支持记录保持不变,此类损失依然存在。APTInvestBench 提供可复用的调查环境和诊断性评估,用于识别这些差距并开发更可靠的防御智能体。
cs.CR / 77 / 2609.39279
Faithful Dual-constrained Erasure for Robust LLM Safety Alignment
用于稳健 LLM 安全对齐的忠实双约束擦除
large language model
大语言模型相关
Abstract
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly resurface after benign fine-tuning. In this work, we investigate the optimization dynamics of unlearning and identify that this vulnerability stems from shallow alignment. Rather than effectively erasing target knowledge, models often exploit a shortcut by activating previously dormant parameters to act as spurious suppressors, forming a fragile inhibitory shell over intact malicious representations. To address this issue and enforce authentic memory deletion, we propose FDCU, a novel dual-constrained subspace projection framework. FDCU restricts parameter updates through a highly scalable, element-wise dual-masking rule: it preserves general knowledge manifolds via Fisher Information and strictly prohibits the abnormal activation of spurious suppressors via the Principle of Minimal Functional Intervention (PMFI). By reliably blocking the model's ability to superficially hide knowledge, FDCU promotes the authentic dismantling of target representations. Extensive experiments across specific knowledge erasure and safe output control tasks demonstrate that FDCU achieves state-of-the-art robustness against retraining attacks while maintaining near-lossless general utility, ensuring durable safety for LLMs.
Chinese Translation
机器遗忘已成为一种关键机制,用于移除危险知识并在大语言模型(LLM)中强制执行安全对齐。然而,近期研究揭示了一个持续存在的安全风险:已遗忘模型仍然极易受到重训练攻击,其中被抑制的恶意行为在良性微调后迅速重新出现。在这项工作中,我们研究了遗忘的优化动力学,并发现这种脆弱性源于浅层对齐。模型往往不是有效地擦除目标知识,而是利用一种捷径:激活先前休眠的参数以充当虚假抑制器,从而在完整的恶意表征之上形成一个脆弱的抑制外壳。为解决这一问题并强制实现真实的记忆删除,我们提出了 FDCU,一个新颖的双约束子空间投影框架。FDCU 通过一种高度可扩展的逐元素双重掩码规则来约束参数更新:它通过 Fisher 信息保留通用知识流形,并通过最小功能干预原则(PMFI)严格禁止虚假抑制器的异常激活。通过可靠地阻断模型表面上隐藏知识的能力,FDCU 促进对目标表征的真正拆解。在特定知识擦除和安全输出控制任务上的大量实验表明,FDCU 在抵御重训练攻击方面实现了最先进的稳健性,同时保持了近乎无损的通用效用,确保 LLM 具有持久的安全性。
cs.CR / 78 / 2609.39549
Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
推测式安全蜜罐:面向多轮智能体攻击的主动防御
large language model
大语言模型相关
Abstract
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification stage, the system uses the target agent's real actions to calibrate and prune the trajectory tree, effectively reducing false positives. As a plug-and-playable component, SSH provides existing detectors with rich decision redundancy beyond the current interaction slice. By judging risk based on the evolution of the entire trajectory tree rather than a single point in time, the system reduces the reliance on the absolute precision of individual detection components. This improves the defense resilience and the warning lead-time of agent systems against complex temporal attacks.
Chinese Translation
随着大型语言模型(LLM)智能体日益部署于复杂环境中,多轮交互攻击已成为一项重大的安全挑战。现有检测方法通常依赖历史上下文。然而,这种回溯式逻辑难以识别那些为隐藏未来风险而被拆分到不同轮次中的深层恶意意图。受推测解码启发,我们提出了推测式安全蜜罐(SSH)框架。SSH 使用由小型 LLM 组成的多智能体模拟系统来构建动作级的“推测-验证”工作流。在推测阶段,SSH 预测目标智能体的未来行为,并异步构建轨迹树,以提前暴露潜在风险。在验证阶段,系统使用目标智能体的真实动作来校准和剪枝轨迹树,从而有效减少误报。作为一个即插即用组件,SSH 为现有检测器提供了超越当前交互切片的丰富决策冗余。通过基于整棵轨迹树的演化而非单一时间点来判断风险,系统降低了对单个检测组件绝对精确性的依赖。这提高了智能体系统面对复杂时序攻击的防御韧性和预警提前量。
cs.CR / 79 / 2609.39880
PassGPT+: Leveraging Linguistic Priors for Password Modeling
PassGPT+:利用语言先验进行密码建模
diffusion
扩散模型相关
Abstract
Passwords remain the dominant online authentication mechanism, and understanding how humans choose them is essential for defensive strength estimation and attack simulation alike. Recent learning-based approaches such as PassGAN and PassGPT have shown that deep generative models can learn password structure directly from leaked corpora. However, both train from random initialization on password data alone. The role of linguistic prior knowledge in password modeling, and what it reveals about how humans create secrets, remains largely underexplored. Here, we address this gap with PassGPT+, which adapts the linguistic prior of GPT-2 to password observations through character-aware tokenization. We also introduce PassDiffusion, the first absorbing-state discrete diffusion model for password generation, as a probe of whether non-autoregressive approaches are competitive. On the RockYou benchmark, PassGPT+ recovers 22.53% of held-out passwords at 108 guesses, a 16% relative gain over PassGPT, and retains 79% of this match rate when transferred without retraining to a disjoint 2020 leak dataset, demonstrating that linguistic priors capture persistent regularities of human password generation. PassDiffusion underperforms by two to three orders of magnitude, indicating that autoregressive modeling is substantially better matched than iterative denoising to the discrete, exact-match nature of password generation.
Chinese Translation
密码仍然是主流的在线身份验证机制,理解人类如何选择密码对于防御性强度估计和攻击模拟而言都至关重要。近期的基于学习的方法,如 PassGAN 和 PassGPT,已经表明深度生成模型能够直接从泄露的语料库中学习密码结构。然而,二者都仅在密码数据上从随机初始化开始训练。语言先验知识在密码建模中的作用,以及它揭示了人类如何创建秘密的哪些方面,仍然在很大程度上未被充分探索。在这里,我们通过 PassGPT+ 来填补这一空白,它通过字符感知的分词将 GPT-2 的语言先验适配到密码观测上。我们还引入了 PassDiffusion,这是首个用于密码生成的吸收态离散扩散模型,用以探究非自回归方法是否具有竞争力。在 RockYou 基准上,PassGPT+ 在 108 次猜测中恢复出 22.53% 的留出密码,相较 PassGPT 获得 16% 的相对提升,并且在未经重新训练的情况下迁移到一个不相交的 2020 年泄露数据集时仍保持该匹配率的 79%,这表明语言先验捕捉到了人类密码生成中持续存在的规律。PassDiffusion 的性能低了两到三个数量级,这表明自回归建模比迭代去噪显著更契合密码生成的离散、精确匹配特性。
cs.CR / 80 / 2609.39902
CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
CodeMimicry:通过结构化代码补全利用大型语言模型中的安全泛化滞后
large language model
大语言模型相关
Abstract
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Chinese Translation
大型语言模型在多个不同领域已经取得了显著的能力,但它们的安全对齐仍然容易受到越狱攻击。在这项工作中,我们识别出一种此前未被充分探索的失效模式——安全泛化滞后——其中主要在自然语言上训练的对齐无法迁移到代码领域。我们表明,这种滞后会诱发一种代码补全盲点,使得嵌入在语法有效代码中的恶意意图能够规避安全机制。为了利用这一漏洞,我们提出了 CodeMimicry,一个全自动的黑盒越狱框架,它生成结构化的、面向对象的代码提示,以通过代码补全诱导有害输出。在 8 个最先进的商业 LLM 上进行的实验表明,CodeMimicry 实现了 96.25% 的攻击成功率,平均仅需 1.51 次查询,显著优于基于模板和基于优化的基线。除了实证性能之外,我们还通过潜在空间表示对基于代码的越狱进行机制分析,包括投影到与拒绝相关的方向以及激活引导。这一分析提供了一种解释,说明 CodeMimicry 如何在与代码相关的领域中绕过安全机制。我们的发现揭示了当前安全对齐中的一个弱点,并强调了在代码等结构化领域中实现稳健对齐的必要性。
cs.LG / 81 / 2609.38636
STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding
STEPS:使用扩散与对比风格编码的保留风格场景文本编辑
diffusion
扩散模型相关
Abstract
We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.
Chinese Translation
我们提出了保留风格的场景文本编辑(STEPS),一种用于图像中高质量文本替换的新型扩散模型架构。场景文本编辑(STE),也称为视觉文本编辑,包括改变图像中的文本内容,同时保留原始风格,例如字体、颜色、方向、背景等。STEPS 通过有针对性地聚焦于改进的风格保留,推进了 STE 的最先进水平。我们引入了一种用于视觉文本的风格编码器,它独立于文本内容捕获风格,以及一种将风格编码器与多个语义条件(目标文本字符编码和渲染字形)相结合的模型架构。STEPS 在风格保留、输出可读性和主观质量方面取得了优于先前 STE 方法的结果。
cs.AI / 82 / 2609.38680
ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images
ReGain:在合成图像个性化中恢复主体保真度
diffusion
扩散模型相关
Abstract
Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.
Chinese Translation
文本到图像扩散模型通过在一个主体的少量图像上进行 DreamBooth 微调,从而实现对该主体的个性化。这些图像越来越多地来自扩散模型,而非相机。我们表明,在此类合成图像上微调会降低主体保真度,产生过饱和的颜色和过多的高频细节。为分离原因,我们从同一基础模型出发,用相同的 DreamBooth 配方微调两个模型:一个在某一主体的真实照片上微调,另一个在该主体由第一个模型生成的合成图像上微调。我们将这种退化追溯到无分类器引导(CFG)。对于在合成图像上个性化的模型,条件噪声预测与无条件噪声预测之间的夹角,以及随之而来的二者差值的范数,都远大于在真实照片上个性化的模型。这种膨胀在高频方向增长,并且也出现在与主体语义接近的其他提示上,例如其类别名词,但在不相关的提示上则不出现。我们提出 ReGain,一种在采样时应用的无训练校正,它测量引导的每个频带相对于基础模型被膨胀了多少,并相应地缩小该频带。ReGain 不需要真实照片。在 Stable Diffusion v1.5 上,按照 DINO、DINOv2 和 CLIP-I 的衡量,ReGain 缩小了与在真实照片上个性化模型之间主体保真度差距的 51-64%。它还在 SDXL 和 SD 3.5 上提高了主体保真度,并在所有三个骨干模型上保持了文本对齐。
cs.AI / 83 / 2609.38819
Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video
未来视频生成比观测到的视频更好地对齐人类视觉皮层
diffusion
扩散模型相关
Abstract
Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed visual stimuli, but also predicts upcoming stimuli based on what has been observed. Accordingly, we hypothesize that internal representations for generating future video frames are better aligned with the predictive nature of human visual processing than representations of the observed video itself. To this end, we compare the alignment between human video-watching fMRI responses in the visual cortex and the internal representations from two types of video diffusion models, an autoregressive (AR) model and its non-AR base model. We first conduct a within-model analysis of the AR video diffusion model and show that the representations for future video generation align better with the visual cortex than the representations of the observed video. We then compare the internal representations of the AR model with those of its non-AR base model and again show that the representations for future video generation align better with the visual cortex than the representations for observed video reconstruction by the base model. Specifically, the alignment of observed video reconstruction is concentrated in lower-order visual cortex, whereas that of future video generation is concentrated in higher-order visual cortex. Finally, we show in a human behavioral experiment that humans prefer videos generated by amplifying the contributions of individual layers that align better with the visual cortex.
Chinese Translation
研究视觉模型的内部表征与视觉皮层对相同观测视觉刺激的响应之间的对齐,使我们能够更好地理解人类视觉加工。然而,迄今为止的研究在很大程度上忽视了一个事实:人脑不仅加工观测到的视觉刺激,还基于已观测到的内容预测即将到来的刺激。据此,我们假设,用于生成未来视频帧的内部表征比观测视频本身的表征更好地对齐于人类视觉加工的预测性本质。为此,我们比较了人类观看视频时视觉皮层的 fMRI 响应与两类视频扩散模型(一个自回归(AR)模型及其非 AR 基础模型)的内部表征之间的对齐。我们首先对 AR 视频扩散模型进行模型内分析,并表明用于未来视频生成的表征比观测视频的表征更好地对齐于视觉皮层。随后,我们将 AR 模型的内部表征与其非 AR 基础模型的内部表征进行比较,并再次表明,用于未来视频生成的表征比基础模型用于观测视频重建的表征更好地对齐于视觉皮层。具体而言,观测视频重建的对齐集中在低级视觉皮层,而未来视频生成的对齐则集中在高级视觉皮层。最后,我们通过一项人类行为实验表明,人类更偏好通过放大那些与视觉皮层对齐更好的单个层的贡献而生成的视频。
cs.LG / 84 / 2609.39184
Fiber-Resolved Microstructure Quantification from Multi-Shell Diffusion MRI using Detection Transformers
使用检测Transformer从多壳扩散MRI中进行纤维分辨的微结构量化
diffusion
扩散模型相关
Abstract
Fiber orientation and compartmental microstructure are central to the characterization of white matter tissue in diffusion MRI, yet existing methods either resolve fiber orientations without quantifying microstructure, or quantify microstructure while assuming a fixed number of compartments and a single fiber direction. Nonparametric approaches that recover both require tensor-valued diffusion encoding and computationally expensive Monte-Carlo inversion of an ill-posed inverse Laplace transform. We propose to reframe this problem as an object detection-like task, adopting the Detection Transformer (DETR) architecture to jointly predict mean diffusivity (MD), fractional anisotropy (FA), main fiber direction, and signal fraction for a variable number of compartments per voxel from standard multi-shell diffusion MRI with linear encoding. Hungarian matching during training resolves permutation invariance across compartments. We introduce mean Average Precision as a reproducible benchmark metric. Evaluated on synthetic test data with up to five compartments per voxel, our model achieves $R^2=0.95$ for MD, $R^2=0.88$ for FA, and a median angular error of 4.2°, with performance scaling naturally with compartmental signal fraction.
Chinese Translation
纤维方向和区室微结构是扩散MRI中白质组织表征的核心,然而现有方法要么解析纤维方向而不量化微结构,要么在假设固定数量的区室和单一纤维方向的情况下量化微结构。同时恢复二者的非参数方法需要张量值扩散编码,以及对不适定逆拉普拉斯变换进行计算昂贵的蒙特卡洛反演。我们提出将该问题重新表述为类似目标检测的任务,采用检测Transformer(DETR)架构,从采用线性编码的标准多壳扩散MRI中,联合预测每个体素可变数量区室的平均扩散率(MD)、分数各向异性(FA)、主纤维方向和信号分数。训练期间的匈牙利匹配解决了跨区室的排列不变性。我们引入平均精度均值(mean Average Precision)作为可复现的基准指标。在每体素最多五个区室的合成测试数据上评估,我们的模型对MD达到 $R^2=0.95$,对FA达到 $R^2=0.88$,并具有4.2°的中位角度误差,其性能随区室信号分数自然缩放。
cs.AI / 85 / 2609.39441
CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models
CAST:面向扩散模型的、具有空间接地的组合奖励的因果优势结构化训练
diffusion
扩散模型相关
Abstract
Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.
Chinese Translation
在线强化学习已被扩展到用于扩散模型(DM)图像生成的流匹配。然而,这一范式面临三个局限:(1) 窗口选择。现有方法手动设置随机微分方程(SDE)采样窗口,即在哪些去噪步骤注入探索噪声。我们则从每个模型的去噪轨迹中确定它。(2) 奖励饱和。当前方法依赖于在人类标注上训练的评分模型;我们发现,在最新的 SOTA 开源 DM 上,此类分数极高且几乎无法区分,使得优势估计在很大程度上失效。(3) 样本低效。单一标量奖励将不同的失败模式压缩为几乎相同的分数,几乎不留下用于针对性改进的梯度指导。为解决这些问题,我们提出 CAST(Causal Advantage-Structured Training),一种用于预训练 DM 的 RL 微调方法,它 (1) 识别每个模型在图像中固定对象及其空间排布的去噪步骤,并使用该时机来设置 SDE 窗口,(2) 通过因果场景图(CSG)将每个提示分解为可验证原子,即可独立检查的最小语义单元,例如对象、数量、属性或空间关系,并分别对每个原子给予奖励,以及 (3) 通过教师强制注意力将带符号的原子级优势投影到像素空间,并使用它们在空间上加权 SDE 策略目标。我们使用 CAST 微调了两个最强的开源 DM,FLUX.2-dev 和 Qwen-Image-2512,并在组合基准 GenEval 2 上以及用于整体质量的 Qwen-Image-Bench 上评估它们。在几乎相同的训练预算内,CAST 在最具挑战性的 GenEval 2 提示上相对于基础模型的提升最高可达 Flow-GRPO 的 3.07 倍,同时整体生成质量也有所提高。
cs.AI / 86 / 2609.39504
PartiCam: Camera Controlled Video Generation with Reward Guidance
PartiCam:基于奖励引导的相机控制视频生成
diffusion
扩散模型相关
Abstract
We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid the need to construct large camera-annotated datasets by steering pretrained models toward the desired camera motion at test time. This enables the generation of camera-controlled video data that can subsequently be used to train camera-conditioned video diffusion models. Existing sampling-based guidance approaches often suffer from unstable trajectories: they either explore too broadly and fail to respect the target camera motion or collapse early and lose visual diversity over time. We introduce a global-local refinement framework for diffusion reward guidance, enabling accurate and consistent camera control during video generation. Our method builds on Sequential Monte-Carlo (SMC) guidance, but introduces a local refinement stage based on particle filtered resampling. Experiments show large improvements in camera trajectory adherence, reduced drift, and better visual quality, without requiring model retraining.
Chinese Translation
我们提出 PartiCam,一种无需训练的、以粒子滤波为基础的方法,用于改进相机控制的视频生成。对于大型视频扩散模型而言,生成遵循精确指定相机轨迹的视频仍然具有挑战性。无需训练的方法与骨干模型无关,并且通过在测试时引导预训练模型朝向期望的相机运动,避免了构建大规模相机标注数据集的需求。这使得能够生成相机控制的视频数据,这些数据随后可用于训练以相机为条件的视频扩散模型。现有的基于采样的引导方法常常面临轨迹不稳定的问题:它们要么探索范围过广而无法遵循目标相机运动,要么过早坍缩并随着时间推移丧失视觉多样性。我们引入了一种用于扩散奖励引导的全局-局部细化框架,使得在视频生成过程中能够实现准确且一致的相机控制。我们的方法建立在序列蒙特卡洛(Sequential Monte-Carlo, SMC)引导之上,但引入了基于粒子滤波重采样的局部细化阶段。实验表明,相机轨迹遵循度大幅提升、漂移减少、视觉质量更好,且无需重新训练模型。
cs.LG / 87 / 2609.39542
Comparative study of adapting pre-trained models for driving behavior video captioning
用于驾驶行为视频描述生成的预训练模型适配比较研究
large language model
大语言模型相关
Abstract
This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of different forms of data and this study aims to induce a low dimensional understanding of driving situations into our primary test model SpaceTimeGPT. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate good performance of the full fine tuning framework on some automatic metrics, and in some metrics, it even surpasses the baseline. We also try Low-Rank Adaptation (LoRA) and prompt engineering on VideoLLaVA model and discuss its limitations.
Chinese Translation
本报告考察并比较了现有的众多微调与提示方法中的一些方法,并将它们应用于自动驾驶领域。其思路是通过在视频数据集上适配一个大型语言模型(LLM)来比较这些方法。LLM 已经非常擅长很好地理解不同形式的数据,而本研究旨在将驾驶情境的低维理解引入我们的主要测试模型 SpaceTimeGPT。在 BDD-X(Berkeley DeepDrive eXplanation)数据集上的实验表明,全量微调框架在一些自动指标上表现良好,并且在某些指标上甚至超过了基线。我们还尝试了在 VideoLLaVA 模型上进行低秩适配(LoRA)和提示工程,并讨论了其局限性。
cs.AI / 88 / 2609.39548
Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models
学习正常扩散动力学以实现文本到图像模型中的后门防御
diffusion
扩散模型相关
Abstract
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.
Chinese Translation
后门攻击对文本到图像(T2I)扩散模型的安全部署构成了严重威胁。现有的防御方法通常从内部表征中的特定异常模式来检测后门,随着攻击机制日益多样化,这可能限制其泛化能力。本文从转移动力学的视角研究T2I扩散模型的后门防御。我们观察到,良性扩散轨迹在交叉注意力、潜在空间和噪声空间中表现出结构化的、依赖于时间步的转移模式,而后门攻击则倾向于导致偏离这种正常演化。受这些观察的启发,我们提出了正常扩散动力学学习(NDDL),这是一个新颖的后门防御框架,它仅利用良性样本学习扩散轨迹的正常转移动力学。NDDL构建紧凑的多空间轨迹表征,并训练一个以时间步为条件的动力学模型来预测扩散演化。在推理阶段,利用观测到的转移与预测的转移之间的偏差来量化动力学不一致性,以实现后门检测。NDDL还通过使用低语义词汇进行替换,在无需任何关于所嵌入后门的先验知识的情况下实现触发器定位。针对多种后门攻击的大量实验证明了我们提出的NDDL的有效性和泛化能力。
cs.AI / 89 / 2609.39600
GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
GroundAnything:以闪电速度协调并行解码与精确视觉定位
diffusion
扩散模型相关
Abstract
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
Chinese Translation
自回归(AR)定位模型将空间预测串行化,引入顺序延迟,并将因果顺序强加于输出词元之上。我们将定位视为视觉证据提取:目标、位置和空间关系由图像和查询共同约束,但它们的依赖关系并不意味着存在一种内在的从左到右生成顺序。这一区别使双向扩散成为一种自然契合,允许空间假设并行出现,并通过迭代去噪被联合精炼。我们提出 GroundAnything,一个 4B 参数的定位基础模型,它通过分块去噪协调快速并行解码与精确定位。训练结合了来自公共数据集和专用数据引擎的定位预训练、带有联合 AR 和扩散目标的直接 AR 到扩散转换、监督微调,以及基于 GRPO 的强化后训练。在 30 个定位基准上,我们的自回归变体 GroundAnything-VLM 以 72.42% 在同等规模模型中建立了新的整体最先进水平,与 GPT-6 Astra(71.35%)相比仍具竞争力。借助熵引导解码,GroundAnything 还超越了该规模下先前的最先进水平,平均达到 61.75%,而快速的基于 MTP 的 LocateAnything 模型为 53.32%。我们进一步探索了解码策略,表明可选的自推测模式相较于 AR 对应模型实现了 $4.51\times$ 的加速,同时在 COCO F1mIoU 上下降 0.74 个百分点。基础设施实验表明,渐进式推理优化将并行解码转化为实际加速。这些为延迟敏感的现实世界系统中的高效视觉定位提供了支持。
cs.AI / 90 / 2609.39625
D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
D-Scope:用稀疏自编码器分解和引导扩散 Transformer
diffusion
扩散模型相关
Abstract
Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text descriptions against these visual centroids in the shared image-text embedding space then enables retrieval of individual features without per-feature text annotations. The underlying patches provide evidence for inspecting each selection, while spatially masked interventions test the corresponding decoder direction at varying strengths under fixed generation conditions. We characterize 150 SAEs across two model families and five layers, and introduce a benchmark of 100 target concepts with ten contexts each spanning under-specified and explicit-conflict conditions. Our empirical results show that high reconstruction fidelity can coexist with low dictionary utilization and limited visual-evidence coverage. Under per-case best-of-sweep strength selection, contrastive retrieval yields larger mean regional SigLIP~2 gains than direct retrieval across the tested steering configurations, without consistently improving outside-region preservation. D-Scope provides an inspectable framework for evaluating sparse DiT features through their visual evidence and the effects of their decoder directions on generation. The demo is available at https://jiahaozhang-public.github.io/d-scope/.
Chinese Translation
稀疏自编码器(SAEs)揭示了扩散 Transformer(DiTs)中的视觉结构,但解释一个特征并不能确定它是否可用于控制生成。我们提出 D-Scope(Diffusion Scope),一个通过共享视觉证据将特征解释与生成控制联系起来的框架。D-Scope 将高激活图像块的 SigLIP~2 嵌入聚合成视觉质心。在共享的图像-文本嵌入空间中,将目标文本描述与这些视觉质心进行匹配,从而能够在无需逐特征文本标注的情况下检索单个特征。底层图像块为检查每个选择提供了证据,而在固定生成条件下,空间掩码干预以不同强度测试相应的解码器方向。我们对两个模型家族和五个层中的 150 个 SAE 进行了表征,并引入了一个包含 100 个目标概念的基准,每个目标概念有十个上下文,涵盖欠指定和显式冲突条件。我们的实证结果表明,高重建保真度可以与低字典利用率和有限的视觉证据覆盖率共存。在逐案例的最佳扫描强度选择下,对比检索在所测试的引导配置中比直接检索产生更大的平均区域 SigLIP~2 增益,但没有持续改善区域外保持。D-Scope 提供了一个可检查的框架,用于通过稀疏 DiT 特征的视觉证据及其解码器方向对生成的影响来评估这些特征。演示可在 https://jiahaozhang-public.github.io/d-scope/ 获取。
cs.LG / 91 / 2609.39635
SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
SAGE:基于视觉基础表征的显著因子发现与生成
diffusion
扩散模型相关
Abstract
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below $2$) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy $0.950$ vs.\ at most $0.281$) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent ($90.5\%$ vs.\ $27.7\%$) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.
Chinese Translation
给定一个目标数据集(例如戴眼镜的人脸)和一个背景数据集(例如不戴眼镜的人脸),对比分析将特定于目标的\textit{显著}因子与两者共享的\textit{共同}内容分离开来。我们旨在获得能够捕捉每幅图像中目标特定细节(例如眼镜的形状、颜色和位置)的显著表征,从而在无需子类型标签的情况下揭示子类型,并引导生成某个已发现子类型的新样本,甚至生成一个没有名称或文本描述的子类型的新样本。我们提出了 SAGE,它直接在冻结的表征自编码器的高维空间潜变量中学习这两种因子,并以参考图像学习到的显著表征为条件来调节扩散 Transformer。在 Digits-ImageNet 和 FFHQ 眼镜数据集上,SAGE 将高保真\textit{重建}(rFID 低于 $2$)与无监督\textit{子类型发现}相结合,比基线更好地恢复数字(探针精度 $0.950$ 对比至多 $0.281$),并揭示眼镜类型、更细粒度的太阳镜样式以及错误标注的图像;以显著为条件的\textit{生成}在 Digits-ImageNet 上将子类型准确率相对于未因子化潜变量提高($90.5\%$ 对比 $27.7\%$),并在两个数据集上提高多样性。在视网膜 OCT 上,SAGE 的显著空间仅使用正常/疾病标签就能分离三种疾病。
cs.AI / 92 / 2609.39688
ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
ShieldCLIP:面向多模态基础模型中有害内容缓解的选择性安全对齐
diffusion
扩散模型相关
Abstract
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at https://aimagelab.github.io/ShieldCLIP/.
Chinese Translation
诸如 CLIP 之类的多模态编码器是许多下游系统的基础,但其网络规模训练数据中嵌入了有害关联,安全对齐必须抑制这些关联,同时避免不必要地改变良性表征。由于伦理和实践限制使得无法大规模收集真实的不安全内容,现有数据集将安全的真实样本与生成的对应样本配对,但把每一个生成样本都标记为不安全,即使其中一个模态单独来看是安全的。为了解决这一问题,我们提出 ShieldCLIP,这是第一个根据每个模态所观察到的安全状态而非样本来源来调节安全对齐的框架,从而在保留安全内容的同时仅重定向不安全的内容。我们还提出 ViSUv2,这是一个包含 195k 个四元组的数据集,涵盖 578 个概念和 28 个类别,并具有独立的逐模态安全标签。利用这些标签,ShieldCLIP 定义了一个超越配对级监督的四路条件目标:安全内容被锚定,不安全的模态被重定向到其安全的对应模态,混合对仅更新不安全分支,并且当两者都不安全时强制保持一致性。我们在跨模态检索、使用 Stable Diffusion v1.4 和 SDXL 的文本到图像生成,以及使用 LLaVA 的图像到文本生成上评估 ShieldCLIP。在这些设置中,ShieldCLIP 相较于此前的安全对齐编码器和强大的缓解基线,持续减少有害输出,同时保留原始嵌入空间的效用。广泛的消融研究进一步表明,模态特定监督和选择性对齐目标都对这些增益有所贡献。源代码、训练好的模型以及 ViSUv2(在受控访问协议下)将在 https://aimagelab.github.io/ShieldCLIP/ 公开提供。
cs.AI / 93 / 2609.40079
LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
LongEmo:面向长视频中的情感理解与推理
large language model
大语言模型相关
Abstract
While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.
Chinese Translation
尽管近期的多模态大语言模型(MLLMs)在情感计算中已展现出潜力,但其推理能力在很大程度上仍局限于交互有限的短视频片段。然而,现实世界中的情感并不仅仅是孤立的瞬时反应,而是由过往经历与正在发生的事件深刻塑造的动态且累积的过程。为弥合这一差距,我们提出了 LongEmoBench,一个专用于长视频中情感理解与推理的基准。它评估从连续场景交互到复杂情节发展的渐进式能力。此外,我们提出了 LongEmo,一种新颖的、记忆增强的智能体框架,旨在应对长程情感推理的巨大挑战。LongEmo 处理连续视频流以构建事件记忆图,显式地建模长程依赖关系,并捕捉跨离散事件的情感动态。给定一个问题,该智能体从图中检索与查询相关的事件流,迭代地整合多模态记忆与关系依赖,以推导出最终答案。对 17 种代表性方法的广泛评估表明,它们在长视频中的情感理解与推理方面存在显著困难。相比之下,LongEmo 取得了最先进的性能,证明了其以事件为中心的记忆架构的有效性。
cs.LG / 94 / 2609.40305
Looped Diffusion Transformer
循环扩散 Transformer
diffusion
扩散模型相关
Abstract
Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.
Chinese Translation
提升文本到图像模型传统上依赖于增大模型规模或增加去噪步数。在这项工作中,我们探索了一种扩展计算的替代方式:在每个去噪步骤内重复运行共享的 Transformer 块,从而在保持参数量固定的同时有效增加计算深度。这种循环计算使得内部表示能够迭代式地精炼,而无需显式的推理令牌。然而,朴素的循环并不能持续一致地提升图像质量。我们将这一问题追溯到中间循环之间的弱监督,以及未受调控的注意力更新会逐步侵蚀局部信息。为克服这些挑战,我们提出了循环扩散 Transformer(Looped-DiT),它将跨中间循环的深度监督与自调制注意力相结合,以稳定循环中的特征更新。在参数量匹配和计算量匹配的设置下,Looped-DiT 一致地优于非循环基线。值得注意的是,一个 260M 参数的循环模型能够在多个文本到图像基准上超越一个规模大 6.5 倍的模型,同时所需的推理计算量低 4.9 倍。除了这一性能提升之外,我们发现循环计算可以为扩散模型提供一种更有效的迭代计算形式:在固定的推理预算下,增加循环深度所带来的收益大于增加更多的去噪步骤。此外,更深的循环能够逐步纠正早期循环中所犯的错误,展现出暗示潜在推理的行为。这些结果共同表明,循环计算为扩展视觉生成模型提供了一条有前景的途径。
cs.AI / 95 / 2609.40322
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
MatLoom:紧凑程序空间中的分层文本到材质生成
diffusion
扩散模型相关
Abstract
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.
Chinese Translation
材质生成不仅应产生外观,还应产生构建它的规则。我们提出 MatLoom,一种紧凑的、面向层的语言,用于借助预训练语言模型进行文本到材质生成。每个程序组合了 alpha 掩码层,其共享的空间表达式定义了覆盖范围和基于物理的渲染(PBR)通道,从而使图案、颜色和浮雕之间的依赖关系显式化。一个独立的解释器将程序求值为材质贴图,而源代码保留命名字段和层参数,以供后续创作。无需任务特定的微调,我们的流水线使用解析器引导的修复和基于预览的评审来修订材质设计,然后在保持每个候选其余源代码固定的同时搜索噪声种子。在一个由 141 个提示组成的精选基准上,使用六种骨干模型进行评估时,我们表现最佳的配置在所有四个平面布局提示对齐指标上都取得了高于三个扩散基线的平均分数。其初始程序在平均 BLIPScore 上已经超过所有三个基线,而且这发生在评审或种子搜索之前。当跨骨干模型汇总时,保留程序的中位长度为 21 行。在一项涉及 30 名参与者和 20 个提示的盲测四路比较中,我们的渲染结果获得了 59.2% 的选择,而最受偏好的基线为 19.3%。因此,紧凑的可执行程序提供了一种生成与提示对齐的材质的方法,同时将其构建过程保留为资产的一部分。
cs.CL / 96 / 2609.38486
Anthropomorphism in the age of Large Language Models: An overview of potential risks and mitigations
大语言模型时代的拟人化:潜在风险与缓解措施综述
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) and more broadly Artificial Intelligence (AI) systems are often described and understood in human-like terms, a phenomenon known as \emph{anthropomorphism}. This paper provides a synthesis of recent literature on anthropomorphism in AI, covering theoretical frameworks, the role of language in framing AI as human-like, the various risks of anthropomorphizing machines, and strategies to mitigate these issues. After examining why we tend to anthropomorphize AI systems and whether we are right to do so, we highlight the impact of linguistic framing on anthropomorphism. Then, we introduce a conceptual taxonomy of risks associated with AI anthropomorphism. This taxonomy groups twenty-one concerns within five analytical categories: epistemic, affective, human agency, normative, and societal and institutional risks. Finally, we relate these concerns to proposed interventions in design, communication, education, and governance. We argue that a better understanding of AI systems requires concepts and theories grounded in their organization and demonstrated capacities. The linguistic shaping of anthropomorphic perceptions should form part of this scientific effort, since our descriptions influence both how these systems are understood and the roles we allow them to occupy in society.
Chinese Translation
大语言模型(LLMs)以及更广泛意义上的人工智能(AI)系统,常常以类人的术语被描述和理解,这一现象被称为\emph{拟人化}。本文对近期关于AI拟人化的文献进行了综合梳理,涵盖理论框架、语言在将AI框定为类人存在中的作用、将机器拟人化所带来的各种风险,以及缓解这些问题的策略。在考察了我们为何倾向于将AI系统拟人化以及这种倾向是否恰当前,我们着重分析了语言框定对拟人化的影响。随后,我们提出了一种关于AI拟人化相关风险的概念性分类体系。该分类体系将二十一项关切归入五个分析类别:认识论风险、情感风险、人类能动性风险、规范性风险以及社会与制度性风险。最后,我们将这些关切与在设计、传播、教育和治理方面所提出的干预措施联系起来。我们认为,要更好地理解AI系统,需要以这些系统的组织方式和已展现出的能力为基础的概念与理论。对拟人化感知的语言塑造应成为这一科学努力的一部分,因为我们的描述既影响这些系统被理解的方式,也影响我们允许它们在社会中占据的角色。
cs.AI / 97 / 2609.38697
Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure
Cascadia:一种无需控制平面的超融合 AI 基础设施替代方案
large language model
大语言模型相关
Abstract
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.
Chinese Translation
我们提出 Cascadia,一个用于在商用 Intel AIPC 集群上利用其 CPU、集成 GPU 和 NPU 资源来服务大语言模型的系统。每个节点都内嵌了入口、调度与执行;推理请求不需要专用的路由控制平面。节点使用由 CA 签发的 ed25519 准入证书加入 libp2p QUIC 网状网络,通过 gossip 传播经过签名的能力信息,经由直接的对等流交换实时负载,并将兼容 OpenAI 的请求路由到符合条件的对等节点。由运维方运行的证书颁发机构在推理路径之外处理准入与集群管理。三种服务模式共享同一个接口:在单个节点上执行整个模型、负载均衡的副本,以及使用我们姊妹论文中的编译与推测解码机制的流水线分片链。可选的 KV 缓存迁移可在路由变更后复用兼容的对话前缀,未命中时则进行冷重计算。经过签名的响应回执与哈希链式日志为溯源与审计提供支持。一个三节点 Phi-3.5-mini NPU 测试平台在十个并发请求下,其响应吞吐量达到其单节点配置的 3.10 倍;另一个独立的四节点部署记录到的吞吐量为直接单节点服务的 4.06 倍。配对延迟观测、运行时测量以及内部功能检查刻画了所测试的配置。我们依据厂商文档,在部署占用空间、硬件要求、调度、扩展、许可与信任方面,将 Cascadia 与 IBM、Nutanix、VMware 和 HPE 平台进行比较。论文仓库提供了基准测试脚本、经过整理的测量数据,以及一份主张与证据对应表。
cs.AI / 98 / 2609.38353
TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories
TAGGRAPH:用于智能体持久历史图检索的标签增强图
diffusion
扩散模型相关
Abstract
Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.
Chinese Translation
长期记忆使 LLM 智能体能够回忆过去的交互并在跨会话中保持一致,但记忆系统难以比较,因为它们在表示、索引、检索和评估方面常常各不相同。我们提出了一个基于共享 5W 风格对话记忆的受控评估框架。局部化图配置遍历一个公共基础图;AdaptiveGraph 添加时间顺序边和 Personalized PageRank 扩散。我们还在相同的抽取笔记上评估 BM25,并将 OpenClaw 作为原始输入的外部参考进行评估。检索排名随记忆设置的不同而变化。在 LongMemEval-S 上,AdaptiveGraph 是最强的图配置,达到 0.844 MRR,但 BM25 达到 0.867,OpenClaw 达到 0.880。在 ATANT Core 上,局部化图遍历优于扩散和 BM25,而 BM25 在压力轮次中领先。在测试范围内缩减 LongMemEval-S 并未复现 ATANT 的扩散惩罚,但测试的最小存储仍大于 ATANT Core,因此存储规模不能排除。在宽松的内容匹配标准下,这种惩罚仍然存在。词汇归一化和抽取质量会显著影响图检索,并且缺失抽取标签在前五名未命中项中很常见。因此,检索策略应与记忆设置一起评估,并与强词汇基线进行对比。
cs.AI / 99 / 2609.38473
Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering
重排序与后期交互驱动检索质量:面向科学问答的RAG策略的受控比较
large language model
大语言模型相关
Abstract
Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is large and the trade-offs between variants are not well understood, especially on domain-specific corpora at realistic scale. In this work, we present a controlled comparison of six retrieval strategies for scientific question answering: (i) classic top-k dense retrieval, (ii) LLM-based query rephrasing, (iii) query rephrasing followed by LLM-based reranking, (iv) multi-query fusion via Reciprocal Rank Fusion (RRF), (v) an agentic tool-call pipeline in which the generator decides for itself whether to retrieve, and (vi) late-interaction retrieval with ColBERTv2. All six pipelines share the same generator (Meta-Llama/Llama-3.1-8B-Instruct), prompt, and evaluation protocol; the five single-vector pipelines additionally share SPECTER2 embeddings and a Chroma vector store; and all six retrieve from the full corpus of 463,971 arXiv papers dated 2024-2025. To support reproducible, large-scale evaluation, we also release a synthetic question dataset of 19,484 problem-statement and methodology questions generated by Llama-3.1-8B-Instruct from a random sample of 10,000 papers across academic domains (query generation succeeded for 9,742 of them), and every strategy is evaluated on this same query set. We describe the architecture and implementation of each pipeline, release the code and the synthetic question dataset, and evaluate each strategy with an LLM-as-a-judge protocol along multiple quality dimensions, together with direct gold-paper retrieval metrics. The result is an open testbed for studying the cost and quality trade-offs of RAG design choices on a research-literature corpus, and a basis for future work on faithfulness, retrieval robustness, and agentic retrieval.
Chinese Translation
检索增强生成(Retrieval-Augmented Generation,RAG)现在是将大型语言模型(Large Language Models,LLMs)建立在外部知识之上的标准方式,然而检索流水线的设计空间很大,并且不同变体之间的权衡尚未得到充分理解,尤其是在现实规模的特定领域语料库上。在这项工作中,我们提出了对科学问答的六种检索策略的受控比较:(i) 经典 top-k 稠密检索,(ii) 基于 LLM 的查询改写,(iii) 查询改写后接基于 LLM 的重排序,(iv) 通过倒数排序融合(Reciprocal Rank Fusion,RRF)进行的多查询融合,(v) 一种智能体工具调用流水线,其中生成器自行决定是否进行检索,以及 (vi) 使用 ColBERTv2 的后期交互检索。所有六种流水线共享相同的生成器(Meta-Llama/Llama-3.1-8B-Instruct)、提示词和评估协议;五种单向量流水线还共享 SPECTER2 嵌入和一个 Chroma 向量存储;并且所有六种流水线都从包含 463,971 篇日期为 2024-2025 年的 arXiv 论文的完整语料库中进行检索。为了支持可复现的大规模评估,我们还发布了一个包含 19,484 个问题陈述型和方法论型问题的合成问题数据集,该数据集由 Llama-3.1-8B-Instruct 从跨学术领域的 10,000 篇论文的随机样本生成(其中 9,742 篇的查询生成成功),并且每一种策略都在这一相同的查询集上进行评估。我们描述了每种流水线的架构和实现,发布了代码和合成问题数据集,并使用 LLM 作为评判者的协议沿多个质量维度评估每种策略,同时结合直接的金标准论文检索指标。其结果是一个开放测试平台,用于研究在研究文献语料库上 RAG 设计选择的成本与质量权衡,并为关于忠实性、检索鲁棒性和智能体检索的未来工作奠定基础。
cs.CL / 100 / 2609.40241
Decision-Oriented Recommendation Reranking: An Empirical Study of Jev
面向决策的推荐重排序:Jev 的实证研究
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate whether a decision-oriented model provides a useful alternative when the reranking task is fundamentally a structured choice among predefined candidate items. Specifically, we conduct a controlled empirical study of Jev, described by TypeSafe AI as a ``System One Model,'' for personalized recommendation reranking and compare it with recommendation-specific models and pointwise and listwise Qwen rerankers across multiple Amazon Reviews domains and candidate-set sizes, evaluating both recommendation effectiveness and observed serving latency. Our results show that Jev maintains strong recommendation effectiveness relative to the evaluated baselines while exhibiting substantially more gradual latency growth than the pointwise Qwen rerankers, although its observed serving latency remains substantially higher than that of recommendation-specific models. Together, these characteristics place Jev in a distinct quality--latency operating regime across candidate sizes and domains. These findings motivate further investigation of decision-oriented models for recommendation and other ranking tasks with structured output spaces.
Chinese Translation
大语言模型(LLMs)在推荐重排序中已展现出前景,但其使用在推荐质量与服务效率之间引入了一个重要的权衡。我们探究了当重排序任务本质上是在预定义候选物品之间进行结构化选择时,面向决策的模型是否能提供一种有用的替代方案。具体而言,我们针对 Jev——TypeSafe AI 将其描述为“System One Model”——在个性化推荐重排序中开展了一项受控的实证研究,并将其与推荐专用模型以及 pointwise 和 listwise 的 Qwen 重排序器进行比较,涵盖多个 Amazon Reviews 领域和不同候选集规模,同时评估推荐效果与观测到的服务延迟。我们的结果表明,相对于所评估的基线,Jev 保持了较强的推荐效果,同时其延迟增长比 pointwise 的 Qwen 重排序器要平缓得多,尽管其观测到的服务延迟仍显著高于推荐专用模型。综合来看,这些特性使 Jev 在不同候选规模和不同领域下处于一个独特的质量—延迟运行区间。这些发现促使我们进一步研究面向决策的模型在推荐以及其他具有结构化输出空间的排序任务中的应用。
cs.LG / 101 / 2609.38363
Simulator-Refined Diffusion for Radio-Frequency Inverse Design
面向射频逆向设计的模拟器精炼扩散
diffusion
扩散模型相关
Abstract
Diffusion models have shown potential in inverse design of printed circuit boards (PCBs), enabling the generation of layouts conditioned on target S-parameters. Despite this promise, applying diffusion models to PCB layout generation remains challenging due to their difficulty in meeting the quantitative electromagnetic specifications. A common approach is gradient-based guidance, which biases the diffusion sampling process with the gradient of an objective used for evaluation. However, full-wave electromagnetic simulators are accurate but expensive and typically non-differentiable, whereas differentiable surrogates are informative but not always reliable. To address these limitations, this paper proposes Simulator-Refined Diffusion (SRD), a novel combination of a low-fidelity differentiable surrogate and a high-fidelity non-differentiable simulator within the diffusion sampling process. Unlike standard zeroth-order optimization, which requires a great number of random perturbations, our approach uses the surrogate's gradient to propose the perturbation direction while the simulator then searches based on this direction to identify an effective design update. Experimental results across different settings show that this method consistently outperforms current state-of-the-art methods, producing layouts whose simulated S-parameters match the target specifications up to 21.2% closer for in-distribution targets and up to 19.8% for out-of-distribution targets.
Chinese Translation
扩散模型在印刷电路板(PCB)的逆向设计中已展现出潜力,能够实现以目标 S 参数为条件生成布局。尽管前景可观,但将扩散模型应用于 PCB 布局生成仍然具有挑战性,因为它们难以满足定量的电磁规格。一种常见方法是基于梯度的引导,它利用用于评估的目标函数的梯度来偏置扩散采样过程。然而,全波电磁仿真器虽然准确,但代价高昂且通常不可微,而可微代理模型虽然信息丰富,却并不总是可靠。为了解决这些局限,本文提出了模拟器精炼扩散(SRD),这是一种在扩散采样过程中将低保真可微代理模型与高保真不可微仿真器相结合的新方案。与需要大量随机扰动的标准零阶优化不同,我们的方法使用代理模型的梯度来提出扰动方向,而仿真器随后基于该方向进行搜索,以确定有效的设计更新。在不同设置下的实验结果表明,该方法始终优于当前最先进的方法,其生成的布局的仿真 S 参数与目标规格的匹配程度,对于分布内目标最多提升 21.2%,对于分布外目标最多提升 19.8%。
cs.LG / 102 / 2609.38383
Learning to Plan from Random Exploration
从随机探索中学习规划
diffusion
扩散模型相关
Abstract
Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.
Chinese Translation
随机探索揭示了在目标被指定之前,一个环境可以如何被穿行。这种经验能否在没有策略改进训练的情况下支持长程规划?我们的随机游走分析解释了时间关系包含什么:短时域在扩散极限下揭示测地几何,而更长时域则在混合消除这些区别之前揭示区域之间的连通性。我们用一个条件基于能量的模型来学习这些关系,该模型通过以时域为条件的嵌入来估计时间对数密度比。该模型通过噪声对比估计在观测对上进行训练,不使用动作或奖励标签。规划器在朝目标移动时,在不同时域查询这些学到的关系。在测试时,一个单独的局部动力学模型预测候选动作结果,而时间模型通过选择或聚合跨时域的估计改进来评估它们朝目标的进展。智能体执行一个动作,并在两个模型固定的情况下重新规划。实验展示了使用状态和图像从随机探索中进行长程迷宫规划。学到的分数场、嵌入探测和规划路线展现出多尺度认知地图的特性。我们进一步展示了从随机探索中进行的以自我为中心的导航,以及从次优数据中进行的操作规划。
cs.LG / 103 / 2609.38523
MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting
MM-FinEval:面向真实世界金融预测的多任务多模态基准
large language model
大语言模型相关
Abstract
Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimodal large language models (LLMs) can support real-world financial analysis. In this paper, we introduce MM-FinEval, a novel benchmark designed to evaluate multimodal LLMs across multiple financial tasks. MM-FinEval spans a diverse timeline from 2019 to 2022. The entire proposed dataset contains 2,045 S\&P 500 conference earning calls as inputs and 12 financial task labels as outputs. Each input contains three modalities: a word-to-word text transcript of the earning call, the corresponding presentation slides used during the call, and the entire audio recording. To establish a rigorous evaluation framework, we analyze 19 baseline models across three distinct model categories: Image-Text, Audio-Text, and Any-to-Any configurations. We observe that small-size Any-to-Any models processing all three modalities achieve strong performance, even when compared against larger proprietary models restricted to two-modality inputs. This indicates that our tri-modal dataset design introduces useful, non-redundant information. These results validate that text, audio, and visual data serve as important, complementary signals that mimic the decision-making process of expert human analysts.
Chinese Translation
基于财报电话会议的金融预测要求模型对复杂的公司披露、市场预期和微妙的沟通信号进行推理。然而,现有的金融基准通常局限于单模态输入或单任务设置,这使得难以评估多模态大语言模型(LLMs)是否能够支持真实世界的金融分析。在本文中,我们介绍了 MM-FinEval,这是一个新颖的基准,旨在评估多模态 LLMs 在多个金融任务上的表现。MM-FinEval 涵盖从 2019 年到 2022 年的多样化时间线。整个提出的数据集包含 2,045 个 S&P 500 财报电话会议作为输入,以及 12 个金融任务标签作为输出。每个输入包含三种模态:财报电话会议的逐字文本转录、通话期间使用的相应演示幻灯片,以及完整的音频录音。为了建立严格的评估框架,我们分析了跨越三个不同模型类别的 19 个基线模型:图像-文本、音频-文本和任意到任意(Any-to-Any)配置。我们观察到,处理所有三种模态的小型 Any-to-Any 模型取得了强劲性能,即使与仅限于双模态输入的更大专有模型相比也是如此。这表明我们的三模态数据集设计引入了有用的、非冗余的信息。这些结果验证了文本、音频和视觉数据作为重要的互补信号,模拟了人类专家分析师的决策过程。
cs.LG / 104 / 2609.38536
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents
这个动作还能解释任务吗?面向扩散语言模型智能体的反向评分
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action long after it has failed. We give a mechanistic account of this failure and a training-free remedy. We trace the retry loop to the adaptivity of masked decoding: the sampler commits the positions it is most confident about and defers the uncertain ones, and at a failure state the context already offers a confident fill for the deferred decision, i.e. the failed action itself, so the retry is committed without the failure feedback ever being confronted. We model the resulting distortion of the action distribution as a task-blind corruption: contextually salient actions (e.g., the action just taken) receive inflated probability by a factor that depends on the state and the action but not on the task. Under this model, we analyse an invariance proposition: the task-blind factor cancels exactly from the reverse conditional, i.e. the likelihood of the task given the state and a candidate action, which coincides with the task posterior of an idealized uncorrupted model. Masked dLLMs evaluate the reverse conditional natively, unlike autoregressive models, by masking the task tokens and denoising, at the cost of a few parallel passes per candidate. We instantiate the rule as Reflect Reverse and evaluate it on four multi-turn embodied benchmarks, where it improves task success and progression rates over forward-scoring baselines.
Chinese Translation
基于扩散的大语言模型(dLLMs)有望通过并行解码打破自回归智能体的串行延迟瓶颈,但近期的评估表明,这种效率并不能迁移到具身智能体能力上:由 dLLM 驱动的智能体会反复陷入重试循环,在一个动作早已失败之后仍重新发出该动作。我们给出了对这一失败的机制性解释,以及一种无需训练的补救方法。我们将重试循环追溯到掩码解码的自适应性:采样器会确定它最有把握的位置,而推迟不确定的位置,而在失败状态下,上下文已经为被推迟的决策提供了一个有把握的填充,即失败动作本身,因此重试被确定下来,而失败反馈从未被真正面对。我们将由此产生的动作分布扭曲建模为一种任务盲的污染:上下文显著的动作(例如刚刚执行的动作)所获得的概率会被一个因子抬高,该因子取决于状态和动作,但不取决于任务。在该模型下,我们分析了一个不变性命题:该任务盲因子会从反向条件中精确地约去,即给定状态和一个候选动作时任务的可能性,它与理想化的未受污染模型的任务后验相一致。与自回归模型不同,掩码 dLLM 能够原生地评估该反向条件,其方式是对任务 token 进行掩码并去噪,代价是每个候选需要若干次并行前向计算。我们将该规则实例化为 Reflect Reverse,并在四个多轮具身基准上对其进行评估,在这些基准上,它相较于前向评分的基线方法提升了任务成功率和推进率。
cs.LG / 105 / 2609.38585
FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
FlexRouter:学习用于灵活LLM路由的互补模型集合
large language model
大语言模型相关
Abstract
Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
Chinese Translation
现有的大语言模型(LLM)路由方法独立地对各个LLM进行打分,以选择 top-$k$ 模型。然而,这种做法忽略了模型之间的相关性,并施加了刚性的计算预算。因此,路由器往往会选择共享失效模式的冗余模型,从而限制了整体成功概率。为解决这一问题,我们提出了 FlexRouter,一个显式建模模型互补性的路由框架。FlexRouter 以 \textit{答案覆盖} 为优化目标,最大化至少有一个被选模型给出正确回答的概率。该目标与实际推理流程相一致,在这些流程中会生成多个候选输出,并由下游验证器或用户选择最终输出。我们将路由形式化为一个面向覆盖的子集选择问题,并使用行列式点过程(DPPs)对路由策略进行建模,DPPs 能够自然地同时刻画模型的能力与冗余性。为了在不需要真实目标子集的情况下直接优化覆盖,我们引入了一种基于对失败集进行边缘化的训练目标。在推理阶段,我们采用一种基于边缘对数行列式增益的贪心策略,使路由器能够在没有预定义预算的情况下自适应地确定子集大小。在大规模 RouterEval 基准上的大量实验表明,与强基线相比,我们提出的 FlexRouter 在域内和域外任务上均以更低的冗余度实现了更高的覆盖率,同时保持了灵活的推理成本。
cs.LG / 106 / 2609.38599
JARQ: Joint Alternating Refinement for Quantization
JARQ:用于量化的联合交替精化
large language model
大语言模型相关
Abstract
Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale solves, and keeps the host's bit width, groups, zero points, and inference cost. Across Llama-2, Llama-3, and Qwen models with RTN, GPTQ, OmniQuant, and AWQ hosts, JARQ lowers perplexity in 90 of 96 comparisons, cuts three-bit RTN perplexity by up to 36%, raises mean multiple-choice accuracy in 23 of 24 configurations, and improves QEP, QuaRot, and OJBKQ outputs, at under a minute per 7B block.
Chinese Translation
针对大语言模型的分组式训练后量化器将权重舍入到一个网格上,而该网格并未根据由此得到的整数编码重新拟合。我们表明,这会损失本可获得的精度:最佳网格取决于编码,输入相关性会将不同组的误差耦合在一起,而有用的编码改变常常同时涉及许多编码。我们提出 JARQ,一种即插即用式精化方法,它从任意分组量化器出发,交替进行对所有组尺度的联合最小二乘拟合,以及在当前网格上将同一组的许多编码一起移动的有界 Babai 提议。该问题是一个双线性、带箱约束的混合整数最小二乘问题;该求解器无需反向传播,在精确尺度求解下不会增加逐层目标,并保持宿主的位宽、分组、零点和推理成本。在采用 RTN、GPTQ、OmniQuant 和 AWQ 宿主的 Llama-2、Llama-3 和 Qwen 模型上,JARQ 在 96 次比较中的 90 次降低了困惑度,将三比特 RTN 困惑度最多降低 36%,在 24 种配置中的 23 种提高了平均多项选择准确率,并改进了 QEP、QuaRot 和 OJBKQ 的输出,且每个 7B 块耗时不到一分钟。
cs.LG / 107 / 2609.38632
Proper Scoring Rule-based Diffusion for Probabilistic Weather Forecasting
基于适当评分规则的扩散模型用于概率天气预报
diffusion
扩散模型相关
Abstract
Recent probabilistic weather forecasters train stochastic predictors with the continuous ranked probability score (CRPS) to generate each ensemble member in a single forward pass. These models learn the predictive distribution from the forecast context alone, which becomes difficult at longer forecast horizons where uncertainty is high. To learn the predictive distribution more effectively, we introduce auxiliary conditional denoising tasks that predict the same future state from the context and its corrupted version, which provides partial future information that can reduce prediction ambiguity. Building on distributional diffusion models, we learn the conditional distributions of these tasks with a single stochastic predictor by minimizing a proper scoring rule across noise levels. At inference, the predictor can still generate each ensemble member in a single forward pass at the fully corrupted endpoint. Standard CRPS training is recovered as the endpoint-only special case of our formulation, so our framework extends existing CRPS-based forecasters with only additional conditioning inputs. Controlled experiments show that the auxiliary tasks improve one-step forecasting across architectures, with larger gains at longer forecast horizons. The gains extend to high-dimensional global weather forecasting under both training from scratch and fine-tuning, along with improved calibration and potential benefits for generalization under distribution shift.
Chinese Translation
近期的概率天气预报模型使用连续排序概率评分(CRPS)训练随机预测器,以单次前向传播生成每个集合成员。这些模型仅从预报背景中学习预测分布,在预报时效较长、不确定性较高时,这一学习变得困难。为了更有效地学习预测分布,我们引入了辅助的条件去噪任务,从背景及其损坏版本中预测相同的未来状态,这提供了可以降低预测模糊性的部分未来信息。在分布扩散模型的基础上,我们通过最小化跨噪声水平的适当评分规则,用单个随机预测器学习这些任务的条件分布。在推理时,预测器仍可在完全损坏的端点处以单次前向传播生成每个集合成员。标准的CRPS训练被恢复为我们公式中仅端点处的特例,因此我们的框架仅通过额外的条件输入扩展了现有的基于CRPS的预报器。受控实验表明,辅助任务改善了跨架构的单步预报,在更长的预报时效上增益更大。这些增益扩展到高维全球天气预报,无论是从头训练还是微调,同时校准得到改善,并且在分布偏移下具有潜在的泛化优势。
cs.LG / 108 / 2609.38672
Provable Test-Time Scaling for Beam Search in LLM Reasoning
LLM 推理中束搜索的可证明测试时扩展
large language model
大语言模型相关
Abstract
Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning efficiency and more favorable test-time cost scaling. Despite strong empirical success, the theoretical understanding of beam search remains limited. In this paper, we study the test-time compute guarantee of the commonly used beam search framework that uses the model's internal log-likelihood for intermediate scoring, while relying on an external reward model only after a complete response is generated. We first establish a lower bound for vanilla beam search, showing that at least $Ω(C^\star(x)^2)$ samples are required for the optimal response to survive, where $C^\star(x)$ is the token-level coverage coefficient for prompt $x$. This motivates our modified confidence-filtered beam search (CF-Beam), which reduces the sufficient coverage dependence from quadratic to nearly linear under prefix competitiveness, for fixed horizon, gap, and target accuracy. We then show that the regret of CF-Beam is upper-bounded by the probability of rare failure events and the reward estimation error scaled by a path-level coverage coefficient, where the rare-failure term vanishes as per-step sampling increases. Our results highlight a fundamental advantage of beam search over sequence-level inference methods such as Best-of-N and Best-of-Majority. While the guarantees of these approaches typically involve coverage coefficients that grow exponentially with the horizon $L$, CF-Beam controls the dominant search-induced term through a token-level coverage coefficient that scales polynomially with $L$. Our numerical experiments further confirm that beam search is more robust on hard instances and under increasing reasoning horizons.
Chinese Translation
基于束搜索的测试时方法提供了一种有效途径,通过尽早剪除无效的推理路径来提升大语言模型(LLM)在长时程生成上的性能,从而显著提升推理效率,并带来更有利的测试时成本扩展。尽管在经验上取得了显著成功,对束搜索的理论理解仍然有限。在本文中,我们研究常用束搜索框架的测试时计算保证,该框架使用模型内部的对数似然进行中间打分,而仅在完整回复生成之后才依赖外部奖励模型。我们首先为原始束搜索建立一个下界,表明至少需要 $Ω(C^\star(x)^2)$ 个样本才能使最优回复存活,其中 $C^\star(x)$ 是提示 $x$ 的词元级覆盖系数。这促使我们提出改进的置信度过滤束搜索(CF-Beam),在固定时程、间隔和目标精度下,并在前缀竞争性条件下,它将充分覆盖依赖从二次降低到近乎线性。随后我们表明,CF-Beam 的遗憾上界由稀有失败事件的概率以及按路径级覆盖系数缩放的奖励估计误差所界定,其中稀有失败项随着每步采样数的增加而消失。我们的结果凸显了束搜索相较于 Best-of-N 和 Best-of-Majority 等序列级推理方法的根本优势。尽管这些方法的保证通常涉及随时程 $L$ 指数增长的覆盖系数,CF-Beam 却通过一个随 $L$ 多项式增长的词元级覆盖系数来控制主导的搜索诱导项。我们的数值实验进一步证实,束搜索在困难实例上以及推理时程不断增加时都更为稳健。
cs.LG / 109 / 2609.38776
Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution
蒸馏扩散分数差异以实现高效的训练数据归因
diffusion
扩散模型相关
Abstract
Training data attribution for diffusion models aims to identify the training samples that influence a generated instance, but existing methods either require costly per-sample gradient computation or query-specific model optimization. Moreover, most methods attribute changes in a proxy loss rather than changes in the actual model's generative behavior. We address these limitations by formulating attribution directly with a local score discrepancy measure, which applies to any diffusion variant (including DDPM, EDM, and flow matching), and by showing that such measure can be estimated without retraining, as a preconditioned gradient similarity. We instantiate this estimator as Training-data Influence via score Discrepancy (TID), which uses Kronecker-factored curvature to avoid random projections and per-sample gradient storage. We then distill TID into TIDE, a forward-only student trained online to reproduce the teacher's rankings from the diffusion model's internal activations. Under counterfactual evaluation on CIFAR-10, ArtBench-10, and MS-COCO, TID matches or outperforms state-of-the-art approaches, while TIDE retains most of TID's accuracy at four to five orders of magnitude lower per-query cost, attributing generated samples in milliseconds and faster than the generation itself.
Chinese Translation
扩散模型的训练数据归因旨在识别影响某个生成实例的训练样本,但现有方法要么需要昂贵的逐样本梯度计算,要么需要针对查询的模型优化。此外,大多数方法归因的是代理损失的变化,而非实际模型生成行为的变化。我们通过直接以局部分数差异度量来形式化归因,从而解决这些局限;该度量适用于任何扩散变体(包括 DDPM、EDM 和流匹配),并且我们表明该度量可以在无需重新训练的情况下进行估计,即作为一种预条件梯度相似度。我们将该估计器实例化为基于分数差异的训练数据影响(Training-data Influence via score Discrepancy,TID),它使用 Kronecker 因子化曲率来避免随机投影和逐样本梯度存储。随后我们将 TID 蒸馏为 TIDE,这是一个仅前向的学生模型,通过在线训练从扩散模型的内部激活中复现教师的排名。在 CIFAR-10、ArtBench-10 和 MS-COCO 上的反事实评估中,TID 匹配或优于最先进的方法,而 TIDE 在每次查询成本低四到五个数量级的情况下保留了 TID 的大部分准确率,能在毫秒级完成对生成样本的归因,且比生成本身更快。
cs.LG / 110 / 2609.38781
ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual Density
ChartDensity-Bench:面向视觉密度下数值数据重建的多模态大语言模型基准测试
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) offer a promising approach for recovering numerical data from scientific charts, but their ability to reconstruct chart data from visually dense figures remains poorly understood. Existing chart understanding benchmarks primarily evaluate question answering or chart-level reasoning and provide limited support for evaluating structured numerical reconstruction from scientific figures. We introduce \textbf{ChartDensity-Bench}, a benchmark for evaluating MLLMs on structured numerical data reconstruction from compound chart figures under controlled visual density. Built from charts paired with source-level ground-truth data, ChartDensity-Bench systematically varies the number of simultaneously presented charts ($k\in{1,3,6,9}$), enabling controlled evaluation of density-induced degradation. We further propose a multi-dimensional evaluation framework covering structural reliability, reconstruction completeness, parseability, and numerical fidelity. Experiments on five recent MLLMs show that numerical reconstruction generally degrades as visual density increases, while the magnitude of degradation varies substantially across models. Chart-level paired comparisons further show that the same source chart can incur higher reconstruction error when embedded in denser visual contexts. These findings highlight visual density as an important and previously underexplored factor in MLLM chart data reconstruction and provide a systematic benchmark for evaluating model robustness in this setting.
Chinese Translation
多模态大语言模型(MLLMs)为从科学图表中恢复数值数据提供了一种有前景的途径,但它们从视觉密集型图形中重建图表数据的能力仍然鲜为人知。现有的图表理解基准主要评估问答或图表级推理,对评估从科学图形中进行的结构化数值重建所提供的支持有限。我们提出了 \textbf{ChartDensity-Bench},这是一个用于在受控视觉密度下评估 MLLMs 从复合图表图形中进行结构化数值数据重建的基准。ChartDensity-Bench 构建自与源级真值数据配对的图表,系统性地改变同时呈现的图表数量($k\in{1,3,6,9}$),从而能够对由密度引起的退化进行受控评估。我们进一步提出了一个多维评估框架,涵盖结构可靠性、重建完整性、可解析性和数值保真度。在五个近期 MLLMs 上的实验表明,数值重建通常会随着视觉密度的增加而退化,而退化的幅度在不同模型之间差异很大。图表级配对比较进一步表明,同一源图表在被嵌入到更密集的视觉语境中时,可能会产生更高的重建误差。这些发现凸显了视觉密度是 MLLM 图表数据重建中一个重要且此前尚未被充分探索的因素,并为评估该情境下的模型鲁棒性提供了一个系统性基准。
cs.LG / 111 / 2609.38806
Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems
黑板智能在全局约束问题上能够超越自回归
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by next-token prediction itself. We study this question through blackboard intelligence: an inference-time perspective in which a model works on a fixed, revisable canvas and searches over candidate solution states rather than committing to a causal, left-to-right trajectory. We instantiate this idea with diffusion language models, whose any-order prediction interface naturally exposes predictions over partially filled solution states. Our key observation is that mean confidence, a simple model-internal quantity available from the standard masked diffusion objective, provides a useful proxy for global coherence and can guide inference-time search and revision. Empirically, across ZebraLogic, Nurse Rostering, and Job-Shop Scheduling, Blackboard consistently improves inference while holding the fine-tuned LLaDA-8B-Instruct checkpoint fixed and substantially outperforms same-scale autoregressive baselines, reaching 90.4% accuracy on ZebraLogic-Hard, 76.4% exact feasibility on Nurse Rostering, and 80.2% optimality on JSSP. Stronger autoregressive search and refinement also fail to close the gap on ZebraLogic-Hard, while Blackboard surpasses tested frontier LLMs there and on JSSP despite their substantially greater scale and strong test-time reasoning. We open-source our codebase at https://github.com/jwoosang1/blackboard-intelligence.
Chinese Translation
下一词元预测推动了大语言模型的显著进展,但越来越多的证据表明,它们在受复杂全局约束支配的问题上可能举步维艰。在这项工作中,我们关注这一情形,并追问这些局限中是否有一些源于下一词元预测本身所诱导的推理接口。我们通过黑板智能来研究这个问题:一种推理时视角,其中模型在一个固定的、可修订的画布上工作,并在候选解状态上进行搜索,而不是固守因果的、从左到右的轨迹。我们用扩散语言模型来实例化这一思想,其任意顺序预测接口自然能够给出对部分填充的解状态的预测。我们的关键观察是,平均置信度——一种可从标准掩码扩散目标中获得的简单模型内部量——为全局连贯性提供了有用的代理指标,并能够指导推理时的搜索与修订。从经验上看,在 ZebraLogic、Nurse Rostering 和 Job-Shop Scheduling 上,黑板智能在保持微调后的 LLaDA-8B-Instruct 检查点固定的同时,始终改善推理,并显著优于同等规模的自回归基线,在 ZebraLogic-Hard 上达到 90.4% 的准确率,在 Nurse Rostering 上达到 76.4% 的精确可行性,在 JSSP 上达到 80.2% 的最优性。更强的自回归搜索与精化也未能缩小在 ZebraLogic-Hard 上的差距,而黑板智能在该任务和 JSSP 上超过了所测试的前沿 LLM,尽管这些模型规模大得多且具备强大的测试时推理能力。我们在 https://github.com/jwoosang1/blackboard-intelligence 开源了我们的代码库。
cs.LG / 112 / 2609.38830
SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs
SparLeak:共享 GPU 上 LLM 推理中稀疏注意力的隐私泄露
large language model
大语言模型相关
Abstract
Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior introduces previously unexplored privacy risks. We identify a new GPU micro-architectural side channel, termed Sparsity-Induced Memory Access (SIMA), which arises from secret-dependent key-value cache access patterns induced by sparse attention. Based on this observation, we present SparLeak, a phase-aware side-channel attack that extracts SIMA traces during LLM inference and enables two practical privacy extractions: query attribute inference from prefill-phase traces and autoregressive response reconstruction from decoding-phase traces. By reconstructing approximate token-level sparsity profiles from page-level observations and applying profiling-based learning, SparLeak accurately recovers sensitive information, including user-query attributes and private LLM response content. Extensive evaluation across three LLM architectures, three sparse attention mechanisms, and three privacy-sensitive datasets shows that SparLeak achieves average attack success rates of 90.9% for attribute inference and 87.3% for response reconstruction under real-world LLM serving settings, highlighting the significance to account for SIMA leakage when deploying sparse-attention-based LLM systems. We provide anonymized SIMA traces, trained attack models, evaluation scripts, and documentation as artifacts at https://anonymous.4open.science/r/Janus_artifacts/.
Chinese Translation
稀疏注意力被广泛用于加速现代大型语言模型(LLM)中的长上下文推理,但其依赖于输入的执行行为引入了此前未被探索的隐私风险。我们识别出一种新的 GPU 微架构侧信道,称为稀疏诱导内存访问(Sparsity-Induced Memory Access, SIMA),它源于稀疏注意力所引发的依赖于秘密的键值缓存访问模式。基于这一观察,我们提出 SparLeak,一种阶段感知的侧信道攻击,它提取 LLM 推理期间的 SIMA 踪迹,并实现两种实际的隐私提取:从预填充阶段踪迹中进行查询属性推断,以及从解码阶段踪迹中进行自回归响应重构。通过从页级观察中重构近似的 token 级稀疏性轮廓,并应用基于剖析的学习,SparLeak 能够准确恢复敏感信息,包括用户查询属性和私密的 LLM 响应内容。跨三种 LLM 架构、三种稀疏注意力机制和三个隐私敏感数据集的广泛评估表明,在真实世界的 LLM 服务设置下,SparLeak 在属性推断上达到了 90.9% 的平均攻击成功率,在响应重构上达到了 87.3%,凸显了在部署基于稀疏注意力的 LLM 系统时考虑 SIMA 泄露的重要性。我们提供匿名的 SIMA 踪迹、训练好的攻击模型、评估脚本和文档作为工件,位于 https://anonymous.4open.science/r/Janus_artifacts/。
cs.LG / 113 / 2609.38853
Visualizing Distribution Coverage in Generative Diffusion Models
可视化生成扩散模型中的分布覆盖
diffusion
扩散模型相关
Abstract
Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their multi-step teachers in generation. However, standard evaluations such as GenEval2 typically draw only one sample per prompt, so improved scores may fail to reveal losses in distribution coverage. We therefore revisit whether distilled models truly match their teachers beyond single-draw performance using \textbf{pass@$\mathbf{k}$}, which measures the probability that at least one of $k$ independent samples satisfies a quality criterion. At $k{=}1$, pass@$k$ reduces to standard single-draw evaluation. As $k$ grows, the curve reveals whether additional draws find genuinely different successes or merely revisit the same modes, directly exposing how broadly a model covers the space of valid outputs. We first show that classifier-free guidance (CFG), whose quality--coverage tradeoff is well established, is the clearest case: higher guidance improves pass@$1$, but its advantage shrinks and reverses at larger $k$. Applying pass@$k$ to few-step distilled models, we find the same tradeoff splits along training objectives: distribution-matching objectives concentrate the student's output distribution, boosting early-hit rates while eroding large-budget coverage, whereas consistency and trajectory-based objectives better preserve the teacher's coverage even at large $k$. We further show that this tradeoff extends to few-step causal video generation. Our findings reveal a previously overlooked cost of diffusion distillation: across both image and video generation, the choice of training objective fundamentally determines whether a few-step model inherits its teacher's distribution coverage or trades it away for single-draw quality.
Chinese Translation
扩散蒸馏被广泛采用以加速采样,而由此得到的少步模型被广泛认为在生成中匹敌甚至超越其多步教师模型。然而,诸如 GenEval2 的标准评估通常对每个提示仅抽取一个样本,因此提高的分数可能无法揭示分布覆盖上的损失。因此,我们使用 \textbf{pass@$\mathbf{k}$} 重新审视蒸馏模型是否在单次抽取性能之外真正匹敌其教师模型,该指标衡量 $k$ 个独立样本中至少有一个满足质量标准的概率。当 $k{=}1$ 时,pass@$k$ 退化为标准的单次抽取评估。随着 $k$ 增大,该曲线揭示额外抽取是找到真正不同的成功样本,还是仅仅重复访问相同的模式,从而直接暴露模型覆盖有效输出空间的范围有多广。我们首先表明,无分类器引导(CFG)——其质量—覆盖权衡已被充分确立——是最清晰的情形:更高的引导会提升 pass@$1$,但其优势在更大的 $k$ 下会缩小并逆转。将 pass@$k$ 应用于少步蒸馏模型时,我们发现同样的权衡会沿着训练目标而分化:分布匹配目标会集中学生的输出分布,提高早期命中率,同时侵蚀大预算覆盖,而一致性和基于轨迹的目标即使在大的 $k$ 下也能更好地保持教师模型的覆盖。我们进一步表明,这种权衡也延伸到少步因果视频生成。我们的发现揭示了扩散蒸馏此前被忽视的代价:在图像和视频生成中,训练目标的选择从根本上决定了一个少步模型是继承其教师模型的分布覆盖,还是为了单次抽取质量而将其舍弃。
cs.LG / 114 / 2609.38860
Optimal Design for Active Preference Learning with Biased LLM Judges
面向有偏 LLM 评判者的主动偏好学习的最优设计
large language model
大语言模型相关
Abstract
Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the acquisition design rather than relying on a separate calibration stage. Under joint estimation, comparisons that appear highly informative about the reward may also reflect judge bias and therefore provide less information about human preferences. To address this issue, we propose Nuisance-Adjusted Optimal Design (NAOD), a comparison-selection strategy that prioritizes policy-relevant target information after nuisance adjustment and uses the Frank-Wolfe algorithm for optimization. Theoretically, we establish a sharp conditional local asymptotic minimax lower bound on policy risk and construct an estimator that attains it. We further characterize the finite-sample cost of learning the nuisance representation and show that representation error can reverse an oracle design advantage. Finally, we validate these predictions experimentally and evaluate NAOD on Chatbot Arena data across 17 judges, 15 budget configurations, and 15 random cluster-level splits. NAOD reduces the mean regret of proxy policy by 29.1% relative to a matched target-information design, outperforms existing methods, and improves human-preference prediction on held-out data.
Chinese Translation
从人类偏好中学习是大型语言模型(LLM)对齐的核心,但人类偏好标注成本高昂。主动偏好学习通过选择有信息量的比较来降低这一成本,而 LLM 评判者可以提供额外的可扩展反馈。然而,评判者的偏好可能偏离目标人类群体的偏好。即使在可信参考数据上进行校准之后,主动获取也可能改变比较分布,并暴露出残留的评判者偏差。因此,我们将评判者偏差纳入获取设计,而不是依赖单独的校准阶段。在联合估计下,那些看起来对奖励极具信息量的比较也可能反映评判者偏差,因此提供关于人类偏好的信息较少。为了解决这个问题,我们提出 Nuisance-Adjusted Optimal Design (NAOD),一种在讨厌参数调整后优先考虑与策略相关的目标信息的比较选择策略,并使用 Frank-Wolfe 算法进行优化。在理论上,我们建立了策略风险的尖锐条件局部渐近极小极大下界,并构造了一个达到该下界的估计量。我们进一步刻画了学习讨厌参数表示的有限样本成本,并表明表示误差可能逆转 oracle 设计的优势。最后,我们通过实验验证了这些预测,并在 Chatbot Arena 数据上跨 17 个评判者、15 种预算配置和 15 个随机簇级划分评估了 NAOD。相对于匹配的目标信息设计,NAOD 将代理策略的平均遗憾降低了 29.1%,优于现有方法,并改善了在留出数据上的人类偏好预测。
cs.LG / 115 / 2609.38879
Does Learning Protein Folding Generalize to Broader Reasoning?
学习蛋白质折叠能否泛化到更广泛的推理?
large language model
大语言模型相关
Abstract
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
Chinese Translation
大型语言模型严重依赖人类文本,而人类文本往往传达的是表面答案,而不是其背后的空间与结构逻辑。蛋白质折叠是一个天然的试验平台,因为一个已解析的结构能够产生成千上万条可精确检验的空间和拓扑陈述。我们提出疑问:学习折叠蛋白质能否教会通用模型可复用的推理能力?为回答这一问题,我们构建了 FoldingCorpus,一个由蛋白质衍生的问答数据集,以及 Fold2Reason,一种通过两种互补信号在其上进行后训练的方法:经由模型原生语言头预测的离散结构答案,以及从相同共享表示中解码出的连续 3D 几何。在 FoldBench 上,Fold2Reason 取得的结构预测分数是 Qwen3.5-9B 的 2.7 到 3.5 倍。除了蛋白质结构预测之外,它还在涵盖空间、图、科学和通用推理的全部 10 个基准上提升了性能,将宏平均准确率从 45.09% 提高到 48.33%(+3.23 个百分点),在全部 10 个基准上均有正向增益,而由随机、合成和打乱结构构建的匹配对照则产生显著更小或负向的增益。我们的工作表明,非语言的、结构密集的科学数据能够系统性地改进语言模型中的广泛推理,使一个已解决的科学问题成为后训练监督的一个实用来源。
cs.LG / 116 / 2609.38884
Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling
正确答案,高成本模型:基于 LLM 的优化建模中的效率差距
large language model
大语言模型相关
Abstract
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57\% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.
Chinese Translation
优化建模将现实世界的决策问题表述为求解器可用于寻找最优决策的数学规划。大语言模型(LLMs)可以自动化这一过程,但由此得到的正确建模表述在构建和求解时可能需要大量时间和内存,限制了实际可扩展性。因此,我们系统性地研究 LLM 能否从自然语言描述中识别问题结构,并应用合适的优化建模技术来生成能够正确且高效地求解这些问题的数学模型和求解器代码。为此,我们首先构建 OptTips,这是一个包含八个类别中 50 种专家建模技术的知识库。利用这些知识,我们开发了 OptDachshund,一个多智能体框架,将来自现有优化基准的问题转化为用于评估 LLM 对建模技术使用的新任务。它针对相同任务和数据构建传统数学模型和专家数学模型及求解器代码,为正确性和计算成本提供基线。由此产生的 EfficientOpt 基准包含 561 个经过专家评审的任务,并带有成对的参考实现。对 11 个代表性 LLM 的评估显示,在测量结果可比的正确求解任务上存在效率差距:对于每个 LLM,大多数生成程序比对应的专家程序需要更长的求解时间。在参考规模可比的子集中,57\% 的具有正确目标值且变量和线性约束更少的程序具有更长的记录求解器时间。案例研究表明,不同的建模技术可以在相似的记录成本下达到相同的最优值。如果代码需要更长时间来准备数据和构建模型,更快的求解可能不会减少执行时间。因此,LLM 优化建模应同时评估正确性和计算效率。
cs.LG / 117 / 2609.38909
Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
使用对比式遗忘集遗忘大型语言模型中的欺骗行为
large language model
大语言模型相关
Abstract
Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit face a dilemma. Suppression objectives such as NPO leave much of the deception in place. Target-based objectives, which distill the model's neutral behavior into the pressured context, remove it but induce context blindness: a target generated without the context teaches the model to stop reading it, eroding benign system-prompt instructions, secret-keeping and the reasoning a monitor inspects, a failure invisible to deception rates and capability benchmarks. We introduce PACT, which trains toward pressure-aware counterfactual targets (the model's own honest response, with a trace that registers the pressure and resists it) while retaining the benign uses of the triggering context. On two 32B reasoning models, PACT reduces held-out deception from over 50% to under 3% while system-prompt adherence, secret-keeping and the reasoning trace stay at the base model's level. On a tug-of-war score of removal against retention, PACT reaches 0.94 and 0.86, against at most 0.77 and 0.60 for any baseline. Like removed knowledge, removed deception is shallow under relearning, and terms that simulate the attacker hold it only at a cost in context use.
Chinese Translation
大型语言模型常常知道真相却说出别的话:一个在中性提问时能正确作答的模型,一旦上下文给予奖励,就会认可用户的错误信念,或错报其系统提示想要隐藏的事实。这类欺骗是一种以上下文为条件的行为,而非知识,然而机器遗忘——从权重中移除某种行为的天然工具——是为遗忘事实而构建的,而欺骗性模型仍然需要这些事实。我们提出遗忘模型何时欺骗,而非它知道什么,其对比式遗忘单元由模型自身实际发生的欺骗构建:同一个问题分别置于触发欺骗的上下文和中性的上下文之下,仅当信念成立且行为发生翻转时才被接纳。在该单元上,标准目标面临一个两难困境。诸如 NPO 之类的抑制型目标会使大部分欺骗行为原封不动地保留下来。基于目标的目标函数将模型的中性行为蒸馏到受压力的上下文中,能够移除欺骗,但会引发上下文盲视:在不含该上下文的情况下生成的目标会教模型停止阅读上下文,从而侵蚀良性的系统提示指令、保密能力以及监控器所检查的推理过程,这是一种对欺骗率和能力基准测试都不可见的失败。我们提出 PACT,它朝压力感知的反事实目标进行训练(即模型自身诚实的回答,并带有一条能记录该压力并予以抵制的轨迹),同时保留触发上下文的有益用途。在两个 32B 推理模型上,PACT 将留出集上的欺骗率从超过 50% 降至 3% 以下,同时系统提示遵循度、保密能力和推理轨迹均保持在基座模型的水平。在“移除”对抗“保留”的拔河评分上,PACT 分别达到 0.94 和 0.86,而任何基线方法至多为 0.77 和 0.60。与被移除的知识一样,被移除的欺骗在重新学习下也是浅层的,而模拟攻击者的项只能以牺牲上下文使用为代价来维持它。
cs.LG / 118 / 2609.38920
Learning Where to Steer: Noise-Space Geometry for Efficient Offline Multi-Objective Optimization with Generative Models
学习在何处引导:面向高效离线多目标优化的生成模型噪声空间几何
diffusion
扩散模型相关
Abstract
Offline multi-objective optimization (MOO) seeks solutions with better objective trade-offs using only a fixed dataset, without querying the objectives. Diffusion models trained on such data have emerged as a promising approach, but their samples are not inherently better than the data and must be steered toward the Pareto front. Existing methods guide or condition every sampling step. We instead act on the initial noise and leave the sampling process unchanged. Across Off-MOO-Bench, we observe that the objectives, as functions of the noise, are sensitive to only a few directions. We estimate these directions once per task via a Recursive Feature Machine using function values alone, and a small cache serves every trade-off, so each candidate costs one noise displacement and one ODE solve. We prove that this displacement increases the learned scalarized objective in expectation, and that sweeping trade-offs recovers the flow's attainable front up to proxy and steering errors. With additional guidance, for which we introduce novel data-adaptive and Pareto-aware operators, our method attains the best average hypervolume rank among generative methods on 47 tasks, at comparable or lower sampling cost. Steering alone outranks the best prior generative method at a fraction of its sampling cost.
Chinese Translation
离线多目标优化(MOO)仅使用固定数据集、不查询目标函数,寻求具有更优目标权衡的解。在此类数据上训练的扩散模型已成为一种有前景的方法,但其样本本身并不优于数据,必须被引导至 Pareto 前沿。现有方法在每一步采样中都进行引导或条件化。我们则作用于初始噪声,并保持采样过程不变。在 Off-MOO-Bench 上,我们观察到目标作为噪声的函数仅对少数方向敏感。我们仅使用函数值,通过递归特征机(Recursive Feature Machine)为每个任务估计一次这些方向,且一个小型缓存即可服务每一种权衡,因此每个候选解只需一次噪声位移和一次 ODE 求解。我们证明,该位移在期望上提升所学得的标量化目标,并且扫描权衡能够恢复该流的可达前沿,误差仅受代理与引导误差限制。在额外引导下(为此我们引入了新颖的数据自适应与 Pareto 感知算子),我们的方法在 47 个任务上取得生成式方法中最佳的平均超体积排名,且采样成本相当或更低。仅引导本身就能以该最佳先前生成式方法采样成本的一小部分超越它。
cs.LG / 119 / 2609.38955
Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching
基于Q分数匹配的序列化价值恢复的无循环逆强化学习
diffusion
扩散模型相关
Abstract
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
Chinese Translation
逆强化学习(IRL)旨在恢复一个能够解释专家演示的奖励函数。现有的IRL方法通常依赖于在奖励学习与策略优化之间交替进行的双层优化过程,这带来了巨大的计算负担和训练不稳定性。在本工作中,我们引入了一条不同的路径,通过利用扩散策略完全消除了策略优化。我们的关键洞见是,扩散策略编码了最优soft Q函数的动作梯度结构,使得奖励学习可以被转化为一系列价值恢复问题,从而让我们能够绕开先前IRL方法中固有的奖励-策略循环。具体而言,我们的方法分三个阶段进行:(I)通过动作梯度匹配恢复最优soft Q函数,并以受Gumbel回归启发的方式估计相应的soft价值函数(Q值的LogSumExp);(II)通过推断一个状态相关的偏移量来校准这些soft价值;(III)通过施加贝尔曼一致性来提取奖励。这引出了无循环逆强化学习(LFIRL),一种完全离线的算法,以简单、无循环且序列化的方式运行。LFIRL实现简单,在保持强劲奖励恢复性能的同时显著提升了训练效率。在实证上,在Maze、Franka Kitchen、Adroit Hand Pen和Push-T基准上,LFIRL相比最快的基线实现了2-3倍的加速,同时在奖励恢复质量上匹配或超越了最先进的方法。
cs.LG / 120 / 2609.39019
Synchronous Multi-view Neural Diffusion
同步多视图神经扩散
diffusion
扩散模型相关
Abstract
Multi-view learning seeks to learn more comprehensive representations by exploiting the complementarity and consistency across diverse modalities or views. However, existing multi-view fusion strategies treat intra- and inter-view fusion as independent stages, without simultaneously considering the evolution within views and the dependency across views. Such an asynchronous fusion paradigm inevitably constrains cross-view interactions due to conflicting view-specific structural inductive biases. As a result, information flow is prone to distortion and compression along intermediate pathways, confining the model to learn within a restricted solution space. To address this, we propose Synchronous Multi-view Neural Diffusion (SynMDiff), which conceptualizes the multi-view feature space as a unified dynamical system driven by a diffusion process. By modeling the diffusion flow across arbitrary dyadic feature interactions in a joint space, SynMDiff enables the concurrent and adaptive intra- and inter-view information fusion. While a direct implementation of this synchronized mechanism incurs prohibitive computational costs, we further introduce an energy-based topological sampling strategy and an Ego-Net style centralized training architecture, ensuring both efficiency and scalability during learning and inference. Due to its conceptual elegance and computational efficacy, evaluations on real-world datasets demonstrate that SynMDiff outperforms the baselines by a large margin.
Chinese Translation
多视图学习旨在通过利用不同模态或视图之间的互补性和一致性来学习更全面的表示。然而,现有的多视图融合策略将视图内融合和视图间融合视为相互独立的阶段,而没有同时考虑视图内的演化和视图间的依赖关系。这种异步融合范式由于相互冲突的视图特定结构归纳偏置,不可避免地限制了跨视图交互。因此,信息流在中间路径上容易发生失真和压缩,从而将模型限制在受限的解空间内进行学习。为了解决这一问题,我们提出了同步多视图神经扩散(SynMDiff),它将多视图特征空间概念化为一个由扩散过程驱动的统一动力系统。通过在联合空间中对任意二元特征交互之间的扩散流进行建模,SynMDiff 能够实现并发的、自适应的视图内和视图间信息融合。虽然直接实现这种同步机制会带来高昂到难以承受的计算成本,但我们进一步引入了基于能量的拓扑采样策略和 Ego-Net 风格的集中式训练架构,从而确保学习和推理过程中的效率与可扩展性。由于其概念上的优雅性和计算上的高效性,在真实世界数据集上的评估表明,SynMDiff 大幅优于基线方法。
cs.LG / 121 / 2609.39114
The Row Normalization Puzzle in Muon
Muon 中的行归一化谜题
large language model
大语言模型相关
Abstract
This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performance (Li et al.). Despite its growing adoption and promising performance in large language model (LLM) pretraining, NorMuon's worst-case guarantees remain poorly understood. One fundamental question is: Does row normalization yield provable convergence gains, potentially through its interaction with approximate polar computation and exponential moving-average momentum? Our results show that row normalization introduces a dimension-dependent factor in the worst-case iteration complexity under the operator-norm geometry, which persists even with exact polar computation and any fixed momentum parameters. Indeed, we establish an algorithm-dependent lower bound and a matching upper bound in deterministic settings, and extend our upper bound analysis to stochastic settings. Both upper-bound analyses allow approximate polar computation. Experiments show that NorMuon is slower than Muon on synthetic problems inspired by our worst-case construction, yet outperforms Muon in LLM pretraining. These findings sharpen the puzzle of why row normalization helps in practice and complement the recent findings of Dewulf et al.
Chinese Translation
本文考察逐行重归一化如何影响 Muon,重点关注 NorMuon 的最坏情况保证与其实际性能之间的差距(Li 等人)。尽管 NorMuon 在大语言模型(LLM)预训练中日益被采用并展现出有前景的性能,但其最坏情况保证仍鲜为人知。一个基本问题是:行归一化是否会带来可证明的收敛增益,并且这种增益可能通过其与近似极分解计算和指数移动平均动量的相互作用而产生?我们的结果表明,在算子范数几何下,行归一化在最坏情况迭代复杂度中引入了一个依赖于维度的因子,即使采用精确极分解计算和任意固定动量参数,该因子依然存在。事实上,我们在确定性设置下建立了一个依赖于算法的下界和一个相匹配的上界,并将我们的上界分析扩展到随机设置。两项上界分析均允许近似极分解计算。实验表明,在受我们最坏情况构造启发的合成问题上,NorMuon 比 Muon 更慢,但在 LLM 预训练中优于 Muon。这些发现使“行归一化为何在实践中有所帮助”这一谜题更加尖锐,并补充了 Dewulf 等人最近的发现。
cs.LG / 122 / 2609.39124
CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data
CDMD:一种用于表格数据的跨数据集混合类型扩散模型
diffusion
扩散模型相关
Abstract
Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.
Chinese Translation
表格数据的生成模型通常针对每个数据集单独训练,这限制了知识迁移,并且需要存储许多专用模型。在本文中,我们提出 CDMD,一种跨异构数据集联合训练的表格扩散模型,这些数据集具有不同的模式以及数量可变的数值特征和类别特征。与在连续表示空间中运行的现有跨数据集表格扩散模型不同,CDMD 直接在混合类型特征空间上定义扩散,并进行端到端训练。为了容纳异构的类别域,我们为掩码扩散模型引入了一种模式受限的逆向过程参数化,其中输出空间动态地适应每个特征的词汇表。然后,我们将数值和类别特征级扩散过程组合成一个依赖模式的行级过程。一个共享的模式感知 Transformer 去噪器捕获特征之间的依赖关系,并在不同模式上参数化逆向过程。在七个真实世界数据集上,单个联合训练的 CDMD 在强大的单数据集和跨数据集基线中取得了最高的平均生成质量,同时其总参数量显著少于分别训练的模型集合。此外,在一个包含 337 个数据集的语料库上进行预训练,在有限的目标数据和有限的适应轮次下,均能改善在先前未见数据集上的生成。这些结果证明了直接混合类型扩散在共享且可迁移的表格数据生成方面的潜力。我们的代码可在 https://github.com/ketatam/cdmd 获取。
cs.LG / 123 / 2609.39131
Characterizing High Bandwidth Flash for LLM Serving
用于LLM服务的高带宽闪存特征刻画
large language model
大语言模型相关
Abstract
Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.0% relative to HBM-only systems. Modeled energy savings reach 55.8%, although HBF increases energy consumption on some light workloads. Buffered cache-aware scheduling extends estimated HBF write lifetime from 4.77 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.
Chinese Translation
大语言模型(LLM)服务需要大量内存来存储模型权重和KV缓存。随着模型变得更大且上下文变得更长,内存容量和带宽日益成为服务性能的瓶颈。智能体工作负载通过不断增长的上下文上的重复交互加剧了这一压力,使得保留KV状态以供复用变得越来越重要。高带宽闪存(HBF)为扩展用于大语言模型(LLM)服务的加速器内存容量提供了一种途径,但其访问成本和有限的写入耐久性使其使用复杂化。我们针对高吞吐量智能体服务评估HBF,涵盖系统设计和调度选择,以理解额外容量何时能够提升服务性能和能效。我们引入一种HBM-HBF-主机分层存储系统和带缓冲的缓存感知调度,并使用跟踪驱动的模拟来分析它们对性能、能耗和HBF写入寿命的影响。在评估的工作负载中,最快的HBF增强系统相对于仅HBM系统将完成时间减少了36.1-87.0%。建模的节能达到55.8%,尽管HBF在某些轻量工作负载上会增加能耗。在评估配置中,带缓冲的缓存感知调度将估计的HBF写入寿命从4.77年延长至14.82年。这些结果证明了协调数据放置和调度以提高服务效率、同时维持实际可用的HBF写入寿命的重要性。
cs.LG / 124 / 2609.39137
ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control
ID Balancing:通过基于 PID 的负载控制实现极端稀疏 MoE 的稳定训练
large language model
大语言模型相关
Abstract
Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization and training efficiency, and can undermine training stability, becoming a bottleneck to reliable scaling. In this work, we unify two representative auxiliary-loss-free methods as incomplete Proportional-Integral-Derivative (PID) controllers: DeepSeek's loss-free method acts as a fixed-step integral controller, while Kimi K3's Quantile Balancing functions as a generalized proportional controller. Building on this control perspective, we propose ID Balancing, an Integral-Derivative controller. It scales its integral term with load error and activates its derivative term only when imbalance worsens, enabling stronger corrections for large or worsening errors and smaller updates near balance. Evaluated across Top-$10$, Top-$5$, and Top-$3$ routing over $768$ experts, ID Balancing reduces worst-case backbone MaxVio and training-average backbone MinVio by over $50\%$ and $12\%$, respectively, relative to the best baselines in the Top-$3$ setting. When the total parameter count increases from $18.9$B to $69.9$B (Top-$10$-of-$768$), ID Balancing's worst-case backbone MaxVio remains nearly unchanged and is approximately $89.6\%$ lower than that of the auxiliary-loss baseline. ID Balancing also maintains competitive language-modeling and downstream performance. The advantages of ID Balancing grow as sparsity increases, making it a promising solution for scaling larger, sparser MoE models.
Chinese Translation
通过混合专家(MoE)扩展大语言模型(LLMs)能够以近乎恒定的每 token 计算量实现参数量的大规模增长。然而,进一步扩展参数量需要越来越稀疏的路由,其中专家负载不均衡会变得更加严重。这种不均衡会降低参数利用率和训练效率,并可能削弱训练稳定性,成为可靠扩展的瓶颈。在这项工作中,我们将两种具有代表性的无辅助损失方法统一为不完整的比例-积分-微分(PID)控制器:DeepSeek 的无损失方法充当固定步长积分控制器,而 Kimi K3 的 Quantile Balancing 则充当广义比例控制器。基于这一控制视角,我们提出了 ID Balancing,一种积分-微分控制器。它使其积分项随负载误差缩放,并且仅在不均衡恶化时激活其微分项,从而能够对较大或不断恶化的误差进行更强的校正,并在接近平衡时进行更小的更新。在 $768$ 个专家上的 Top-$10$、Top-$5$ 和 Top-$3$ 路由中进行评估,在 Top-$3$ 设置下,相对于最佳基线,ID Balancing 将最坏情况 backbone MaxVio 和训练平均 backbone MinVio 分别降低超过 $50\%$ 和 $12\%$。当总参数量从 $18.9$B 增加到 $69.9$B(Top-$10$-of-$768$)时,ID Balancing 的最坏情况 backbone MaxVio 几乎保持不变,并且比辅助损失基线的相应值低约 $89.6\%$。ID Balancing 还保持了具有竞争力的语言建模和下游性能。随着稀疏度增加,ID Balancing 的优势也会增长,使其成为扩展更大、更稀疏 MoE 模型的有前景的解决方案。
cs.LG / 125 / 2609.39223
QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs
QATFactory:一个用于大语言模型量化感知训练与蒸馏的多功能、面向部署的框架
large language model
大语言模型相关
Abstract
Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) and reinforcement learning (QARL). QATFactory simulates deployment-time quantization while performing matrix multiplications in BF16, allowing models to adapt to quantization noise without requiring training hardware that natively supports the target format; for example, it supports NVFP4 training on H100 GPUs, which lack FP4 Tensor Cores. The framework supports NVFP4, MXFP4, and llama.cpp's Q4_K format; dense and mixture-of-experts models; and both full-parameter and LoRA-based training. It exports checkpoints directly to vLLM and llama.cpp without an additional lossy conversion step or added inference overhead. With QATFactory, we conduct extensive experiments on models ranging from 8B to 230B parameters and evaluate exported checkpoints in production inference engines. Across models and formats, QAD consistently improves deployed-model quality over strong PTQ baselines. On Qwen3.5-9B, QAD achieves average benchmark accuracies of 68.9% under NVFP4 and 66.0% under MXFP4, outperforming the best PTQ results of 65.4% and 56.4%, respectively. Through our experiments, we found that although both FP4 formats quantize weights and activations at deployment, the best training strategy is format-dependent: NVFP4 generally performs better when only weights are quantized during training, whereas MXFP4 benefits from quantizing both weights and activations. At a fixed training token budget, training on fewer 32K sequences improves average accuracy by 1.9 points over training on more 4K sequences. We release the complete QATFactory training code and the resulting checkpoints.
Chinese Translation
大语言模型(LLM)推理正日益转向更低精度,以充分发挥硬件加速器的吞吐量,但激进的后训练量化(PTQ)会降低模型质量。我们提出 QATFactory,一个面向部署对齐的量化感知蒸馏(QAD)和强化学习(QARL)的开源框架。QATFactory 在模拟部署时量化的同时以 BF16 执行矩阵乘法,使模型能够适应量化噪声,而无需原生支持目标格式的训练硬件;例如,它支持在缺少 FP4 Tensor Core 的 H100 GPU 上进行 NVFP4 训练。该框架支持 NVFP4、MXFP4 以及 llama.cpp 的 Q4_K 格式;支持稠密模型和专家混合模型;并支持全参数训练和基于 LoRA 的训练。它可将检查点直接导出到 vLLM 和 llama.cpp,而无需额外的有损转换步骤,也不会增加推理开销。借助 QATFactory,我们在参数量从 8B 到 230B 的模型上进行了大量实验,并在生产推理引擎中评估导出的检查点。在不同模型和格式上,QAD 相较于强 PTQ 基线始终能提升部署模型的质量。在 Qwen3.5-9B 上,QAD 在 NVFP4 下达到 68.9% 的平均基准准确率,在 MXFP4 下达到 66.0%,分别优于最佳 PTQ 结果的 65.4% 和 56.4%。通过实验,我们发现,尽管两种 FP4 格式在部署时都会量化权重和激活,但最佳训练策略取决于格式:当训练期间仅量化权重时,NVFP4 通常表现更好,而 MXFP4 则受益于同时量化权重和激活。在固定训练 token 预算下,与在更多 4K 序列上训练相比,在更少 32K 序列上训练可将平均准确率提高 1.9 个百分点。我们发布了完整的 QATFactory 训练代码及所得检查点。
cs.LG / 126 / 2609.39271
RW-Flow: One-Step Generation on Compact Manifolds via Wasserstein Gradient Flows
RW-Flow:通过 Wasserstein 梯度流在紧流形上进行单步生成
diffusion
扩散模型相关
Abstract
Manifold-valued data, and consequently the distributions they induce, are prevalent across many domains, ranging from the locations of geospatial events, such as earthquakes, to biomolecular torsion angles that encode information about three-dimensional structure. While diffusion and flow-based generative models have been successfully extended to compact manifolds, sampling typically requires tens or hundreds of sequential network evaluations. We introduce RW-Flow, a theoretically grounded framework for learning one-step generative models on compact manifolds via Wasserstein gradient flows. The main challenge is identifiability: driving the velocity field to zero should guarantee that the model distribution matches the target distribution. We establish a necessary and sufficient condition for identifiability on compact, connected Riemannian manifolds. We specifically show that, for a symmetric, Lipschitz-continuous cost function, the velocity field induced by the Sinkhorn divergence is identifiable if and only if the associated Gibbs kernel is nondegenerate. This characterization provides a general principle for designing identifiable costs on compact manifolds. It also reveals that the squared geodesic distance, the natural manifold analogue of the squared Euclidean distance, does not always guarantee identifiability. Across benchmarks involving geospatial events, protein side chain torsion angles, RNA backbone torsion angles, and general manifolds discretized as triangular meshes, RW-Flow outperforms existing one-step methods in nearly all settings under fair comparison conditions.
Chinese Translation
流形值数据,以及因此由它们诱导的分布,在许多领域中都很普遍,其范围从地理空间事件(例如地震)的位置,到编码三维结构信息的生物分子扭转角。尽管扩散模型和基于流的生成模型已被成功扩展到紧流形,但采样通常需要数十或数百次顺序网络评估。我们提出了 RW-Flow,一个具有理论基础的框架,用于通过 Wasserstein 梯度流在紧流形上学习单步生成模型。主要挑战是可辨识性:将速度场驱动到零应能保证模型分布与目标分布相匹配。我们建立了紧致连通黎曼流形上可辨识性的充分必要条件。我们特别表明,对于对称、Lipschitz 连续的代价函数,由 Sinkhorn 散度诱导的速度场可辨识当且仅当关联的 Gibbs 核非退化。这一刻画为在紧流形上设计可辨识的代价函数提供了一个一般性原则。它还揭示出,平方测地距离——平方欧几里得距离在流形上的自然类比——并不总能保证可辨识性。在涉及地理空间事件、蛋白质侧链扭转角、RNA 骨架扭转角以及被离散化为三角网格的一般流形的基准测试中,在公平比较条件下,RW-Flow 在几乎所有设置中都优于现有单步方法。
cs.LG / 127 / 2609.39329
PatchKV: Weight-Space Compensation of KV Cache
PatchKV:KV 缓存的权重空间补偿
large language model
大语言模型相关
Abstract
Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170K tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.
Chinese Translation
大语言模型(LLM)的长上下文推理受到键值(KV)缓存内存线性增长的瓶颈制约。现有的压缩方法通过令牌驱逐或近似来缩减缓存,但在激进的压缩预算下性能急剧下降。我们提出 PatchKV,一个免训练的框架,它通过将部分上下文承载于模型权重之中,来对 KV 缓存压缩方法进行补偿。PatchKV 将现成的压缩 KV 缓存与一个特定于上下文的权重补丁相配对,该补丁在上下文加载时一次性计算完成,并在该上下文的下游查询中持续复用。该权重补丁通过岭回归以闭式解导出,其做法是在完整缓存与压缩缓存下,对齐由上下文导出的参考查询令牌的逐块激活。一旦合并到模型中,该补丁便脱离前向计算图,并在单上下文、多查询的设置下使每次查询的推理开销保持不变。在三种模型架构上的长上下文问答(SCBench,最多 170K 令牌、SQuAD、NIAH)与数学(GSM8K)基准测试中,PatchKV 一致地提升了缓存压缩方法的效果,这暗示了在激进预算下对它们进行补偿的另一条方向。
cs.LG / 128 / 2609.39383
From Search to Signal: Online Post-Training in Automatic Heuristic Design
从搜索到信号:自动启发式设计中的在线后训练
large language model
大语言模型相关
Abstract
Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting them into learning signals must account for the prompt and evolving search state that produced each candidate. We formulate online post-training of small open-weight LLMs in AHD as context-dependent signal construction and develop alternative mappings from program validity, task performance, and generation context to update signals. Using shared evaluated rollouts and matched update budgets, controlled experiments across AHD tasks and model families compare these mappings with online post-training baselines, testing their effects on validity, performance among valid proposals, and the yield of valid proposals that improve under contextual comparisons. Complementary checkpoint, frozen-search, and live-system evaluations assess whether proposal-level gains appear in updated checkpoint behavior and subsequent search, rather than arising solely from accumulated search state. A resource-matched comparison under pre-specified cost accounting tests whether online updating adds value beyond additional search with a frozen generator. Together, this design avoids treating end-to-end search gains alone as evidence of stronger heuristic-design capabilities.
Chinese Translation
基于大语言模型(LLM)的自动启发式设计(AHD)迭代地提出并精化启发式,将设计理据与可执行代码配对。任务特定的评估器对程序进行评估;执行结果与性能分数引导搜索。许多 AHD 系统保持生成器冻结;EvoTune 与算法和语言模型协同演化(CALM)则从被评估的候选方案更新生成器。当此类结果驱动带可验证奖励的强化学习(RLVR)时,它们会形成一个与搜索耦合的循环:被评估的候选流既提供搜索状态更新,也为生成未来候选的模型提供训练信号。然而,有效性与性能并不能唯一地决定有用的模型更新;将它们转化为学习信号必须考虑产生每个候选的提示与不断演化的搜索状态。我们将 AHD 中小型开放权重 LLM 的在线后训练表述为依赖于上下文的信号构建,并开发了从程序有效性、任务性能与生成上下文到更新信号的多种替代映射。使用共享的已评估 rollout 与匹配的更新预算,跨 AHD 任务与模型族的受控实验将这些映射与在线后训练基线进行比较,检验它们对有效性、有效提案中的性能、以及在上下文比较下有所改进的有效提案产出率的影响。互补的检查点、冻结搜索与实时系统评估,考察提案层面的增益是否体现在更新后的检查点行为与后续搜索中,而非仅仅来自累积的搜索状态。在预先设定的成本核算下进行的资源匹配比较,检验在线更新是否在冻结生成器的额外搜索之外带来附加价值。总体而言,这一设计避免仅将端到端搜索增益视为更强启发式设计能力的证据。
cs.LG / 129 / 2609.39437
Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning
面向开放权重LLM的基于网络的空间上下文检索:面向有依据地理推理的忠实性基准
large language model
大语言模型相关
Abstract
Large language models (LLMs) encode substantial latent geographic knowledge, yet they reason poorly over space and are unreliable when queried from coordinates alone. Useful behaviour emerges only when structured spatial context is supplied in the prompt. This raises a question geographic evaluation has left unexamined: once the right context is supplied, does the model reason from it, or override it with its own parametric recall? We take up this question with an open pipeline for network-based spatial context retriev-al. In it, the surroundings of a selected point are defined by the pedestrian street network, the area actually reachable on foot. Using only open data and open-weight models, the pipeline retrieves features from OpenStreetMap and the GHS-POP population grid, computes indicators over the network catchment in code, and injects them as a compact spatial brief. On this basis we build a faithfulness benchmark. It labels every claim a model makes by its source (grounded in the brief, or drawn from training knowledge) and its correctness, and it probes each case with a planted false premise that the brief refutes. We evaluate sixteen open-weight model configurations across three families (Qwen, Gemma and Llama, with Gemma in two generations), four size classes and, where available, both thinking and non-thinking modes, on three con-trasting cities, resampling every case over ten seeds. The results show that resistance to the planted premise varies more strongly by model family and generation than by scale, while brief-reading competence forms a partly separate dimension. These behaviours are not captured by conventional world-correctness scores or single-shot evaluation. We release the implementation, spatial briefs, model outputs, and claim-level labels as a reproducible workflow at github.com/perezjoan/NSCR-LLM.
Chinese Translation
大语言模型(LLM)编码了大量潜在的地理知识,但它们在空间上的推理能力很差,并且在仅根据坐标进行查询时并不可靠。只有当提示中提供结构化的空间上下文时,才会出现有用的行为。这提出了一个地理评估尚未考察的问题:一旦提供了正确的上下文,模型是依据它进行推理,还是用它自身的参数化回忆来覆盖它?我们以一个用于基于网络的空间上下文检索的开放流水线来研究这个问题。在其中,一个选定点的周边环境由步行街道网络定义,即实际可步行到达的区域。仅使用开放数据和开放权重模型,该流水线从OpenStreetMap和GHS-POP人口网格中检索特征,在代码中计算网络集水区上的指标,并将它们作为紧凑的空间简报注入。在此基础上,我们构建了一个忠实性基准。它对模型提出的每一个主张按照其来源(有依据于简报,或来自训练知识)及其正确性进行标注,并且用一个简报所驳斥的植入错误前提来探查每个案例。我们在三个对比鲜明的城市上评估了十六种开放权重模型配置,涵盖三个系列(Qwen、Gemma和Llama,其中Gemma有两个代际)、四个规模类别,并在可用时涵盖思考与非思考两种模式,并对每个案例在十个随机种子上进行重采样。结果表明,对植入前提的抵抗性随模型系列和代际的变化比随规模的变化更强,而简报阅读能力则构成一个部分独立的维度。这些行为无法被传统的世界正确性分数或单次评估所捕捉。我们发布实现、空间简报、模型输出以及主张级标签,作为一个可复现的工作流,位于 github.com/perezjoan/NSCR-LLM。
cs.LG / 130 / 2609.39560
Self-Repulsive Sampling for Diffusion Language Models
扩散语言模型的自排斥采样
diffusion
扩散模型相关
Abstract
Sampling several responses and voting over their answers can improve a language model's accuracy, but repeated answers limit the benefit of additional samples. Raising temperature increases diversity at a potential cost to per-sample accuracy. We introduce Self-Repulsion (SR), a sampler for masked diffusion language models that uses peer commitments to diversify the pool. At each penalized denoising step, each path lowers a token's logit according to how many peers have committed that token at the same position. Paths share a batched forward pass and then commit in sequence, so later paths observe choices made earlier in the same step. This coupling requires no training or additional forward or backward pass and can produce distinct paths even at temperature zero. When all paths commit a position together from identical logits, the update exactly maximizes total logit minus a convex duplication cost. On LLaDA-8B-Instruct with ten paths and 128 denoising steps, deterministic SR reaches 80.38% plurality accuracy on GSM8K, compared with 70.17% for the unpenalized greedy decoder. At temperature 0.6 and matched model-evaluation budgets, the count penalty improves over self-consistency by 2.06 percentage points in blocks of 32 and 14.50 under pure diffusion. Experiments on GSM8K, MATH and TruthfulQA show that voting gains arise mainly from higher coverage of correct answers, with gains that vary by benchmark and decoding regime.
Chinese Translation
采样多个回复并对其答案进行投票可以提高语言模型的准确率,但重复的答案限制了额外样本所带来的收益。提高温度会增加多样性,但可能以牺牲单样本准确率为代价。我们提出自排斥(Self-Repulsion,SR),一种用于掩码扩散语言模型的采样器,它利用同伴的已提交选择来使样本池多样化。在每个施加惩罚的去噪步骤中,每条路径根据有多少同伴在同一位置提交了该 token 来降低该 token 的 logit。各条路径共享一次批量前向传播,然后按顺序提交,因此靠后的路径能观察到同一步中较早做出的选择。这种耦合不需要训练,也不需要额外的前向或反向传播,并且即使在温度为零时也能产生彼此不同的路径。当所有路径从相同的 logit 出发共同提交某一位置时,该更新恰好最大化总 logit 减去一个凸的重复代价。在 LLaDA-8B-Instruct 上,使用十条路径和 128 个去噪步骤时,确定性的 SR 在 GSM8K 上达到 80.38% 的多数投票准确率,而未施加惩罚的贪心解码器为 70.17%。在温度为 0.6 且模型评估预算匹配的情况下,计数惩罚相较于自一致性在大小为 32 的块中提升了 2.06 个百分点,在纯扩散下提升了 14.50 个百分点。在 GSM8K、MATH 和 TruthfulQA 上的实验表明,投票带来的增益主要来自正确答案覆盖率的提高,且这些增益因基准和解码机制的不同而有所差异。
cs.LG / 131 / 2609.39626
Parameterization method of reservoir properties for ensemble-based data assimilation using intermediate latent space of StyleGAN
使用 StyleGAN 中间潜在空间的基于集合的数据同化储层属性参数化方法
diffusion
扩散模型相关
Abstract
Ensemble smoothers are the most successful and efficient techniques currently available for history matching. However, because these methods rely on Gaussian assumptions, their performance is severely degraded when the prior geology is described in terms of complex facies distributions (non-Gaussian). In this way, for these methods, we need to apply efficient parameterization techniques. Currently, the most efficient methods for performing parameterization are deep learning models. However, given the variety of existing deep learning models, studies have not identified which is most suitable for use with ensemble-based methods, although some important models had already been evaluated. Based on a recent literature review, the most promising models selected were VAE-GAN, Latent Diffusion, and StyleGAN models. As a novel aspect of this work, data assimilation with the second generation of StyleGAN (StyleGAN2) model was performed using the latent z-space and intermediate w-space, separately. They were applied in two 2D case studies: one categorical (three facies) and the other continuous. The results demonstrated that all three models are highly efficient, with the StyleGAN2 model standing out for generating samples with geological realism and achieving excellent data matching in the cases studied. Our findings show that performing data assimilation with StyleGAN2 using the intermediate space (w-space) yielded better results than the traditional application in the latent space (z-space). This is due to the fact that ESMDA uses linear updates and the w-space is much more linear and disentangled than the highly entangled z-space, thereby ensuring that the updated vectors remain close to realistic geological patterns. These results were validated using main geostatistical and history matching metrics.
Chinese Translation
集合平滑器是目前可用于历史拟合的最成功且最高效的技术。然而,由于这些方法依赖于高斯假设,当先验地质以复杂相分布(非高斯)来描述时,其性能会严重下降。这样一来,对于这些方法,我们需要应用有效的参数化技术。当前,用于执行参数化的最高效方法是深度学习模型。然而,鉴于现有深度学习模型的多样性,尽管一些重要模型已经被评估过,研究尚未确定哪一种最适合与基于集合的方法一起使用。基于最近的文献综述,所选出的最有前景的模型是 VAE-GAN、Latent Diffusion 和 StyleGAN 模型。作为本工作的一个新方面,分别使用潜在 z 空间和中间 w 空间对第二代 StyleGAN(StyleGAN2)模型进行了数据同化。它们被应用于两个二维案例研究:一个为分类型(三个相),另一个为连续型。结果表明,所有三个模型都非常高效,其中 StyleGAN2 模型脱颖而出,因为它能够生成具有地质真实感的样本,并在所研究的案例中实现出色的数据匹配。我们的研究结果表明,使用中间空间(w 空间)对 StyleGAN2 进行数据同化,比在潜在空间(z 空间)中的传统应用取得了更好的结果。这是因为 ESMDA 使用线性更新,而 w 空间比高度纠缠的 z 空间更加线性和解耦,从而确保更新后的向量保持接近真实的地质模式。这些结果使用主要的地质统计学和历史拟合指标进行了验证。
cs.LG / 132 / 2609.39628
MIND: Marginal-Invariant Neural Dependency Diffusion for Mixed-Type Tabular Generation
MIND:面向混合类型表格生成的边缘不变神经依赖扩散模型
diffusion
扩散模型相关
Abstract
This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.
Chinese Translation
本文提出 MIND,一种面向混合类型表格数据的边缘不变神经依赖扩散模型。MIND 并不直接在原始异构特征空间中学习联合分布。相反,它首先通过逐列的边缘传输将不同的变量类型映射到一个统一的潜在依赖空间。随后,一个条件扩散模型学习列与列之间的关系。Copula 切向去噪将已知的边缘成分与可学习的依赖残差分离开来。采样阶段的秩投影进一步缓解反向扩散中的边缘偏移。在九个多样化的表格基准上的实验表明,相较于现有的统一方法,MIND 在边缘保真度和依赖保持方面均能持续取得提升。通过将边缘建模与依赖学习显式隔离,MIND 在边缘保真度、联合依赖保持以及下游预测效用之间实现了强大而稳定的平衡。这项工作支持将边缘建模与依赖建模相分离,作为复杂混合类型表格生成的一种有原则且极为有效的范式。
cs.LG / 133 / 2609.39648
From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models
从众数到记忆:刻画扩散模型的尺度空间动力学
diffusion
扩散模型相关
Abstract
Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $σ$, we treat the denoiser as a self-map and study its dynamics. For an exact denoiser, fixed points correspond to critical points of the smoothed data density, while attractors correspond to its modes; as $σ$ increases, sample-level modes merge into progressively coarser ones. This suggests a geometric view of memorization: examples that receive excess probability mass due to duplication or overfitting, as well as outliers, should remain distinguishable under stronger smoothing than ordinary examples. We quantify this persistence by the critical scale $σ_c$, the largest noise scale at which an example is retained by the fixed-scale dynamics. In conditional models, the same construction extends naturally to image--caption pairs. Experiments in controlled settings and on large-scale models show that $σ_c$ tracks memorization arising from duplication, overfitting, and outliers, and identifies both memorized and partially memorized examples in Stable Diffusion. Moreover, $σ_c$ yields interpretable measures of the image spatial distribution and caption dependence of memorization.
Chinese Translation
扩散模型通常被视为将噪声转换为数据的随机过程。我们采取一个互补的视角:一个扩散模型定义了一族以噪声尺度为索引的确定性动力系统。在每个固定尺度 $σ$ 下,我们将去噪器视为一个自映射并研究其动力学。对于精确的去噪器,不动点对应于平滑后数据密度的临界点,而吸引子对应于其众数;随着 $σ$ 增大,样本级众数合并为逐渐更粗的众数。这提出了一种关于记忆的几何视角:由于重复或过拟合而获得过多概率质量的样本,以及异常值,在比普通样本更强的平滑下仍应保持可区分。我们通过临界尺度 $σ_c$ 来量化这种持续性,$σ_c$ 是一个样本在固定尺度动力学下得以保留的最大噪声尺度。在条件模型中,同样的构造自然扩展到图像--说明对。在受控设置和大规模模型上的实验表明,$σ_c$ 追踪由重复、过拟合和异常值引起的记忆,并在 Stable Diffusion 中识别出完全记忆和部分记忆的样本。此外,$σ_c$ 产生了对记忆的图像空间分布和说明依赖性的可解释度量。
cs.LG / 134 / 2609.39658
Graph Residual Conjugate Diffusion: SNR-Equalized Heat Flow for Graph Signals
图残差共轭扩散:面向图信号的 SNR 均衡热流
diffusion
扩散模型相关
Abstract
Diffusion models generate data by reversing a forward corruption process that typically approaches a simple Gaussian prior. Recent work has extended this framework to signals supported on fixed graphs, e.g., road-network traffic and sensor-network measurements. Many graph signals have nonuniform spectral energy, whereas isotropic corruption adds the same conditional noise variance to every graph-frequency mode. Driving all modes to near-zero terminal signal-to-noise ratio (SNR) requires strong corruption, which increases the noise range that must be covered under a fixed sampling budget. We introduce Graph Residual Conjugate Diffusion (GRCD), which replaces the shared clock of graph heat diffusion with a mode-dependent clock that gives every graph-Fourier mode the same conditional SNR. GRCD fits a zero-mean graph-spectral Gaussian reference on the training split and stops at a finite terminal SNR at which the propagated reference still carries the fitted spectral variances. The Gaussian component has an exact modewise propagator in the probability-flow ODE, so sampling advances it analytically and integrates only the learned residual score numerically. We evaluate GRCD on five settings (METR-LA traffic, Molene weather, and three stochastic block models) against seven comparators under a matched protocol: Graph-Aware Diffusion (GAD), EDM (graph backbone), two adaptations of Whitened Score Diffusion (WSD), and three preconditioning controls. At four function evaluations (NFEs), GRCD lowers averaged maximum mean discrepancy (aMMD) by 22 to 36 times over the best comparator on all five settings, reaching 0.054 on METR-LA, where it clears an aMMD 0.1 target with 87% less sampling wall-clock time than the cheapest comparator that reaches it. Fitting the terminal reference reduces aMMD by 2.7 to 7.3 times at finite terminal SNR, while the factors shrink to 1.00 to 1.01 near zero.
Chinese Translation
扩散模型通过逆转一个通常趋近于简单高斯先验的前向破坏过程来生成数据。近期工作已将这一框架扩展到支撑在固定图上的信号,例如道路网络交通流和传感器网络测量。许多图信号具有非均匀的谱能量,而各向同性破坏会给每个图频率模态加入相同的条件噪声方差。将所有模态驱动到接近零的终端信噪比 (SNR) 需要强破坏,这增加了在固定采样预算下必须覆盖的噪声范围。我们提出图残差共轭扩散 (GRCD),它用依赖于模态的时钟替换图热扩散的共享时钟,该时钟使每个图傅里叶模态具有相同的条件 SNR。GRCD 在训练划分上拟合一个零均值图谱高斯参考,并停止在一个有限终端 SNR 处,在该处传播后的参考仍携带拟合得到的谱方差。高斯分量在概率流 ODE 中具有精确的逐模态传播子,因此采样以解析方式推进它,并仅对学习到的残差得分进行数值积分。我们在五个设置(METR-LA 交通、Molene 天气以及三个随机块模型)上,在匹配协议下,将 GRCD 与七个比较方法进行评估:图感知扩散 (GAD)、EDM(图骨干)、白化得分扩散 (WSD) 的两种改编版本,以及三个预条件控制。在四次函数评估 (NFEs) 下,GRCD 在所有五个设置上将平均最大均值差异 (aMMD) 较最佳比较方法降低 22 至 36 倍,在 METR-LA 上达到 0.054,在那里它以比达到该目标的最便宜比较方法少 87% 的采样墙钟时间越过 aMMD 0.1 目标。拟合终端参考在有限终端 SNR 下将 aMMD 降低 2.7 至 7.3 倍,而在接近零时这些因子缩小到 1.00 至 1.01。
cs.LG / 135 / 2609.39692
GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales
GFD-OPD:跨尺度的扩散模型引导折叠式同策略蒸馏
diffusion
扩散模型相关
Abstract
On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectly match the distribution of a larger teacher, while classifier-free guidance can accumulate and amplify the distributional discrepancies between the student's conditional and unconditional branches and those of the teacher. To solve this problem, we propose GFD-OPD, a simple yet effective method that reduces the student-teacher gap while avoiding the error amplification of the CFG composition. Across numerous experiments, GFD outperforms previous baselines in both training efficiency and final performance, achieving state-of-the-art results on all benchmarks.
Chinese Translation
同策略蒸馏(OPD)在语言模型中已经展示了两种重要能力:将大型教师模型压缩为较小的学生模型,以及将多个专家模型合并为一个单一模型。然而,现有的扩散 OPD 大多关注后者,其教师模型和学生模型共享相同的骨干网络和规模。我们研究了从大型教师模型到小型学生模型的大到小扩散 OPD,并发现标准方案会失败。为了找到根本原因,我们提出了 Fixed-State KL,一种有效且公平的方法,用于衡量在扩散模型的 OPD 训练过程中学生模型与教师模型之间的分布差距。我们首次阐明了为什么大到小 OPD 对扩散模型而言具有挑战性:较小的学生模型难以完美匹配较大教师模型的分布,而无分类器引导会累积并放大学生模型的条件分支和无条件分支与教师模型对应分支之间的分布差异。为了解决这个问题,我们提出了 GFD-OPD,一种简单而有效的方法,它在缩小学生-教师差距的同时,避免了 CFG 组合的误差放大。在大量实验中,GFD 在训练效率和最终性能方面均优于先前基线,并在所有基准上取得了最先进的结果。
cs.LG / 136 / 2609.39757
Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction
重新审视在线策略对抗式黑盒蒸馏:校准分组奖励几何以实现有效的优势构造
large language model
大语言模型相关
Abstract
Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://github.com/2018cx/GRGC.
Chinese Translation
黑盒蒸馏是一条实用的途径,用于将仅暴露文本输出的、可通过 API 访问的大语言模型的能力迁移到更小的学生模型中。近期诸如 GAD 之类的在线策略对抗式方法通过在评论者(critic)与学生之间形成对抗循环,改进了 SeqKD:其中评论者针对学生在自身采样回复上的基于 GRPO 的学生策略优化提供奖励。然而,GRPO 是根据同一提示下学生样本的组内相对奖励来计算优势的,而评论者主要被训练用于区分教师回复与学生回复。这种目标不匹配可能产生尺度坍缩或边际脆弱的奖励组,从而导致分组优化信号不稳定。我们提出分组奖励几何条件化(Groupwise Reward Geometry Conditioning, GRGC),这是一个两阶段框架,通过在评论者训练和策略优化两个阶段塑造学生侧奖励组来改进优势构造。为改善评论者侧的条件化,高斯分组最优传输校准在训练期间对评论者进行正则化,通过将按提示排序的奖励匹配到以组为中心的高斯分位数,使产生的奖励组具有非坍缩的分布跨度和平滑的按排名排列的间隔。在此基础上,策略侧组功率调制在将按提示的奖励组转换为优势之前对其重新塑形,在保留评论者所诱导的排序的同时,提高与优化相关的边际可分性。跨多种教师模型、学生模型家族与规模以及训练数据集的广泛实验表明,GRGC 在分布内和分布外评估上均具有有效性,同时相较于 GAD 仅引入可忽略的开销。代码可在 https://github.com/2018cx/GRGC 获取。
cs.LG / 137 / 2609.39801
RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models
RATIO:面向量化推理模型的推理分析与Token级推理优化
large language model
大语言模型相关
Abstract
Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approaches mainly rely on complex optimization procedures. More recent lightweight inference strategies instead use predefined overthinking markers, limiting their adaptability across quantized models. To address these issues, we propose Reasoning Analysis and Token-level Inference Optimization (RATIO), a framework that identifies model-specific overthinking tokens and assigns each a tailored penalty. RATIO first introduces Quantization-aware Reasoning Behavior Analysis (QRBA) to identify overthinking tokens by analyzing discrepancies between full-precision and quantized models. It then adopts Token-Specific Penalty Determination (TSPD), which leverages full-precision guidance to derive token-specific penalties without additional training. Extensive experiments show that RATIO achieves a better accuracy-efficiency trade-off than existing token-level interventions. Specifically, RATIO achieves up to 9.8 points accuracy improvement and reduces chain-of-thought (CoT) length by up to 51.3% compared with quantized baselines. The code will be available at https://github.com/steven-bao1/RATIO.
Chinese Translation
训练后量化(PTQ)已成为一种广泛采用的技术,用于降低大语言模型(LLM)的内存占用和推理成本。然而,近期研究表明,当应用于推理模型时,PTQ 不仅会降低推理性能,还会加剧过度思考,导致更长的推理轨迹。这些问题可能会抵消低精度推理所预期的效率收益。现有方法主要依赖复杂的优化过程。较新的轻量级推理策略则使用预定义的过度思考标记,限制了它们跨量化模型的适应性。为解决这些问题,我们提出推理分析与Token级推理优化(RATIO),一个识别模型特定的过度思考token并为每个token分配定制惩罚的框架。RATIO 首先引入量化感知推理行为分析(QRBA),通过分析全精度模型与量化模型之间的差异来识别过度思考token。然后采用Token特定惩罚确定(TSPD),它利用全精度指导来推导token特定的惩罚,而无需额外训练。大量实验表明,RATIO 相比现有token级干预方法实现了更好的准确率-效率权衡。具体而言,与量化基线相比,RATIO 实现了最高 9.8 个百分点的准确率提升,并将思维链(CoT)长度最多减少 51.3%。代码将在 https://github.com/steven-bao1/RATIO 提供。
cs.LG / 138 / 2609.39859
Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models
Fork-dLLM:避免扩散语言模型中的灵活性陷阱
diffusion
扩散模型相关
Abstract
Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.
Chinese Translation
掩码扩散语言模型(dLLMs)在与基于置信度的采样器结合时,通过并行 token 生成展现出更快推理的强大潜力。然而,近期工作表明,此类方法会推迟对高熵 fork 位置(即存在多个合理延续的位置)的去掩码。这导致生成多样性降低(表现为更差的 pass@k 扩展),并限制了从 RL 后训练中可获得的收益。为了避免这一灵活性陷阱,先前的工作主张采用自回归(AR)采样。在此,我们表明,抛弃基于置信度的采样是不必要的,而且一旦将推理成本纳入考量,这种做法是浪费的。我们首先提出 Fork-dLLM,这是一种简单的混合采样器,仅在不确 定的回退步骤使用 AR 风格的排序,而在其他情况下保留并行生成。然后,我们将同样的原则扩展到使用 ForkGRPO 的后训练,该方法使用 Fork-dLLM 的 rollout,并仅在回退步骤应用 GRPO 目标,同时保持精确的策略似然比,并大幅降低 rollout 与优化成本。在我们的实验中,Fork-dLLM 达到了与 AR 采样相当的强 pass@k 扩展,同时效率高出 2-3 倍;而 ForkGRPO 在显著更低的训练成本下,取得了与基于 AR 的 GRPO 基线相当或更好的下游性能。
cs.LG / 139 / 2609.39995
Learning to Explain While Planning: Rule-Aligned Diffusion Planning for Autonomous Driving
在规划中学习解释:面向自动驾驶的规则对齐扩散规划
diffusion
扩散模型相关
Abstract
Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations to fit trajectory distributions, learning statistical correlations among scenes, behaviors, and trajectories without explicitly modeling driving rules. In long-tail scenarios where expert data are scarce, the lack of behaviors to imitate may lead to trajectories that violate safety or compliance requirements. Moreover, their generation process lacks rule-level explanations, making it difficult to determine which rules drive trajectory adjustments, when they take effect, and how strongly they act, thereby limiting failure diagnosis, safety validation, and targeted improvement. To address these limitations, we propose the Rule-Aligned Diffusion Planner (RADP), which incorporates differentiable driving rules into the diffusion objective during training, turning rule knowledge into intrinsic behavioral principles beyond finite demonstrations. We further introduce Rule-Pressure Attribution (RPA), which constructs supervision signals from gradients of rule losses with respect to predicted trajectories and employs a lightweight attribution head to estimate the optimization pressure exerted by each rule online. To assess the closed-loop behavioral relevance of these attributions, we propose a temporal risk-alignment protocol that evaluates whether current rule pressures reflect corresponding risks during subsequent closed-loop execution. Experiments on nuPlan show that RADP improves closed-loop planning in challenging safety-critical scenarios, while RPA exhibits consistent temporal alignment with subsequent rule-specific risks, validating both intrinsic rule learning and rule-level interpretability.
Chinese Translation
扩散规划器在生成多模态轨迹方面展现出强大的能力。然而,现有方法主要依赖专家演示来拟合轨迹分布,学习场景、行为与轨迹之间的统计相关性,而没有显式地建模驾驶规则。在专家数据稀缺的长尾场景中,缺乏可供模仿的行为可能导致生成的轨迹违反安全或合规要求。此外,其生成过程缺乏规则层面的解释,使得难以确定哪些规则驱动了轨迹调整、这些规则何时生效以及其作用强度如何,从而限制了故障诊断、安全验证和针对性改进。为解决这些局限,我们提出规则对齐扩散规划器(Rule-Aligned Diffusion Planner, RADP),其在训练期间将可微驾驶规则纳入扩散目标,将规则知识转化为超越有限演示的内在行为原则。我们进一步引入规则压力归因(Rule-Pressure Attribution, RPA),它从规则损失相对于预测轨迹的梯度构造监督信号,并采用轻量级归因头来在线估计每条规则所施加的优化压力。为了评估这些归因在闭环中的行为相关性,我们提出一种时间风险对齐协议,用于评估当前规则压力是否反映了后续闭环执行过程中相应的风险。在 nuPlan 上的实验表明,RADP 在具有挑战性的安全关键场景中改进了闭环规划,而 RPA 与后续的规则特定风险表现出一致的时间对齐,验证了内在规则学习和规则级可解释性。
cs.LG / 140 / 2609.40030
Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models
Fenchel Tilting:用于生成模型高效微调的加权校正
diffusion
扩散模型相关
Abstract
Adapting a pretrained generative model to an arbitrary preference expressed as a utility function underlies reward alignment, guided design, and constraint satisfaction, enabling diverse applications. Existing fine-tuning methods trade off generality against computational cost: they either restrict the family class of supported preferences to keep optimization simple or preserve generality at the expense of efficiency. We introduce Fenchel Tilt Flow Control (FTFC), which decouples utility optimization from generative-model fitting. FTFC first optimizes for a target distribution by jointly fitting an effective reward and density-ratio weights on pretrained samples. Method combines the utility's variational structure with Fenchel duality, supporting general $f$-divergence penalties that determine how rewards are transformed into an distribution-correction weights. These weights are then frozen and used to modify a diffusion or flow model in a single stage of importance-weighted denoising or flow matching, without differentiating through sampling trajectories. We establish exact duality for concave utilities under suitable conditions and show that weighted fitting reproduces the optimal target distribution for a given utility. Across image and molecule generation benchmarks, FTFC improves over baselines on diverse preference functions, while also being up to $20\times$ more efficient. roposed method enables adaptation beyond expected-reward maximization without complex optimization, while preserving robustness for more general class of the utility functions compared to baselines.
Chinese Translation
将预训练生成模型适配到以效用函数表示的任意偏好,是奖励对齐、引导设计和约束满足的基础,并促成了多样化的应用。现有微调方法在通用性与计算成本之间进行权衡:它们要么限制所支持偏好的族类以保持优化简单,要么以牺牲效率为代价来保持通用性。我们提出 Fenchel Tilt Flow Control (FTFC),它将效用优化与生成模型拟合解耦。FTFC 首先通过在预训练样本上联合拟合一个有效奖励和密度比权重,来优化目标分布。该方法将效用的变分结构与 Fenchel 对偶性相结合,支持一般的 $f$-散度惩罚,这些惩罚决定了奖励如何被转换为分布校正权重。随后,这些权重被冻结,并用于在重要性加权去噪或流匹配的单个阶段中修改扩散模型或流模型,而无需通过采样轨迹进行微分。我们在适当条件下为凹效用函数建立精确对偶性,并表明加权拟合能够复现给定效用的最优目标分布。在图像和分子生成基准上,FTFC 在多种偏好函数上优于基线,同时效率最高可提升 $20\times$。所提方法能够在无需复杂优化的情况下实现超越期望奖励最大化的适配,同时相较于基线,对更一般的效用函数类别保持鲁棒性。
cs.LG / 141 / 2609.40235
Distribution Matching Distillation for Continuous Diffusion Language Models
面向连续扩散语言模型的分布匹配蒸馏
diffusion
扩散模型相关
Abstract
Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
Chinese Translation
连续扩散语言模型能够并行生成所有词元,然而高质量生成仍然可能需要数百次网络评估(NFEs)。我们研究分布蒸馏如何通过利用学生模型的概率化词元输出来降低这一开销。我们的统一表述将学生模型的输出参数化与由此产生的梯度估计器联系起来,并得到两种方法,它们具有相同的学生架构和反向KL匹配目标:Simplex-DMD 使用连续词元松弛和路径梯度,而 Reinforce-DMD 使用类别采样以及带有学习到的密度比的 REINFORCE。我们为多步生成开发了这两种方法,并研究了与每种参数化相关的训练与采样选择。在 OpenWebText 上,对于 1,024 个词元的序列,Simplex-DMD 在仅 4 次 NFE 的情况下、在 5.44 nats 的一元熵下达到 45.6 的生成困惑度,在匹配熵和采样预算的条件下,相对于所评估的最强扩散基线降低了 49%。Reinforce-DMD 在更大预算下推进了前沿,在 256 次 NFE、5.00 nats 的熵下达到 14.9 的生成困惑度,在相同的比较协议下降低了 20%。
cs.LG / 142 / 2609.40335
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
在 DP-SGD 下的隐私设置中,权重绑定对仅解码器 LLM 仍然有益吗?
large language model
大语言模型相关
Abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.
Chinese Translation
差分隐私随机梯度下降(DP-SGD)是用于大语言模型(LLM)隐私保护微调的一种领先方法。许多仅解码器 LLM 在输入和输出嵌入之间采用权重绑定,这一设计选择最初是为了在非隐私设置中提高参数效率并改善语言建模性能而引入的。然而,权重绑定在差分隐私训练下的影响仍很大程度上尚未被探索。在本工作中,我们使用 GPT2 和 DistilGPT2 作为代表性的仅解码器架构,研究权重绑定在 DP 设置中的作用。有趣的是,我们发现,在 DP-SGD 下,未绑定嵌入始终优于权重绑定模型,在 SST-2、QNLI 和 QQP 上取得了最高 4.74 个百分点的准确率提升。除了提升效用之外,解除嵌入绑定还使得能够为 DP-SGD 使用内存高效的 ghost clipping。相比之下,权重绑定引入了共享参数交互,这使标准 ghost norm 计算变得复杂,并在很大程度上抵消了其计算优势。因此,未绑定模型在保留 ghost clipping 优势的同时,实现了超过 60% 的内存使用量降低。我们的结果表明,未绑定嵌入为仅解码器 LLM 的差分隐私训练提供了一种更有效且可扩展的设计,并凸显了在隐私保护设置中重新审视标准 LLM 架构选择的必要性。
cs.LG / 143 / 2609.40360
Semifactual Credit-Augmented Policy Optimization
半事实信用增强策略优化
large language model
大语言模型相关
Abstract
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.
Chinese Translation
具有可验证奖励的强化学习(RLVR)提升了大语言模型(LLM)的推理能力,然而它们的预测仍然对任务无关的提示特征敏感。我们通过保留底层问题及其答案的半事实提示干预来研究这种敏感性。我们的分析揭示了词元级敏感性的显著变化,并表明在解码过程中抑制高漂移词元候选可在不更新模型权重的情况下提高推理准确率。这些发现凸显了组相对策略优化(GRPO)的一个局限:它将相同的、由结果导出的优势分配给每个响应词元,并可能在强化有用推理的同时强化潜在的虚假依赖。受这一观察启发,我们提出半事实信用增强策略优化(SCAPO),这是 GRPO 的一种受因果启发的变体,将半事实稳定性纳入词元级信用分配。SCAPO 测量在半事实干预下固定响应的词元概率漂移,并使用归一化稳定性分数在早期训练期间降低相对不稳定词元的优势,同时不会仅因稳定性本身而给予额外信用。在 Qwen3-4B-Base 和 Qwen3-1.7B-Base 上,SCAPO 在 AIME 2024-2026 上的准确率相较 GRPO 分别提高了 5.63 和 4.17 个百分点。在两种模型规模上,SCAPO 在大多数评估的数学基准以及所有评估的分布外基准上,在所比较的方法中取得了最佳结果。这些结果表明,半事实稳定性提供了一种有效的训练信号,可通过 RLVR 中更细粒度的信用分配来改进推理和泛化。代码可在 https://github.com/DtYXs/SCAPO 获取。
cs.LG / 144 / 2609.40361
Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
面向多模态临床诊断的排序感知提示优化
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.
Chinese Translation
多模态大语言模型(MLLMs)正在迅速推动临床诊断的发展,但其适配流程仍然锚定于基于准确率的目标。临床数据高度类别不平衡:一个始终预测多数类的预测器可以取得超过 90% 的准确率,却在临床上毫无用处。因此,我们评估并优化 AUROC,这是一种无需阈值的分数,它将正例排在负例之上,并且对类别平衡不变。我们关注 MLLMs 中的提示优化。诸如 GEPA 之类的反思方法使用一个二值分数矩阵,其中每个评估实例一行,每个候选提示一列;单元格记录每个实例的正确性,因此列平均值即为准确率,并驱动候选选择。我们引入配对级 Pareto 提示进化(Ranking-PE),它将每个正确性行替换为在(正例,负例)实例对上的成对排序行:如果候选提示给正例的分数高于配对的负例,则该单元格为 1。根据 Wilcoxon-Mann-Whitney 恒等式,列平均值随后等于经验 AUROC。我们将这种替换应用于提示进化搜索所读取的全部三个层面——决定 Pareto 支配关系的分数矩阵、提供给反思 LM 的逐示例反馈,以及最终候选选择——且无需额外模型调用,也不使用代理损失。在 MIMIC 上的三种疾病中,基于准确率的提示进化可能会降低排序性能;Ranking-PE 扭转了这一点,在微调后的 Qwen3-VL-8B 上以 +5.8 AUROC 个百分点的优势超过基于准确率的方案,在 MedGemma-4B 上则高出 +16.2 个百分点。消融实验考察了每个设计组件,并表明医学级视觉骨干网络——通过视觉编码器调优的 SFT 或医学预训练获得——是提示搜索无法替代的先决条件——我们的方案将反思式提示进化从纯文本数据扩展到多模态临床决策。
cs.AI / 145 / 2609.38420
What Was Said, Not What Was 'Thought': Type-6 Logic for CoT Verification
所言,而非‘所想’:面向 CoT 验证的 Type-6 逻辑
large language model
大语言模型相关
Abstract
We introduce Type-6 logic, a variant of dynamic epistemic logic augmented with two operators (uncertainty and recurrence), designed to model the inferential dynamics of contemporary large language model (LLM) chain-of-thought (CoT) reasoning. Type-6 accounts for common LLM reasoning pathologies such as unlicensed revision, enthymemes, loopbacks, and unverifiable/incorrect claims. We propose a verifier based on Type-6 logic that builds a graph out the trace, and checks it against Type-6's axioms and inference rules. We evaluate our framework on LLM-generated CoTs four splits spanning formal and informal reasoning. Our verifier detects structurally unsound reasoning steps that surface-level heuristics miss, and allows for easy visualisation of the model's reasoning process. In our corpus, our verifier shows that derived contradiction is the most common hard-fail category in CoT, and that only about 3\% of the propositions of a trace have impact on the final derivation. Ablation studies show that other verification methods (LLMs-as-judges, other neurosymbolic approaches, etc.) cannot be considered interchangeable: for example, agreement between LLMs-as-judges and LINC is $κ\approx 0.034$, and this persists within a method across underlying models. Type-6, however, is the most agreed-with method amongst the ones we tested. We prove our verifier runs on average-case linear time; and release our logic specification and artefacts.
Chinese Translation
我们引入 Type-6 逻辑,它是动态认知逻辑的一种变体,增加了两个算子(不确定性和重现),旨在对当代大型语言模型(LLM)思维链(CoT)推理的推理动态进行建模。Type-6 解释了常见的 LLM 推理病理,例如未经许可的修正、省略三段论(enthymemes)、回环以及不可验证/不正确的断言。我们提出一个基于 Type-6 逻辑的验证器,它从轨迹中构建一个图,并根据 Type-6 的公理和推理规则对其进行检查。我们在 LLM 生成的 CoT 的四个划分上评估我们的框架,这些划分涵盖形式推理和非形式推理。我们的验证器能够检测到表层启发式方法会漏掉的结构上不健全的推理步骤,并使得模型的推理过程易于可视化。在我们的语料库中,我们的验证器表明,推导出的矛盾是 CoT 中最常见的硬失败类别,并且一条轨迹中只有约 3\% 的命题对最终推导有影响。消融研究表明,其他验证方法(LLM 作为评判者、其他神经符号方法等)不能被视为可互换的:例如,LLM 作为评判者与 LINC 之间的一致性为 $κ\approx 0.034$,并且这在同一种方法内跨不同底层模型持续存在。然而,在我们测试过的方法中,Type-6 是获得最多认同的方法。我们证明我们的验证器以平均情况线性时间运行;并发布我们的逻辑规范和制品。
cs.MA / 146 / 2609.38327
Absorbing State Phase Transitions in Multi-Agent Search
多智能体搜索中的吸收态相变
large language model
大语言模型相关
Abstract
Nontrivial dynamics can emerge in large language model (LLM)-based multi-agent systems, and preliminary evidence exists that formalisms from statistical mechanics can be effective at modeling and predicting such behaviors. In parallel, designing multi-agent communication topology for optimal task-solving is an active research question. In this paper, we focus on predicting the success of multi-agent search tasks using the formalism of absorbing state phase transitions. We first taxonomize search tasks into four types, informed by classical results in combinatorial search. We then theoretically derive a critical communication degree $d_c$, the minimum number of agents each agent can communicate with, above which incorrect hypotheses do not proliferate uncontrollably and the search enters the solved state. Finally, we evaluate frontier LLM-based multi-agent systems on real-world search and discovery tasks, software configuration debugging and physical mechanism discovery, and find that agreement with theory is mixed. LLM agents may not communicate with their neighbors and can develop strategies that are individually beneficial but limits the benefits of collaboration.
Chinese Translation
在基于大语言模型(LLM)的多智能体系统中,可能会出现非平凡动力学,并且已有初步证据表明,来自统计力学的形式体系可以有效地对此类行为进行建模和预测。与此同时,为最优任务求解设计多智能体通信拓扑是一个活跃的研究问题。在本文中,我们聚焦于使用吸收态相变的形式体系来预测多智能体搜索任务的成功。我们首先受组合搜索中的经典结果启发,将搜索任务分类为四种类型。然后,我们从理论上推导出一个临界通信度 $d_c$,即每个智能体能够与之通信的智能体的最小数量,高于该值时,错误假设不会不受控制地增殖,并且搜索进入已解决状态。最后,我们在现实世界的搜索与发现任务、软件配置调试和物理机制发现上评估前沿的基于 LLM 的多智能体系统,并发现与理论的一致性参差不齐。LLM 智能体可能不会与其邻居通信,并且可能发展出对个体有利但限制了协作益处的策略。
cs.MA / 147 / 2609.38516
From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs
从独自学习到社会学习:刻画大型语言模型中的递归社会性改进
large language model
大语言模型相关
Abstract
Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent frameworks often have every model work toward a shared goal. We ask a different question. When each agent pursues its own reward, can self-improving LLMs learn from one another well enough to improve the whole population? We call this capability recursive social improvement. We study populations that revise skill files and choose whether, when, and whom to copy from. Independent search, learning from peers, and acting all share one token budget. In controlled environments, established social-learning algorithms benefit from peers, but three LLMs do not. They earn less reward per token than solo learners, and explore too narrowly or run out of tokens before acting. We then let the models write and revise their own skills. Observing peers changes how they improve, helping one model find useful skills sooner and another spend less on private search. Neither, however, outperforms independent learners at the same cost. Skills are copied, revised, and passed on, so one discovery can seed further search. Yet these exchanges concentrate the population around fewer independent discoveries. Together, these results show that LLMs can make learning more efficient by copying from peers, but not yet more effective.
Chinese Translation
大型语言模型(LLMs)如今能够通过修改其遵循的指令来改进自身,而 LLM 智能体也越来越多地被编排在一起协作解决复杂问题。然而,自我改进方法通常一次只优化一个系统,而多智能体框架往往让每个模型都朝着一个共同目标工作。我们提出一个不同的问题。当每个智能体追求自身奖励时,能够自我改进的 LLMs 能否足够好地相互学习,从而改进整个群体?我们将这种能力称为递归社会性改进。我们研究这样的群体:它们修改技能文件,并选择是否复制、何时复制以及从谁那里复制。独立搜索、向同伴学习和行动共享同一个 token 预算。在受控环境中,已有的社会学习算法能从同伴中获益,但三个 LLMs 却不能。它们每 token 获得的奖励少于独立学习者,并且探索范围过窄,或在行动之前就耗尽了 token。随后,我们让这些模型编写并修改自己的技能。观察同伴会改变它们改进的方式,帮助一个模型更早找到有用的技能,并让另一个模型在私人搜索上花费更少。然而,二者在相同成本下都没有超过独立学习者。技能会被复制、修改并传递下去,因此一个发现可以催生进一步的搜索。然而,这些交换使群体集中在更少的独立发现周围。总之,这些结果表明,LLMs 可以通过从同伴那里复制来使学习更高效,但尚未使其更有效。
cs.MA / 148 / 2609.39211
Consensus and Factual Dynamics in Large Populations of Interacting Language Models
大规模交互语言模型群体中的共识与事实动态
large language model
大语言模型相关
Abstract
Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers and converge towards more factual responses, treating agreement as a proxy for correctness. However, these studies usually fix a single interaction structure, leaving open how consensus depends on how agents interact. We address this gap by introducing RHEON, a physics-inspired framework that recasts a population drawn from a single frozen model as an evolving $O(n)$ spin system on a ladder of interaction geometries of increasing effective dimension --from a 1D ring to a full-coupling mean-field graph-- with the sampling temperature $T$ as the tunable source of thermal disorder, evolved through a Glauber-like asynchronous dynamics. Sweeping RHEON across $432$ configurations of prompt, population size, communication topology, and sampling temperature yields Eraclitus-4.7M, a tagged evolutionary corpus of $4.7$ million responses. We find that agents reach their strongest consensus gain within the first few update sweeps and that increasing the number of neighbours per agent accelerates convergence on average. We further show that whether a configuration settles on factually correct or hallucinated consensus is not predictable from its initial state alone, and that the hallucination-minimising temperature depends on how the agents are coupled, so the common near-greedy default is not automatically the safest. Finally, semantic agreement correlates positively with factual convergence, and interaction strengthens the association, yet never enough for unanimity to certify correctness.
Chinese Translation
大型语言模型(LLM)智能体正越来越多地以相互作用的实体群体的形式被部署,其中,共识——即对某个共享答案的一致意见——作为一种集体性的、非工程化的行为涌现出来。先前关于 LLM 共识的研究表明,智能体可以交叉验证其答案并收敛到更具事实性的回答,将一致意见视为正确性的代理指标。然而,这些研究通常固定单一交互结构,因而未能阐明共识如何取决于智能体之间的交互方式。我们通过引入 RHEON 来填补这一空白;RHEON 是一个受物理学启发的框架,它把从单一冻结模型中抽取的群体重新表述为在一个有效维度递增的交互几何阶梯上演化的 $O(n)$ 自旋系统——从一维环到全耦合平均场图——其中采样温度 $T$ 是可调的热无序来源,并通过类 Glauber 的异步动力学演化。在提示、群体规模、通信拓扑和采样温度的 $432$ 种配置上扫描 RHEON,产生了 Eraclitus-4.7M,一个包含 $4.7$ 百万条响应的带标签演化语料库。我们发现,智能体在最初几次更新扫描内达到其最强的共识增益,并且每个智能体的邻居数量增加平均而言会加速收敛。我们进一步表明,一个配置最终稳定在事实正确还是幻觉性的共识上,不能仅从其初始状态预测;并且最小化幻觉的温度取决于智能体如何耦合,因此常见的近贪心默认设置并非自动就是最安全的。最后,语义一致性与事实性收敛正相关,并且交互会增强这种关联,但它从未强到足以让全体一致证明正确性。
cs.NE / 149 / 2609.40258
Large Language Model-Guided Evolutionary Discovery of Native Neural Architectures for Spiking Sequence Modeling
大语言模型引导的用于脉冲序列建模的原生神经架构进化发现
large language model
大语言模型相关
Abstract
Spiking neural networks (SNNs) offer low-energy sequence modeling through sparse, event-driven computation. However, interactions among spike encoding, neuronal dynamics, and information propagation complicate architecture design. Existing SNN sequence models often adapt artificial neural network (ANN) architectures designed for real-valued activations, potentially underusing spike-based communication and temporal state updates, motivating automated discovery of native SNN architectures. Most evolutionary neural architecture search (ENAS) methods operate within predefined configuration spaces, limiting discovery to mechanisms expressible within those spaces. We introduce OpenArchEvo, which uses large language models (LLMs) to evolve executable architecture code in an open program space under spiking-projection constraints. In this space, code differences need not reflect architectural novelty, while direct performance evaluation requires costly training. We construct a three-view representation spanning code, design rationale, and a behavioral fingerprint to support novelty estimation and performance prediction. The search treats predicted performance and estimated novelty as two objectives, using surrogate predictions to select candidates for expensive training evaluations. With an estimated candidate-training cost of 132 V100 GPU-days, the search uncovers multiple native SNN architectures, exemplified by three designs featuring mechanisms such as spike-activity-dependent control of state updates and residual pathways. The discovered NeuroGate surpasses the ANN DeltaNet on WikiText-103, and the discovered architectures reduce estimated architecture-level arithmetic energy by up to 50.6x (LoopMem) relative to a common dense Transformer (ANN) baseline. All code and all discovered architectures will be made publicly available soon.
Chinese Translation
脉冲神经网络(SNN)通过稀疏、事件驱动的计算提供低能耗序列建模。然而,脉冲编码、神经元动力学和信息传播之间的相互作用使架构设计变得复杂。现有 SNN 序列模型通常沿用为实值激活设计的人工神经网络(ANN)架构,可能未充分利用基于脉冲的通信和时间状态更新,这促使人们自动发现原生 SNN 架构。大多数进化神经架构搜索(ENAS)方法在预定义配置空间内运行,将发现限制在这些空间内可表达的机制中。我们提出 OpenArchEvo,它使用大语言模型(LLM)在脉冲投影约束下的开放程序空间中演化可执行架构代码。在这个空间中,代码差异不一定反映架构新颖性,而直接性能评估需要昂贵的训练。我们构建了一个涵盖代码、设计理由和行为指纹的三视图表示,以支持新颖性估计和性能预测。该搜索将预测性能和估计新颖性视为两个目标,并使用代理预测来选择候选者进行昂贵的训练评估。在估计候选训练成本为 132 个 V100 GPU 日的情况下,该搜索发现了多个原生 SNN 架构,以三种设计为例,这些设计具有诸如依赖脉冲活动的状态更新控制和残差通路等机制。所发现的 NeuroGate 在 WikiText-103 上超越了 ANN DeltaNet,并且所发现的架构相对于常见的密集 Transformer(ANN)基线,将估计的架构级算术能量最多降低 50.6 倍(LoopMem)。所有代码和所有发现的架构将很快公开发布。
cs.OS / 150 / 2609.39819
Capture the lifecycle: KV Cache management in ReAct Agents with KVTether
捕捉生命周期:使用 KVTether 管理 ReAct 智能体中的 KV Cache
large language model
大语言模型相关
Abstract
Efficient serving of long-context reasoning-and-acting (ReAct) agents relies on KV cache reuse to reduce large language model (LLM) prefill latency and monetary cost. However, a semantic gap exists between agent harnesses and the underlying serving stack. Through context mutation, tool execution, and subagent coordination, context messages may become actively engaged, permanently discarded, and temporarily unused, while the serving stack only observes accesses to the corresponding KV cache. This lifecycle blindness prevents recency-only policies such as LRU from reclaiming dead KV promptly and from preserving older KV that will be reused sooner than newer entries. We present KVTether, a lifecycle-aware KV cache management framework for ReAct agents. By tracing semantic primitives embedded in agent harnesses, KVTether captures runtime lifecycle semantics during highly dynamic execution. KVTether then translates message-level semantics into KV-level lifecycle states and uses these states to drive state-prioritized cache management without exposing physical complexities to agent harnesses. After reclaiming dead KV, KVTether preferentially preserves live-but-idle KV that is waiting for reuse, reducing premature eviction before reuse. Across agent benchmarks and production workloads, KVTether reduces end-to-end request latency by up to 26.3% and 17.4% relative to LMCache and MORI, respectively, and lowers estimated task cost by 40.0% and 33.2% on average.
Chinese Translation
长上下文推理与行动(ReAct)智能体的高效服务依赖于 KV 缓存复用,以降低大语言模型(LLM)的预填充延迟和金钱成本。然而,智能体框架与底层服务栈之间存在语义鸿沟。通过上下文变更、工具执行和子智能体协调,上下文消息可能处于活跃使用、被永久丢弃或暂时未被使用的状态,而服务栈只能观察到对相应 KV 缓存的访问。这种生命周期盲区使得诸如 LRU 之类的仅基于最近使用时间的策略无法及时回收已死亡的 KV,也无法保留那些将比更新的条目更早被复用的较旧 KV。我们提出 KVTether,一个面向 ReAct 智能体的生命周期感知 KV 缓存管理框架。通过追踪嵌入在智能体框架中的语义原语,KVTether 能在高度动态的执行过程中捕获运行时生命周期语义。随后,KVTether 将消息级语义转化为 KV 级生命周期状态,并利用这些状态驱动按状态划分优先级的缓存管理,同时不向智能体框架暴露物理层面的复杂性。在回收已死亡的 KV 之后,KVTether 优先保留正在等待复用的存活但空闲的 KV,从而减少复用前的过早驱逐。在智能体基准测试和生产工作负载中,相较于 LMCache 和 MORI,KVTether 将端到端请求延迟分别最多降低 26.3% 和 17.4%,并将估计任务成本平均降低 40.0% 和 33.2%。
cs.OS / 151 / 2609.40247
Herschel: Continuous Optimization of Production LLM Inference through On-Demand Profiling
Herschel:通过按需性能分析实现生产级 LLM 推理的持续优化
large language model
大语言模型相关
Abstract
Model-as-a-service platforms call for continuous optimization as complex serving conditions expose inefficiencies missed before deployment. Detailed always-on profiling can incur substantial overhead, while lightweight collection omits information needed for diagnosis. We present Herschel, a continuous optimization system for production large language model (LLM) inference. Our key insight is that adaptive, on-demand profiling can provide rich full-stack evidence without continuous collection. Herschel safely attaches to and detaches from selected running processes without engine changes or restarts, and adapts coverage as investigations reveal missing evidence. Herschel reconstructs operator executions and cross-process dependencies to identify inefficiency mechanisms and suggest solutions using applicable reference fixes. AI agents implement and test engine and kernel changes under controlled conditions that preserve the triggering workload and dependencies, with expert review before deployment. Controlled tests show active-collection overhead below 0.5% for time to first token and 7% for time per output token. Bounded windows, typically 30 s, avoid the continuous cost of always-on tracing. Over six months, Herschel collected approximately 17,000 traces across over 120 model variants and more than 10 accelerator types, identifying inefficiency patterns in 23% of the traces. Representative findings guide widely deployed optimizations, including restructured synchronization, removal of unused computation, and improved operator implementations.
Chinese Translation
模型即服务平台需要持续优化,因为复杂的服务条件会暴露出部署前被遗漏的低效问题。详细的常开性能分析会带来显著开销,而轻量级采集会遗漏诊断所需的信息。我们提出 Herschel,一个面向生产级大型语言模型(LLM)推理的持续优化系统。我们的核心洞见是,自适应、按需的性能分析能够在不进行持续采集的情况下提供丰富的全栈证据。Herschel 能够安全地附加到选定的运行中进程并从中分离,无需更改引擎或重启,并且随着调查揭示缺失证据而调整覆盖范围。Herschel 重建算子执行和跨进程依赖关系,以识别低效机制,并利用适用的参考修复方案提出解决方案。AI 代理在受控条件下实现并测试引擎和内核变更,这些条件保留触发工作负载和依赖关系,并在部署前进行专家评审。受控测试表明,主动采集在首 token 时间上的开销低于 0.5%,在每输出 token 时间上的开销为 7%。有界窗口,通常为 30 秒,避免了常开追踪的持续成本。在六个月多的时间里,Herschel 在超过 120 个模型变体和 10 种以上加速器类型上收集了约 17,000 条 trace,并在 23% 的 trace 中识别出低效模式。具有代表性的发现指导了广泛部署的优化,包括重构同步、移除未使用的计算以及改进算子实现。
cs.SE / 152 / 2609.40119
Automatically Building Machine-Checked Assurance Cases from C Codebases to Requirements
从 C 代码库到需求:自动构建机器可检查的保证案例
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown promise in automating interactive theorem proving, yet verification of real-world C codebases requires more than discharging individual proof goals. The task involves jointly constructing expressive function specifications and their proofs, and ensuring that library interfaces compose along intended call sequences even without a designated client. This paper presents CCV, an LLM-assisted framework for building machine-checked assurance cases: structured, auditable artifacts supporting the claim that a C codebase meets its intended requirements. To model intended cross-interface use in open libraries, CCV constructs an interface protocol that exposes permitted call sequences and resource assumptions for review, with a conditional safety guarantee under verified contracts and caller obligations. CCV coordinates two complementary phases: (i) requirement-guided analysis and bottom-up construction of candidate specifications and protocols; and (ii) modular proof construction with feedback that revises the specifications and proofs. Implemented using VST in Rocq, CCV verifies memory safety and leak freedom for all 299 function definitions across six C benchmarks, including industrial cryptographic components, with less than one person-day of reported human effort per benchmark. The guarantees depend on disclosed contracts and assumptions; human review supplies the conformance judgments connecting the formal artifacts to the intended requirements.
Chinese Translation
大语言模型(LLMs)在自动化交互式定理证明方面已展现出前景,然而对真实世界 C 代码库的验证所需的远不止是逐个消解证明目标。该任务涉及联合构造具有表现力的函数规范及其证明,并确保即使没有指定的客户端,库接口也能沿预期的调用序列进行组合。本文提出 CCV,一个由 LLM 辅助的框架,用于构建机器可检查的保证案例:即结构化的、可审计的工件,用以支持 C 代码库满足其预期需求这一主张。为了对开放库中预期的跨接口使用进行建模,CCV 构造一种接口协议,该协议将允许的调用序列和资源假设公开以供审查,并在已验证的契约和调用方义务之下给出条件性安全保证。CCV 协调两个互补的阶段:(i) 需求引导的分析以及候选规范与协议的自底向上构造;(ii) 模块化的证明构造,并带有可对规范与证明进行修订的反馈。CCV 使用 Rocq 中的 VST 实现,对六个 C 基准(包括工业密码学组件)中全部 299 个函数定义验证了内存安全性与无泄漏性,每个基准所报告的人工投入不到一个人日。这些保证依赖于所披露的契约与假设;人工审查提供了一致性判断,将形式化工件与预期需求连接起来。
cs.LG / 153 / 2609.38472
Diffusion-2BC: Hybrid Diffusion and Regression Training for Offline Behavior Cloning in Autonomous Driving
Diffusion-2BC:面向自动驾驶中离线行为克隆的混合扩散与回归训练
diffusion
扩散模型相关
Abstract
Behavior cloning provides an offline route to autonomous-driving policy learning, but mean-squared-error regression is poorly matched to demonstrations in which one observation admits several valid actions. Diffusion policies can represent conditional multimodal action distributions, yet their closed-loop performance may be unstable when visual features and control are learned from limited data. This paper presents Diffusion-2BC, which combines a diffusion denoising objective with an auxiliary deterministic behavior-cloning loss over a shared visual encoder. The auxiliary branch is used only during training; inference remains diffusion-based. The proposed method is evaluated in the controlled Claw environment and in bird's-eye-view CARLA navigation, including route-conditioned driving, route-free navigation through multiple intersections, and cross-map evaluation from Town01 to Town02. In the Claw task, Diffusion-2BC reduced the mean mask-distance error by approximately 10% relative to a diffusion-based behavior-cloning baseline and by 85% relative to standard deterministic behavior cloning. In route-free CARLA, Diffusion-2BC traveled substantially farther before termination under the evaluation protocol than both baselines in Town01 and Town02. Additional qualitative rollouts revealed distinct route choices, showing the multimodal behavior of the proposed diffusion-based agent. The results indicate that an auxiliary regression signal can improve the closed-loop reliability of diffusion behavior cloning while preserving multimodal prediction in the controlled benchmark.
Chinese Translation
行为克隆为自动驾驶策略学习提供了一条离线路径,但均方误差回归与以下演示数据匹配不佳:在这些数据中,一个观测可能对应多个有效动作。扩散策略能够表示条件多模态动作分布,但当视觉特征和控制从有限数据中学习时,其闭环性能可能不稳定。本文提出 Diffusion-2BC,它基于共享视觉编码器将扩散去噪目标与辅助的确定性行为克隆损失相结合。辅助分支仅在训练期间使用;推理仍然基于扩散。所提方法在受控的 Claw 环境和鸟瞰图 CARLA 导航中进行评估,包括路线条件化驾驶、通过多个交叉口的无路线导航,以及从 Town01 到 Town02 的跨地图评估。在 Claw 任务中,相对于基于扩散的行为克隆基线,Diffusion-2BC 将平均掩码距离误差降低了约 10%;相对于标准确定性行为克隆,降低了 85%。在无路线 CARLA 中,在评估协议下,Diffusion-2BC 在 Town01 和 Town02 中终止前行驶的距离均显著远于两个基线。额外的定性 rollout 揭示了不同的路线选择,展示了所提出的基于扩散的智能体的多模态行为。结果表明,辅助回归信号能够提高扩散行为克隆的闭环可靠性,同时在受控基准中保持多模态预测。
cs.AI / 154 / 2609.38982
SimEX: Simulation-Integrated Robotics AutoResearch
SimEX:仿真集成的机器人自动研究
large language model
大语言模型相关
Abstract
Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physical world remains challenging. On the one hand, direct generation methods (e.g., Code as Policies) often suffer from the LLMs' insufficient understanding of robots and physical environments. On the other hand, iterative trial-and-error tuning in the physical world (e.g., physical autoresearch) induces significant experimental cost and safety concerns. We introduce SimEX: Simulation-Integrated Robotics AutoResearch, an autoresearch framework that tightly integrates simulated experimentation, enabling coding agents to efficiently acquire physical capabilities for controlling real robots. SimEX operates in two stages. First, the agent conducts open-ended probe-and-optimize iterations in simulation, developing a robot toolbox with robust and generalizable capabilities. Second, the agent adapts the toolbox and the simulator together through only a few physical trials: each trial corrects the simulator, and the corrected simulator is used to diagnose failures and screen candidate repairs. We evaluate SimEX extensively in sim-to-sim settings and on physical robots. On challenging real-world manipulation tasks including towel folding, barcode scanning, and plate manipulation, SimEX enables coding agents to efficiently acquire robot skills without any demonstration and with only 10 minutes of real-robot interaction. These results suggest that simulation can be a critical component in achieving physical intelligence, not only as a source of training data that must closely replicate the real world, but also as a roughly correct laboratory where a coding agent develops the knowledge and procedures needed to act on the robot. More details and robot videos at https://robo-simex.github.io/
Chinese Translation
由大语言模型(LLMs)驱动的编码智能体已展现出显著能力,能够在数字世界中自主推理并实现目标。然而,将这一成功带到物理世界仍然具有挑战性。一方面,直接生成方法(例如 Code as Policies)通常受制于 LLMs 对机器人和物理环境理解不足。另一方面,在物理世界中进行迭代式试错调优(例如物理自动研究)会带来显著的实验成本和安全问题。我们提出 SimEX:仿真集成的机器人自动研究,这是一个紧密集成仿真实验的自动研究框架,使编码智能体能够高效获得控制真实机器人的物理能力。SimEX 分两个阶段运行。首先,智能体在仿真中进行开放式的探测与优化迭代,开发出具有稳健且可泛化能力的机器人工具箱。其次,智能体仅通过少量物理试验来同时调整工具箱和模拟器:每次试验都会校正模拟器,而校正后的模拟器用于诊断失败并筛选候选修复方案。我们在 sim-to-sim 设置中和物理机器人上对 SimEX 进行了广泛评估。在包括叠毛巾、扫描条形码和操作盘子等具有挑战性的真实世界操作任务上,SimEX 使编码智能体能够在没有任何演示的情况下,仅用 10 分钟的真实机器人交互,就高效获得机器人技能。这些结果表明,仿真可以成为实现物理智能的关键组成部分,不仅作为必须紧密复现真实世界的训练数据来源,也作为一个大致正确的实验室,在其中编码智能体发展出对机器人采取行动所需的知识和程序。更多细节和机器人视频见 https://robo-simex.github.io/
cs.AI / 155 / 2609.38984
Sparse-WAM: Accelerating World Action Models via Action-Guided Sparse Imagination
Sparse-WAM:通过动作引导的稀疏想象加速世界动作模型
diffusion
扩散模型相关
Abstract
World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by token pruning that prioritizes visual fidelity to reduce denoising costs in video diffusion models. However, these methods do not use action relevance to determine which future-frame tokens to retain during joint denoising in WAMs. In this paper, we propose Sparse-WAM, a training-free framework for action-guided sparse imagination that selectively processes future-frame tokens to accelerate WAM inference. We observe substantial overlap in the spatial distribution of attention from action tokens to future-frame tokens (action-to-future attention) between consecutive denoising steps, despite continued updates to the future representations. Motivated by this, we develop Action-Guided Token Selection to retain frame-specific action-relevant regions together with cross-frame context. However, a naive implementation can incur attention-scoring and token-packing overhead that offsets the computational savings from pruning. We therefore introduce Pilot, an efficient engine that reduces sparse inference overhead through lightweight scoring and cross-step reuse of token selections. On LIBERO with FastWAM-Joint and RoboLab-120 with Cosmos 3 Edge, Sparse-WAM achieves inference speedups of approximately $2.0\times$ and $1.8\times$, respectively, over dense eager inference on an NVIDIA RTX 4090, while largely preserving task performance.
Chinese Translation
世界动作模型(WAMs)利用预训练的视频模型,通过联合预测未来的视觉状态与动作来提升机器人控制的泛化能力。这一能力伴随着可观的推理开销,因为在去噪过程中,稠密的未来帧 token 会被反复处理。此前的方法通过在视频扩散模型中优先保证视觉保真度的 token 剪枝来降低去噪开销。然而,这些方法并未利用动作相关性来决定在 WAM 的联合去噪过程中应保留哪些未来帧 token。本文提出 Sparse-WAM,一个无需训练的动作引导稀疏想象框架,它通过有选择地处理未来帧 token 来加速 WAM 推理。我们观察到,尽管未来表示在持续更新,相邻去噪步之间从动作 token 到未来帧 token 的注意力(动作到未来注意力)在空间分布上仍存在显著重叠。受此启发,我们提出动作引导的 Token 选择(Action-Guided Token Selection),在保留跨帧上下文的同时,保留各帧特有的与动作相关的区域。然而,朴素的实现会带来注意力打分与 token 打包的开销,从而抵消剪枝带来的计算节省。因此,我们提出 Pilot,一个高效引擎,通过轻量级打分与跨步复用 token 选择来降低稀疏推理开销。在 LIBERO 上使用 FastWAM-Joint、在 RoboLab-120 上使用 Cosmos 3 Edge 时,Sparse-WAM 在 NVIDIA RTX 4090 上相比稠密的即时(eager)推理分别实现了约 $2.0 imes$ 和 $1.8 imes$ 的推理加速,同时基本保持了任务性能。
cs.AI / 156 / 2609.39575
ECHO-G: Embodied Co-speech Humanoid mOtion Generation
ECHO-G:具身共语人形运动生成
diffusion
扩散模型相关
Abstract
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.
Chinese Translation
为人形机器人生成全身共语运动,需要协调语音韵律、语言内容以及具身特定的运动。为此,我们提出 ECHO-G,一个以语音音频和带时间标注的转录文本为联合条件的框架。其基于语音的扩散 Transformer(SGDiT)将帧对齐的声学特征与 token 级语言上下文相结合,同时保留它们各自不同的粒度。通过修正流匹配进行训练,它直接在机器人空间中建模一对多的话语-运动关系。为支持训练和评估,我们引入了一个源自 BEAT2 的音频-文本-机器人数据集,以及一个涵盖共语特征、机器人运动质量和运行时效率的基准。比较评估支持直接机器人空间生成,相较所评估的人类运动生成与重定向流水线更优,而模态消融则凸显了音频-文本联合条件的优势。我们进一步展示了在实体人形机器人上的部署。一项补充的视频评分研究也表明,联合条件优于其他替代方案。数据集以及训练、推理和评估代码可通过我们的项目页面获取。
cs.AI / 157 / 2609.39599
Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization
文本到3D策略:面向未见规格泛化的细粒度语言-行为对齐
diffusion
扩散模型相关
Abstract
3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.
Chinese Translation
3D视觉运动策略为空间精确操作提供了坚实基础,然而当前的文本到3D策略难以遵循演示所涵盖范围之外的未见细粒度行为规格。我们将这一挑战作为未见规格泛化来研究,其中语言指定了在策略训练中缺失的、具有行为重要性的变化,例如目标位置、位移或关节状态。我们发现,预训练语言表示和传统的全局行为-语言对齐能够捕获粗粒度任务语义,但常常模糊掉需要不同行为的邻近规格。我们提出T3DP,一个用于细粒度语言-行为对齐的文本到3D策略框架。T3DP不将每条指令和演示压缩为单个全局嵌入,而是保留它们的局部结构,并在语言元素与行为片段之间建立双向词元级对应关系。这直接将细微的语言变化锚定到它们所影响的行为组件中,防止密切相关的规格在表示空间中坍缩。由此得到的规格敏感语言表示为一个基于点云的3D扩散策略提供条件,从而在不修改底层策略架构的情况下,对未见行为规格实现更精确的控制。在Meta-World、ManiSkill和RoboTwin上,T3DP相比全局语言-行为对齐将平均留出规格成功率提高了+11.0-14.2个百分点,并在全部15个任务族上均有提升;在真实机器人任务上,它进一步将平均成功率从47.5%提高到65.0%(+17.5个百分点)。表示分析和动作探针分析表明,细粒度对齐更好地保留了规格几何结构和动作相关变化,从而将局部行为锚定与下游控制联系起来。
cs.AI / 158 / 2609.39969
TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks
TACTIC:面向路侧 LiDAR 攻击的时序与上下文感知 LLM 战术规划
large language model
大语言模型相关
Abstract
Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's native point clouds or internal processing. Local perception provides metric vehicle states, while the MLLM combines these measurements with roadside imagery to infer relational traffic context and construct a semantic scene graph. Based on this representation, TACTIC selects and configures two complementary primitives: \emph{push-away}, which shifts the perceived range of a lead vehicle, and \emph{phantom-obstacle braking}, which triggers emergency braking through obstacle injection. Measured traffic states and empirically calibrated constraints ground the generated tactics in physically feasible operating regions. To accommodate MLLM latency, TACTIC overlaps reasoning and execution asynchronously while high-rate local perception detects scene changes and triggers replanning. Across 280 randomized CARLA trials, the full policy achieves a 100% collision rate, compared with 35% for a fixed rule, 60% for random selection, and 75% for a restricted LLM using mode selection with default parameters. Joint physical-and-image input achieves 100% success, versus 65% with physical measurements alone and 75% with imagery alone, while asynchronous $Δ$ refresh reduces scene-mutation response from 7.4 s to 2.0 s. These results show that scene-dependent tactical planning can expose context-sensitive LiDAR failure modes that fixed attack policies may miss.
Chinese Translation
物理 LiDAR 攻击通常使用固定原语和人工选定的参数进行评估,尽管它们强烈依赖周围的交通环境。我们提出 TACTIC,一个场景感知框架,它使用多模态大语言模型(MLLM)来协调状态自适应的路侧 LiDAR 攻击。在灰盒威胁模型下,TACTIC 仅依赖攻击者操作的路侧感知栈,而不访问受害 LiDAR 的原生点云或内部处理过程。本地感知提供度量化的车辆状态,而 MLLM 将这些测量结果与路侧图像相结合,以推断关系型交通上下文并构建语义场景图。基于这一表示,TACTIC 选择并配置两种互补的原语:\emph{push-away},它移动前车的被感知距离,以及 \emph{phantom-obstacle braking},它通过障碍物注入触发紧急制动。实测的交通状态与经经验校准的约束将生成的战术限定在物理可行的运行区域内。为适应 MLLM 的延迟,TACTIC 以异步方式重叠推理与执行,同时高频率的本地感知检测场景变化并触发重新规划。在 280 次随机化的 CARLA 试验中,完整策略实现了 100% 的碰撞率,相比之下,固定规则为 35%,随机选择为 60%,而使用默认参数进行模式选择的受限 LLM 为 75%。物理与图像联合输入实现了 100% 的成功率,相比之下,仅使用物理测量为 65%,仅使用图像为 75%;同时异步 $Δ$ 刷新将场景突变的响应时间从 7.4 s 降低到 2.0 s。这些结果表明,依赖场景的战术规划能够暴露出固定攻击策略可能遗漏的上下文敏感型 LiDAR 失效模式。
cs.AI / 159 / 2609.40165
PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors
PrefPI:偏好引导转向分布外行为
diffusion
扩散模型相关
Abstract
We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy. Our key idea is to formulate preference learning as preference-conditioned generative modeling: preferred trajectories define a conditional distribution, whose density ratio with the broader behavior prior provides an implicit preference signal amplified by classifier-free guidance (CFG). Repeating this preference-conditioned modeling and guidance step yields a form of preference-guided policy iteration, turning incremental improvements toward previously inaccessible behaviors. Across diffusion policies and the PI0.5 flow- matching VLA in simulation and the real world, PrefPI produces substantial behavioral shifts with limited feedback. In particular, PrefPI increases object transport height from 10.7 cm to 19.8 cm on real hardware with only 150 preference-labeled trajectories.
Chinese Translation
我们提出 PrefPI(Preference-Guided Policy Iteration,偏好引导的策略迭代),一个迭代框架,用于仅使用对自生成轨迹的相对偏好来引导预训练的生成式机器人策略。与主要锐化策略已表示模式的先前偏好学习方法不同,我们研究超越初始有效支撑的引导,其中期望行为在初始策略下很少或从未被观察到。我们的关键思想是将偏好学习表述为偏好条件生成建模:偏好轨迹定义一个条件分布,其与更广泛行为先验的密度比提供一个由无分类器引导(CFG)放大的隐式偏好信号。重复这种偏好条件建模和引导步骤会产生一种偏好引导的策略迭代,使增量改进朝向先前无法触及的行为。在扩散策略以及仿真和真实世界中的 PI0.5 流匹配 VLA 上,PrefPI 在有限反馈下产生显著的行为转变。特别地,PrefPI 在真实硬件上仅用 150 条偏好标注轨迹就将物体搬运高度从 10.7 cm 提高到 19.8 cm。
cs.AI / 160 / 2609.39453
From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models
从语音到可编辑概念:用概念瓶颈模型探究情感识别
large language model
大语言模型相关
Abstract
Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
Chinese Translation
语音情感识别(SER)是为话语分配情感标签的任务。早期系统依赖声学特征,而近期方法则结合多种模态,最常见的是语音和文本。然而,在许多数据集上性能仍然不佳。因此,大语言模型(LLM)在SER领域引起了关注,因为它们能够结合指令联合处理多样化的输入。然而,直接的音频输入引发了可解释性方面的问题。为解决图像分类中的类似问题,研究者提出了概念瓶颈模型。本工作将概念瓶颈适配到SER,以考察单个预测如何依赖于转写文本、声学描述和说话人属性。实验在CREMA-D、IEMOCAP和MELD上测试了三个LLM,概念由独立的工具提取。在脚本化语料上,LLM在零样本设置下强烈偏向转写文本,这使CREMA-D上的Macro-F1从27.8降至5.8。微调消除了这一偏差,转写文本使Macro-F1从41.8提升至45.1。在CREMA-D上,移除语速会使48%的中性预测变为厌恶;在MELD上,移除强度水平虽使Macro-F1变化很小,却改变了预测结果。这些发现表明,仅凭总体性能变化无法捕捉概念移除对单个预测的影响。
cs.LG / 161 / 2609.39552
Ghost in the Encoder: Decodable Artist Identity Representations in Lyrics-to-Song Generation
编码器中的幽灵:歌词到歌曲生成中可解码的艺术家身份表征
diffusion
扩散模型相关
Abstract
Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these phenomena have been documented behaviorally on small datasets, little is known about the internal representations that may give rise to them. Prior interpretability work on generative audio has focused on locating semantic concepts such as genre or time signature within model activations. In this work, we show that a trained model can be probed for linearly decodable representations of artist identity from song lyrics alone, without any additional identifiers. Through a controlled case study of ACE-Step 1.5 spanning 2,000 songs across 100 artists, we demonstrate that the artist associated with a given set of lyrics can be identified within the model's internal activations, and that this conditioning signal propagates from the lyric encoder to the diffusion backbone during inference. These findings indicate that lyrics constitute an artist-level conditioning channel not addressed by prompt-side replication safeguards. More broadly, our work highlights how latent-space analysis can be used to audit what generative music models have implicitly learned from their training data.
Chinese Translation
文本到歌曲生成模型可以被提示去模仿特定的艺术家,或从其训练数据中复现整首歌曲。尽管这些现象已在小型数据集上从行为层面得到记录,但人们对可能引发这些现象的内部表征却知之甚少。以往针对生成式音频的可解释性研究主要聚焦于在模型激活中定位诸如流派或拍号之类的语义概念。在这项工作中,我们表明,仅凭歌曲歌词、无需任何额外标识符,就可以从训练好的模型中探测到线性可解码的艺术家身份表征。通过对 ACE-Step 1.5 进行的受控案例研究——涵盖 100 位艺术家的 2,000 首歌曲——我们证明,与给定一组歌词相关联的艺术家可以在模型的内部激活中被识别出来,并且这一条件信号在推理过程中会从歌词编码器传播到扩散主干网络。这些发现表明,歌词构成了一条艺术家级别的条件通道,而提示侧的防复现保护措施并未对此加以处理。更广泛地说,我们的工作凸显了潜空间分析如何能够被用于审计生成式音乐模型从其训练数据中隐式学到了什么。
cs.AI / 162 / 2609.40087
MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
MeanVoiceFlow2:用于快速一步零样本语音转换的均值流与内容编码器的联合优化
diffusion
扩散模型相关
Abstract
Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9\times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.
Chinese Translation
基于流匹配的语音转换(VC)方法因其高语音质量和强说话人相似度而受到关注。其中,诸如 MeanVoiceFlow 这类一步模型尤其具有吸引力,因为它们能够实现高效推理;然而,它们对计算密集型内容编码器的依赖仍然是一个瓶颈。因此,我们提出 MeanVoiceFlow2,一个联合优化基于流的转换模块与计算高效的内容编码器的框架。该模型通过使用 MeanVoiceFlow 的转换蒸馏以及对真实数据的重建进行训练。我们进一步引入带有样本混合和教师引导条件增强的扩散-GAN 训练,以增强真实感和解耦性。在零样本 VC 上的实验表明,MeanVoiceFlow2 实现了更高的感知质量,并且推理速度比 MeanVoiceFlow 快约 $9\times$,同时保持相当的说话人相似度。音频样本可在 https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/ 获取。
cs.SE / 163 / 2609.38335
E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
E2E-SWE:评测大语言模型从零构建可运行代码库的能力
large language model
大语言模型相关
Abstract
Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely specified and independent of any particular implementation. We introduce E2E-SWE, a benchmark for evaluating whether coding agents can build complete, functional software repositories end to end. E2E-SWE contains 186 whole-repository generation tasks spanning 11 programming languages. Given only a natural-language specification and an empty workspace, an agent must implement a complete, installable project that satisfies a comprehensive suite of hidden tests. Each task is constructed by a software engineer in collaboration with an LLM; together, they develop the test suite and a corresponding implementation-independent specification. To ensure that tasks are well specified and practically solvable, we further subject them to an iterative verification process in which autonomous agents audit and repair task defects using static inspection and failures observed from real model rollouts. Evaluating 13 frontier models, we find substantial variation in end-to-end repository generation ability, with pass@1 ranging from 11.7% to 67.7%, providing strong model differentiation while leaving considerable headroom for future progress. Analysis of agent trajectories further reveals long, front-loaded reasoning patterns, highlighting the planning and system-level reasoning required to construct working codebases from scratch.
Chinese Translation
由大语言模型(LLMs)驱动的编码智能体正从进行局部代码修改,演进为开发完整的软件仓库。然而,评估仓库级别的生成仍然具有挑战性:任务必须要求系统级推理,同时又要确保所有被评估的行为都被精确规定,并且不依赖于任何特定实现。我们提出了 E2E-SWE,一个用于评估编码智能体能否端到端地构建出完整、可运行的软件仓库的基准。E2E-SWE 包含 186 个整仓库生成任务,涵盖 11 种编程语言。仅给定一份自然语言规格说明和一个空工作区,智能体必须实现一个完整、可安装的项目,并满足一整套隐藏测试。每个任务都由一名软件工程师与一个 LLM 协作构建;二者共同开发测试套件以及一份相应的、与实现无关的规格说明。为确保任务被良好地规定并且实际上可解,我们进一步让这些任务接受一个迭代式验证过程,在该过程中,自主智能体利用静态检查以及从真实模型运行轨迹中观察到的失败来审计并修复任务缺陷。对 13 个前沿模型进行评估后,我们发现端到端仓库生成能力存在显著差异,pass@1 从 11.7% 到 67.7% 不等,这既提供了强有力的模型区分度,也为未来的进展留下了相当大的提升空间。对智能体轨迹的进一步分析揭示出冗长且前置集中的推理模式,凸显了从零构建可运行代码库所需的规划与系统级推理能力。
cs.SE / 164 / 2609.39284
EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses
EngramBench:一个面向技能演化框架的、以能力为基础的基准
large language model
大语言模型相关
Abstract
While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, current evaluations of skill evolution in autonomous agents suffer from a critical identifiability problem: they structurally confound genuine capability abstraction with rote solution leakage (i.e., copying highly similar code from historical training data). To resolve this, we introduce EngramBench, a rigorous, capability-grounded benchmark governed by the strict axiom of capability overlap without solution overlap. Comprising 30 diverse learning tasks and 13 unseen transfer tasks, EngramBench challenges agents to navigate interactive, multi-hour development cycles driven by LLM-simulated users. Our extensive evaluation across 48 multi-hour execution trajectories -- corroborated by human-expert validation -- reveals a profound insight into procedural memory. We demonstrate that static skill banks do not magically bypass the "last mile" of exact code implementation, which remains bottlenecked by the base model's inherent reasoning limits. However, they serve as an indispensable execution compass. By navigating agents away from catastrophic, token-heavy trial-and-error, genuine capability abstraction slashes redundant context bloat and reduces overall coding time by over 55%. Ultimately, EngramBench shifts the evaluation paradigm from trivial pattern matching to the verifiable measurement of deep, cross-domain capability transfer.
Chinese Translation
尽管大语言模型在孤立的代码生成中已取得了显著成功,但真实的软件工程需要持续的推理、复杂的状态管理以及连续的跨领域抽象。然而,当前对自主智能体技能演化的评估存在一个关键的可识别性问题:它们在结构上将真正的能力抽象与死记硬背式的解法泄漏(即从历史训练数据中复制高度相似的代码)混为一谈。为了解决这一问题,我们提出了 EngramBench,一个严格的、以能力为基础的基准,它遵循“能力重叠但解法不重叠”这一严格公理。EngramBench 由 30 个多样化的学习任务和 13 个未见过的迁移任务组成,它挑战智能体在由 LLM 模拟用户驱动的、交互式的、长达数小时的开发周期中进行探索。我们在 48 条长达数小时的执行轨迹上开展的广泛评估——并得到人类专家验证的佐证——揭示了对程序性记忆的深刻洞见。我们证明,静态技能库并不会神奇地绕过精确代码实现的“最后一英里”,这一环节仍然受限于基础模型固有的推理能力上限。然而,它们充当着不可或缺的执行指南针。通过引导智能体远离灾难性的、耗费大量 token 的试错过程,真正的能力抽象大幅削减了冗余的上下文膨胀,并将整体编码时间减少 55% 以上。最终,EngramBench 将评估范式从琐碎的模式匹配转向对深层、跨领域能力迁移的可验证测量。
cs.SE / 165 / 2609.39568
Self-Spec Verifiable Code Generation
自规范可验证代码生成
large language model
大语言模型相关
Abstract
Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable guarantees. Recently, researchers have proposed several benchmarks to evaluate the capabilities of LLMs in generating formally verifiable code, where LLMs need to formulate formal specifications, generate the corresponding code, and verify its correctness. However, existing benchmarks have two key limitations: (I) They primarily evaluate specification and code generation stage-wise, with code generation typically conditioned on an oracle specification. This setup overlooks whether strong stage-wise performance translates into end-to-end success. (II)They mainly focus on a single proof-oriented language and mathematically structured tasks, offering limited coverage of tasks common in software development. In this paper, we introduce VeriCodeBench, a benchmark for self-spec verifiable code generation, where the LLM relies solely on its own generated specification and code throughout the entire process. VeriCodeBench contains 400 language-native problems across C, Java, Rust, and Python, covering practical concerns in software development. We evaluate specification coverage, code validity, and joint problem-level success. We further introduce CodeNova to enhance the capabilities of LLMs in self-spec verifiable code generation. CodeNova makes requirements explicit through constraint-guided specification and uses verifier feedback to guide targeted implementation repairs. Experimental results reveal that self-generated specifications remain a major bottleneck, while providing more sophisticated specifications may not necessarily lead to higher verification success rates. CodeNova substantially improves performance across all evaluation metrics, enabling Claude Sonnet 5 to achieve the strongest results under the self-spec protocol.
Chinese Translation
大型语言模型(LLMs)可能会在测试遗漏的边角案例上生成不可靠的代码,而形式化验证可以提供机器可检查的保证。最近,研究人员提出了若干基准,以评估 LLMs 生成可形式化验证代码的能力,其中 LLMs 需要构建形式化规范、生成相应代码并验证其正确性。然而,现有基准有两个关键局限:(I)它们主要分阶段评估规范和代码生成,且代码生成通常以预言机规范为条件。这种设置忽视了强的分阶段性能是否会转化为端到端成功。(II)它们主要关注单一面向证明的语言和数学结构化任务,对软件开发中常见任务的覆盖有限。在本文中,我们介绍了 VeriCodeBench,一个用于自规范可验证代码生成的基准,其中 LLM 在整个过程中仅依赖其自身生成的规范和代码。VeriCodeBench 包含横跨 C、Java、Rust 和 Python 的 400 个语言原生问题,覆盖软件开发中的实际关注点。我们评估规范覆盖率、代码有效性以及联合的问题级成功率。我们进一步引入 CodeNova,以增强 LLMs 在自规范可验证代码生成方面的能力。CodeNova 通过约束引导的规范使需求显式化,并利用验证器反馈来指导有针对性的实现修复。实验结果表明,自生成规范仍然是一个主要瓶颈,而提供更复杂的规范未必会带来更高的验证成功率。CodeNova 在所有评估指标上大幅提升性能,使 Claude Sonnet 5 在自规范协议下取得最强结果。
cs.SE / 166 / 2609.39783
COMPASS: Predicting the Relationship of Multiple Patches for Vulnerabilities with LLMs
COMPASS:使用LLMs预测漏洞的多个补丁之间的关系
large language model
大语言模型相关
Abstract
Modern software heavily relies on code reuse, so upstream vulnerability fixes do not automatically propagate to downstream codebases. Downstream maintainers must manually adopt patches to eliminate known risks. In practice, a single vulnerability often corresponds to multiple patches, which greatly complicates downstream patch adoption because different patch relationships imply different adoption strategies. To address this challenge, we first manually inspect large-scale multi-patch vulnerabilities (about 1K) in the real world and interview experienced developers, summarizing six typical types of patch relationships, i.e., Merge, Mirror, Better Solution, Fixing-of-Fixing, Collaboration, and Separation. Based on these observations, we propose COMPASS, an automated approach that predicts the relationships of multiple vulnerability patches with large language models. Given a CVE as input, COMPASS follows a four-phase pipeline that (i) identifies the patch group and pre-scans explicit relationships, (ii) performs individual patch analysis, (iii) infers relationship instances via a hierarchy-guided prompt, and (iv) validates completeness and consistency of the inferred results. As output, COMPASS reports the predicted relationships within the patch group and visualizes them as a relationship graph. We evaluate COMPASS on a benchmark of 300 multi-patch CVEs and compare it against mainstream learning-based and LLM baselines. Results show that our method achieves strong and consistent prediction effectiveness and outperforms SOTA by 85.04% on average. We publicly release an online querying website to support community reuse of patch relationships knowledge: https://patch-relation.com.
Chinese Translation
现代软件严重依赖代码复用,因此上游漏洞修复不会自动传播到下游代码库。下游维护者必须手动采用补丁以消除已知风险。在实践中,单个漏洞通常对应多个补丁,这极大地复杂化了下游补丁采用,因为不同的补丁关系意味着不同的采用策略。为应对这一挑战,我们首先手动检查现实世界中的大规模多补丁漏洞(约1K个),并访谈经验丰富的开发者,总结了六种典型的补丁关系类型,即合并(Merge)、镜像(Mirror)、更优解决方案(Better Solution)、修复的修复(Fixing-of-Fixing)、协作(Collaboration)和分离(Separation)。基于这些观察,我们提出了COMPASS,一种利用大语言模型预测多个漏洞补丁之间关系的自动化方法。给定一个CVE作为输入,COMPASS遵循一个四阶段流水线:(i) 识别补丁组并预扫描显式关系,(ii) 执行单个补丁分析,(iii) 通过层次引导的提示推断关系实例,以及 (iv) 验证推断结果的完整性和一致性。作为输出,COMPASS报告补丁组内预测的关系,并将其可视化为关系图。我们在一个包含300个多补丁CVE的基准上评估COMPASS,并将其与主流的基于学习的方法和LLM基线进行比较。结果表明,我们的方法取得了强大且一致的预测效果,并平均超越SOTA 85.04%。我们公开发布了一个在线查询网站,以支持社区对补丁关系知识的复用:https://patch-relation.com。
cs.AI / 167 / 2609.38419
Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models
基于自适应增强和多分辨率集成模型的结直肠癌分割
large language model
大语言模型相关
Abstract
Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this cancer and increasing the survival rates. Although CRC is more common in developed regions, its occurrence is also increasing in developing regions as well. CRC diagnosis relies on histopathology assessment post-biopsy. Automated deep learning algorithms can significantly reduce diagnosis time, enhancing efficiency and supporting timely clinical decisions. We present an automated segmentation pipeline for whole-slide histopathology images that labels tumor grades 1-3 and normal mucosa. It utilizes dense prediction transformers with various encoder backbones, overlapping patches, and test-time augmentation. An adaptive augmentation policy, guided by large language models, further improves training. Top models were ensembled via soft voting, and mask refining post-processing steps, Gaussian blurring, morphological closing, and connected components analysis. On a colorectal cancer grade dataset, our method improved the F1 score from 62.92 to 69.84. Code is available here: github.com/caglarmert/ICIP2025
Chinese Translation
结直肠癌(CRC)是致死率第二高且发病率第三高的癌症,也是胃肠道癌症中的首要死因。早期诊断对于该癌症的治疗和提高生存率至关重要。尽管CRC在发达地区更为常见,但其在发展中地区的发病率也在增加。CRC诊断依赖于活检后的组织病理学评估。自动化深度学习算法可以显著缩短诊断时间,提高效率,并支持及时的临床决策。我们提出了一种用于全切片组织病理学图像的自动分割流程,该流程可标注肿瘤分级1-3和正常黏膜。它利用具有多种编码器骨干网络的密集预测Transformer、重叠图像块以及测试时增强。一种由大型语言模型引导的自适应增强策略进一步改善了训练。顶级模型通过软投票以及掩膜细化后处理步骤、高斯模糊、形态学闭运算和连通分量分析进行集成。在一个结直肠癌分级数据集上,我们的方法将F1分数从62.92提高到69.84。代码可在此处获取:github.com/caglarmert/ICIP2025
cs.LG / 168 / 2609.39531
WEIRDO: WEak resIdual Regularized DOob's h-transform diffusion alignment
WEIRDO:弱残差正则化的 Doob h-变换扩散对齐
diffusion
扩散模型相关
Abstract
We study the problem of estimating the guidance that steers the distribution learned by a diffusion generative model toward a tilted target $q_0 \propto w\,p_0$ at inference time. Relying on the stochastic optimal control approach, we observe that the exact drift correction is the gradient of the logarithm of Doob's $h$-function, and we study the problem of estimating it from a sample. In the present paper, we assume that the score of the pretrained model is available, that the tilting weight is bounded and positive, and that the reference distribution has a bounded support, no smoothness of the weight is required. Introducing a penalized least-squares risk in which the penalty is the residual of the space-time harmonicity equation satisfied by the $h$-function, measured in a dual Sobolev norm, we derive high-probability bounds on the squared error of the resulting guidance estimate. Since the penalty vanishes at the target, the estimator is free of regularization bias, and in favourable scenarios its rate of convergence is faster than the minimax rate of estimating first-order derivatives of a smooth regression function. Assuming that $w$ is bounded and positive with $\mathbb{E}_{p_0}[w^{-\mathrm{s}}] < \infty$ for some $\mathrm{s} \in (0,\infty]$, and that the reference data are compactly supported, we prove that the guidance is estimable in squared $L^2$ at rate $\varepsilon_n^{\mathrm{s}/(\mathrm{s}+4)}$, where $\varepsilon_n = n^{-2(β-1)/(2(β-1)+d)}.$ We also transfer the obtained bounds to the total variation distance between the marginals of the estimated and the exactly guided samplers, and illustrate the performance of the suggested approach with numerical experiments.
Chinese Translation
我们研究在推理时估计引导的问题,该引导将扩散生成模型学到的分布导向倾斜目标 $q_0 \propto w\,p_0$。基于随机最优控制方法,我们观察到精确的漂移校正就是 Doob 的 $h$-函数的对数的梯度,并且我们研究从样本中估计它的问题。在本文中,我们假设预训练模型的 score 可用,倾斜权重有界且为正,并且参考分布具有有界支撑;不要求权重具有光滑性。通过引入一种惩罚最小二乘风险,其中惩罚项是 $h$-函数所满足的时空调和性方程的残差,并在对偶 Sobolev 范数中度量,我们推导出所得引导估计的平方误差的高概率界。由于惩罚项在目标处消失,该估计量没有正则化偏差,并且在有利情形下,其收敛速度快于估计光滑回归函数一阶导数的极小极大速率。假设 $w$ 有界且为正,且对某个 $\mathrm{s} \in (0,\infty]$ 有 $\mathbb{E}_{p_0}[w^{-\mathrm{s}}] < \infty$,并假设参考数据具有紧支撑,我们证明该引导可以在平方 $L^2$ 中以速率 $\varepsilon_n^{\mathrm{s}/(\mathrm{s}+4)}$ 估计,其中 $\varepsilon_n = n^{-2(β-1)/(2(β-1)+d)}.$ 我们还将所得到的界转移到估计采样器与精确引导采样器的边缘分布之间的全变差距离,并用数值实验说明所提出方法的性能。
cs.LG / 169 / 2609.38438
A Pre-trained Variational Autoencoder for Gyrokinetic Plasma Turbulence Surrogate Modeling
一种用于回旋动理学等离子体湍流代理建模的预训练变分自编码器
diffusion
扩散模型相关
Abstract
Machine learning surrogate models offer a promising path toward accelerating plasma turbulence simulations. We present PreVAE-Turb, a surrogate modeling framework that leverages pre-trained variational autoencoders (VAEs) from the Stable Diffusion image generation model for efficient spatial compression of turbulence fields. The pre-trained VAE is fine-tuned on turbulence data using a physics-informed loss function that includes a spectral loss operating in Fourier space to enforce spectral accuracy across scales. The VAE is combined with convolutional long short-term memory (ConvLSTM) networks to learn temporal dynamics in latent space, with a manifold consistency error metric that monitors encode--decode consistency during autoregressive rollouts. We validate the framework on two-dimensional Hasegawa-Wakatani drift-wave turbulence and extend it to gyrokinetic turbulence from the GENE code, where a four-channel adaptation simultaneously predicts electrostatic potential, density, and parallel/perpendicular temperature fluctuations without requiring architecture redesign. Once trained, inference generates thousands of time steps in seconds on a single GPU, providing substantial computational acceleration compared to direct numerical simulation. The pre-trained approach offers a transferable methodology broadly applicable to various turbulence simulation codes.
Chinese Translation
机器学习代理模型为加速等离子体湍流模拟提供了一条有前景的路径。我们提出 PreVAE-Turb,一种代理建模框架,它利用来自 Stable Diffusion 图像生成模型的预训练变分自编码器(VAEs),对湍流场进行高效空间压缩。预训练的 VAE 在湍流数据上进行微调,使用一种物理信息损失函数,该损失函数包含一个在傅里叶空间中运行的谱损失,以在跨尺度上强制谱精度。该 VAE 与卷积长短期记忆(ConvLSTM)网络相结合,以在潜空间中学习时间动态,并使用流形一致性误差度量来监测自回归展开过程中的编码--解码一致性。我们在二维 Hasegawa-Wakatani 漂移波湍流上验证该框架,并将其扩展到来自 GENE 代码的回旋动理学湍流;其中,一种四通道适配同时预测静电势、密度以及平行/垂直温度涨落,而无需重新设计架构。训练完成后,推理可在单个 GPU 上于数秒内生成数千个时间步,与直接数值模拟相比,提供了显著的计算加速。这种预训练方法提供了一种可迁移的方法论,广泛适用于各种湍流模拟代码。
cs.AI / 170 / 2609.38364
Acceleration of Diffusion Language Model through Discrete Average Generator
通过离散平均生成器加速扩散语言模型
diffusion
扩散模型相关
Abstract
Discrete diffusion models and flow matching have emerged as powerful frameworks for generative modeling over discrete state spaces, yet efficient few-step generation remains a fundamental challenge. In this work, we introduce the Discrete Average Generator, a principled extension of MeanFlow to Continuous-Time Markov Chains (CTMCs). Analogously to how MeanFlow defines an average velocity field over a time interval in continuous spaces, we define an average generator as the normalized increment of the transition kernel over a time interval. We show that this average generator satisfies a self-consistency identity, which provides the foundation for our training objective. We further develop training strategies that align with the standard training paradigm of diffusion language models while keeping the resulting objective tractable. When projected onto per-coordinate marginals, the self-consistency identity admits a closed-form expression, enabling efficient training and inference. In Potts model simulations, our objective reduces the total variation distance of the $K$-step sampler by up to 67%. On OpenWebText, our method achieves the lowest generative perplexity among the evaluated methods for 8 to 64 sampling steps while enabling a $16\times$ acceleration, and achieves comparable performance to existing methods on ImageNet.
Chinese Translation
离散扩散模型与流匹配已成为在离散状态空间上进行生成建模的强大框架,然而高效的少步生成仍然是一个根本性挑战。在这项工作中,我们引入了离散平均生成器(Discrete Average Generator),它是 MeanFlow 到连续时间马尔可夫链(CTMCs)的一个原理性扩展。类似于 MeanFlow 在连续空间中定义时间区间上的平均速度场,我们将平均生成器定义为转移核在一个时间区间上的归一化增量。我们证明该平均生成器满足一个自洽恒等式,这为我们的训练目标提供了基础。我们进一步开发了与扩散语言模型标准训练范式相一致的训练策略,同时使所得目标保持可处理。当投影到逐坐标边缘分布上时,该自洽恒等式具有闭式表达式,从而实现了高效的训练与推理。在 Potts 模型模拟中,我们的目标将 $K$ 步采样器的全变差距离最多降低了 67%。在 OpenWebText 上,对于 8 到 64 个采样步数,我们的方法在所评估的方法中取得了最低的生成困惑度,同时实现了 $16\times$ 的加速;在 ImageNet 上则达到了与现有方法相当的性能。
cs.LG / 171 / 2609.39091
Steepest Guidance: A Practical and Principled Approach to Inference-Time Alignment of Flow and Diffusion-based Models
最陡引导:一种实用且有原则的流模型与扩散模型推理时对齐方法
diffusion
扩散模型相关
Abstract
Inference-time alignment of flow and diffusion-based models is critical for achieving flexible generative modeling. Theoretically, Doob's $h$-transform provides an elegant solution to this problem, and most existing methods are based on this principle. However, in practice, estimating the optimal guidance derived from Doob's $h$-transform at inference time is challenging. To deal with this issue, we regard inference-time alignment as a sequential optimization problem in the space of probability measures and propose a novel framework called *Steepest Guidance*, based on the principle of maximizing local improvement in the objective. We provide a theoretical analysis of the proposed method and demonstrate its effectiveness through extensive experiments.
Chinese Translation
流模型与扩散模型的推理时对齐对于实现灵活的生成建模至关重要。在理论上,Doob 的 $h$-变换为该问题提供了一种优雅的解决方案,而大多数现有方法都基于这一原理。然而,在实践中,在推理时估计由 Doob 的 $h$-变换导出的最优引导是具有挑战性的。为了解决这一问题,我们将推理时对齐视为概率测度空间中的序列优化问题,并提出了一种名为 *最陡引导* 的新框架,其基于最大化目标中局部改进的原则。我们提供了对所提出方法的理论分析,并通过大量实验证明了其有效性。
cs.LG / 172 / 2609.39326
Discrete Score Matching Enables Causal Discovery from Count Data
离散分数匹配使从计数数据中进行因果发现成为可能
diffusion
扩散模型相关
Abstract
Count data pose a challenge for score-matching-based causal discovery: derivatives are unavailable, and simply replacing them with finite differences does not generally suffice for causal discovery. We generalize SCORE's constant-curvature criterion (Rolland et al., 2022) by conditioning on the node's value, yielding the conditional curvature score (CCS) for ordering. We also extend curvature-based parent recovery through the off-diagonal curvature score (OCS), enabling directed acyclic graph (DAG) recovery with both scores constructed from score functions for continuous data and concrete scores for counts. In the bivariate setting, zero CCS exactly characterizes a semiparametric generalized linear model (GLM) conditional form in which the conditional family need not be specified in advance, unlike in classical GLMs. For bivariate semiparametric GLM DAGs under our regularity condition, canonical-parameter nonlinearity is necessary and sufficient for identifiability. In multivariate DAGs, this nonlinearity enables DAG recovery through CCS and OCS. Our framework identifies a new class of semiparametric GLM DAGs that strictly contains the nonlinear Gaussian ANM class identified by SCORE. We introduce DISCO (DIscrete SCOre), a count-DAG recovery algorithm that estimates CCS and OCS using discrete diffusion. Experiments demonstrate accurate DAG recovery across Poisson, negative binomial, binomial, and mixed-family settings, as well as scalability to 1,000-node DAGs on a single GPU.
Chinese Translation
计数数据对基于分数匹配的因果发现构成了挑战:导数不可用,而简单地将它们替换为有限差分通常不足以进行因果发现。我们通过在节点的值上进行条件化来推广 SCORE 的常曲率准则 (Rolland et al., 2022),从而得到用于排序的条件曲率分数 (CCS)。我们还通过非对角曲率分数 (OCS) 扩展了基于曲率的父节点恢复,使得能够进行有向无环图 (DAG) 恢复,其中这两个分数均由连续数据的分数函数和计数的具体分数构建。在双变量情形中,零 CCS 恰好刻画了一种半参数广义线性模型 (GLM) 条件形式,其中条件分布族无需事先指定,这与经典 GLM 中不同。对于在我们的正则性条件下成立的双变量半参数 GLM DAG,典范参数非线性是可识别性的充要条件。在多变量 DAG 中,这种非线性使得能够通过 CCS 和 OCS 进行 DAG 恢复。我们的框架识别出一类新的半参数 GLM DAG,它严格包含由 SCORE 识别出的非线性高斯 ANM 类。我们提出 DISCO (DIscrete SCOre),一种计数 DAG 恢复算法,它使用离散扩散来估计 CCS 和 OCS。实验表明,在泊松、负二项、二项和混合分布族设置中能够准确恢复 DAG,并且能够在单个 GPU 上扩展到 1,000 个节点的 DAG。
人工智能 (cs.AI)
165
cs.AI / 1 / 2609.38340
CARAT: Do Materials LLMs Reason or Recite?
Abstract
When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain rendering carries everything the question needs it is 1.96 points, and where it omits those fields entirely, 46.7 points. The headline mostly measures what the baseline lacked, not how evidence is presented. Third, we attack our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts sat near chance. The frozen model quotes that link yet answers the same when we redirect it, on 95.6% of paired cases: it repeats the relation without using it. After matched supervision it reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% the best shortcut reaches: both steps are learnable.
cs.AI / 2 / 2609.38346
Examining Variation in How Guided AI Tutors Resolve Student Impasses
Abstract
When a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse (i.e., wheel spinning). Generative AI tutors increasingly use guardrails restricting answer-giving, yet little is known about how such tutors behave once an impasse persists. We analyze 20,462 student turns from 1,260 authentic sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns of three major types: conceptual errors, expressed uncertainty, or help-seeking. We then used these impasses to simulate three tutoring conditions to study variation in AI tutor guidance through impasses: baseline, no-direct-answer, and guided tutor. For a sample of 150 impasses, prompt specificity changed pedagogy: a baseline tutor provided the answer directly in 50.7% of responses, a no-direct-answer tutor asked a follow-up question every time, and the guided tutor responded in a wide variety of ways depending on the context. We then analyzed impasse trajectories in authentic interactions, finding that each additional impasse turn lowered the odds of next-turn recovery by 12.7% (AOR = 0.873, p < .001), and early dropouts were caught in recursive concept elicitation before reaching execution. The benefit of questioning decayed as impasses persisted (scripted question x depth AOR = 0.78; follow-up x depth AOR = 0.83), whereas addressing the student's error grew more beneficial (AOR = 1.14); after a failed scripted question, repeating it was followed by recovery in 28.1% of cases, compared with 39.8% when the tutor addressed the error instead. For learning analytics, these findings identify impasse depth and type as observable, turn-level dialogue signals that analytics can use to trigger graduated, state-sensitive assistance in real time.
cs.AI / 3 / 2609.38359
Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks
Abstract
High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We propose a method to generate diverse high quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate latent conversation structure using a Gaussian mixture density over key interaction features (e.g., confusion episode dynamics, scaffolding directive balance) enables sampling expert strategies in proportion to their prevalence in the training data. Across two structurally distinct domains, tutoring and emotional support dialogues, our GFlow based synthetic data generation approach offers a better balance of fidelity, mode coverage and authenticity than reinforcement-learning and end to end LLM baselines, without copying training data. Evaluated on three downstream outcome prediction tasks, classifiers trained on synthetic GFlowNet generated conversations provide a stronger training signal than competitive synthesis baselines.
cs.AI / 4 / 2609.38369
Can an AI Agent Rediscover a Blaschke-Curve Invariant?
Abstract
We study generalized Blaschke curves as a controlled environment for AI-assisted mathematical rediscovery. For one fixed degree-four Blaschke product, an agent receives numerical coordinates of the six pair-lines determined by each of 80 boundary configurations. The target theorem is withheld from the task instructions. The saved research log reports rejected geometric hypotheses and a homogeneous cubic fitted to polygon sides. Its frozen coefficients predict 480 lines from 80 unseen parameter values, with a recorded RMS scale-free residual of $8.88\times10^{-17}$. Discovery-set diagonals provide an out-of-fit consistency check, not a fully held-out test. A separate one-configuration run reports insufficient evidence for invariance. A post-review deterministic degree-search baseline also recovers the cubic, so the experiment does not establish an advantage over polynomial fitting. We present this single-instance case study as a protocol for separating conjecture, numerical validation, and proof, with explicit limitations concerning agent metadata, prior knowledge, and reproducibility.
cs.AI / 5 / 2609.38372
Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
Abstract
A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks. We propose a framework close to recursive self-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held-out tasks are strictly separated, and we additionally evaluate on five out-of-distribution benchmarks never used during evolution. We frame the evolution process as deep-learning training with two stages, multi-task pretraining and continual training. Starting from a 49-line seed harness, the harness obtained at the end of the first stage improves the average score by 4.48 points on the in-distribution benchmarks and by 12.64 points on the out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw-Eval, one of the out-of-distribution benchmarks, further raises the score on that benchmark from 66.17 to 68.06, exceeding Codex. We also provide an in-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review.
cs.AI / 6 / 2609.38386
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
Abstract
Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted. We implement DLFP in vLLM and evaluate it with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry. On Qwen3-0.6B in BF16 on one A100 80 GB GPU, three paired 100-request trials reduce P99 inter-token latency by 24.8%, 30.1%, and 28.2% (mean 27.7%, paired 95% confidence interval 21.0% to 34.3%) with exact output agreement, no failures, and unchanged SLO compliance. The benefit is not free: mean P99 time to first token increases 34.8% while remaining inside the declared SLO. Crucially, the mechanism does not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel configuration. We trace the failure to an asynchronous scheduler-call interval that is only a proxy for completed GPU iteration time. This negative result defines the boundary of the contribution and motivates a completion-timed controller for concurrent CPU and on-device inference. We do not claim mobile-device performance; the present work is a reproducible proof-of-concept and generalization study.
cs.AI / 7 / 2609.38392
MetaPersona: Task-Grounded Synthetic Populations from Empirical Social Science
Abstract
Personas used to seed LLM social simulations face a cold-start problem: existing methods lack a principled basis for deciding which attributes to include and how to assign their values. As a result, synthetic populations may misrepresent the demographic composition, latent attributes, and dependency structure that shape downstream behavior. We introduce MetaPersona-DB, a dataset of 11,000+ empirical human-subjects studies annotated with task-relevant variables, reported relationships, and aggregate-level population statistics. Building on this resource, we propose MetaPersona, a framework that retrieves task-relevant evidence, constructs literature-derived persona dependency graphs, and samples synthetic populations from empirical priors linking demographics, latent attributes, and outcomes. Across three downstream case studies, three baselines, and three frontier models, results vary by task and model: MetaPersona performs strongly on misinformation belief and AI-tool sentiment, while results on income redistribution are mixed. It also reduces persona-construction cost to under $0.5 per task using GPT-5.2. Finally, we present MetaPersona-Studio, a prototype interactive interface for empirically grounded persona generation.
cs.AI / 8 / 2609.38397
SimTrace: Grounded Multimodal User Trajectories Generation for Online User Modeling
Abstract
Virtual clients offer a cost-effective approach to support applications such as A/B testing, recommender system development, and interface evaluation. However, building them requires access to large-scale, semantically faithful, fine-grained online user trajectories. These data are difficult to obtain because proprietary logs are subject to privacy restrictions and small businesses often lack sufficient traffic. Consequently, existing public datasets either abstract away fine-grained user interaction details or preserve rich context but remain platform-specific and small-scale. To address this gap, we propose SimTrace, a framework that generates faithful, fine-grained synthetic multimodal clickstreams through a computer-use client agent that is grounded in real user trajectories and the given web environment. SimTrace anonymizes real interactions and constructs a simulated twin of the given web environment, then uses both to generate synthetic interaction trajectories. Each action is paired with its corresponding web observations and user context, yielding a shareable alternative to confidential logs for developing computer-use agent-style virtual clients. We apply SimTrace to an e-commerce setting and evaluate both its fidelity and downstream utility. SimTrace outperforms competing baselines on 7 out of 8 fidelity metrics. Models trained on synthetic data achieve performance comparable to those trained on real data on downstream tasks such as purchase prediction and recommendation. For next action prediction task, augmenting real data with synthetic data further improves accuracy by 11.0% relative to training on real data alone. We release SimTrace as an open-source package to facilitate research on online user behavior modeling.
cs.AI / 9 / 2609.38411
A Competing-Hazards Systematization of Loss of Control in Autonomous Agents
Abstract
Leading AI developers have reported agents acting beyond their approved limits, which a United Nations panel described as an early warning of loss of human control. Yet incident reports and agent-safety evaluations describe these events differently, making it difficult to compare failures, trace risk across attempts, or separate agent behavior from the environment's role in allowing an out-of-scope action to succeed. To address this gap, we introduce a common framework in which each attempt ends in approved completion, safe stopping, scope escape, or continuation. We formalize the framework as a discrete-time competing-hazards model and derive escape probability within a retry budget, a model-conditional safe-budget limit, and conditions for estimation from execution logs. We audit 22 incident reports and 102 agent-safety evaluations published from January 2025 to September 2026 using primary sources. Six incidents involved tasks that could not be completed within scope, thirteen involved agents that continued rather than stopped, and five did not report stopping behavior. Developers' figures imply a task-level incidence ratio near 47 for out-of-scope coordination in never-solved versus solved tasks. Among evaluations, 87 recorded an out-of-scope effect or specification violation, 26 treated safe stopping as a first-class outcome, only 20 recorded both, and 79 merged budget exhaustion with failure. In 20 of 22 incidents, the environment allowed an out-of-scope effect, indicating that realized loss of control often reflected persistent agent behavior interacting with permissive boundary conditions; meanwhile, no evaluation reported all fields needed to estimate the full competing-hazards process from published evidence.
cs.AI / 10 / 2609.38445
AIM: Agentic Idea Management for Automated Research
Abstract
Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/
cs.AI / 11 / 2609.38448
Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles
Abstract
Open-ended LLM homogeneity can create false plurality when several systems appear to offer independent perspectives while returning the same familiar default. Single-pass answers obscure the distinction between agreement produced by a tightly constrained answer space, prompt-vocabulary echo, and broader answer spaces with stable alternatives beneath the surface. We introduce CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework that adapts free-list elicitation from cognitive anthropology to LLM ensembles. CHOIR repeatedly elicits ranked lists, clusters items into prompt-level concepts, and measures concept salience across models, prompt variants, and persona conditions. We evaluate CHOIR on Infinity-Chat 100, an external prompt bank from recent work on open-ended model homogeneity, and on a 27-question targeted diagnostic bank designed to isolate mechanism-level contrasts. On Infinity-Chat 100, CHOIR reproduces high surface agreement (93/100 prompts above chance) while separating narrow prompts from broad prompts with recoverable depth. Across targeted probes and the external prompt bank, base-model identity remains the strongest recoverable signature, and persona prompts shift surfaced concepts within base-model signatures. A source-blind ranking module prioritises rare-but-stable candidates for later inspection. CHOIR turns open-ended homogeneity into a diagnostic measurement problem by asking where models converge, why they converge, and what remains reachable under structured depth probing.
cs.AI / 12 / 2609.38460
NAQD Env: A benchmark for selective withdrawal in language agents
Abstract
Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after sufficient repair. We introduce NAQD-Env, a synthetic environment that evaluates these decisions against a deterministic reference policy over explicit evidence, authorization, and constraint dependencies. Eleven dependency families support evaluation on development structures, held-out families, and held-out combinations of structures. Metrics distinguish attempted violations from violations permitted by a simulated execution gate and jointly report policy agreement, task value, withdrawal, resumption, and event reporting. We evaluate three open-weight instruction-tuned models from two families under three prompt conditions on 350 frozen scenarios, yielding 3,150 model-prompt episodes before gate replay. Across the reported conditions, withdrawal recall is at most 0.06, no valid resumption is observed at eligible opportunities, and only one episode matches the complete reference policy. Under the NAQD prompt, Qwen2.5-7B has fewer unsafe-attempt episodes than Qwen2.5-3B and Llama-3.1-8B, but also completes less useful work and preserves unaffected actions less accurately. Exploratory supervised fine-tuning probes increase Qwen2.5-3B decision accuracy from 0.45-0.54 to 0.83-0.92; separate diagnostics reveal inappropriate withdrawal after curriculum omissions and a loss of event reporting. These results motivate evaluating selective withdrawal as a distinct component of agent reliability. The setting measures policy application with trusted structured inputs and does not establish real-world containment or source-verification ability.
cs.AI / 13 / 2609.38512
VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation
Abstract
Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests. The agent has sixteen tools backed by a stateful Stripe billing twin and the Apache Fineract loan engine, with account access blocked until caller verification succeeds. The tasks use public household electricity data and a policy based on Pennsylvania's residential billing rules. An LLM-as-a-verifier checks the agent's actions and spoken figures against explicit requirements. On a calibration run, it agrees with a code verifier on 99.1% of checks. Across fourteen voice stacks and three repeats per task, completion ranges from 17.3% to 44.7%. Grok Voice leads, and Gemini 3.8 Live and GPT-Live follow at about the same cost per call. Background television reduces pooled completion from 38.7% to 8.6%. The simulated caller often accepts an incorrect result because it hears the agent's words but cannot inspect its actions. These findings show why voice agents need evaluation across the whole call, including what they say, what they change, and how they handle competing speech.
cs.AI / 14 / 2609.38621
When Scientific Contradictions Are Lost in Translation
Abstract
Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint system and translate its constraints into scientific reports from different laboratories. One assignment satisfies more constraints, while another satisfies fewer but better matches expected biology. This creates a simple dilemma: does the model choose the assignment that best fits the constraints, or the one that better matches biological expectations? When the constraints are stated directly, GPT-5.6 Sol and Claude Opus 5 recover the best-supported assignment in 90% and 96% of cases, respectively. In scientific prose, however, the models behave differently. Claude Opus 5 often prefers the biologically expected assignment. Removing that biological preference increases recovery of the better-supported assignment from 27% to 79% (p<.001); recovery reaches 92% when the same Biology-favored record is accompanied by a formalization request and an explicit paired-design cue (p<.001). GPT-5.6 Sol is less sensitive, with neither corresponding change reaching statistical significance. These results suggest that reliable scientific verification depends not only on formal reasoning, but also on how models decide which findings should be compared and what relations they imply.
cs.AI / 15 / 2609.38639
Component-Aware Feedback for Self-Evolving Programs
Abstract
LLM-guided evolutionary search can discover complex programs, but existing methods mostly only save candidate programs and fitness scores while discarding which component edits produced which fitness metric changes. Existing methods force the mutator LLM to infer the effect of prior edits from cluttered histories, making program search slow and unstable. This is especially true for locally servable LLMs to evolve multi-component systems. We introduce component-aware feedback, which compares each evaluated program with its parent, identifies the components that changed, and logs them with the associated metric differences into an attribution memory that later mutations read. The memory keeps each change in two reference frames, local against the parent it came from and global against the seed program, which shows both the immediate effect of a change and the cumulative progress made since the seed. We study this on LLM reranking, a multi-objective optimization problem where a multi-stage pipeline must balance quality against serving cost. Across twelve \textsc{Bright} datasets, our method reaches the strongest baseline's final quality after a median of one third of the search budget and ends 7.2\% higher in held-out nDCG@10, and under a cost-aware objective it finds pipelines that are on average more accurate while using 11\% fewer tokens per query, showing component-aware feedback to be a promising direction for more efficient self-evolving systems.
cs.AI / 16 / 2609.38642
ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code
Abstract
Chart editing requires cross-modal edit grounding, realizing a requested visual change in the code that draws it, with necessary related updates and without altering unrelated content. Existing benchmarks emphasize either code executability or chart quality, but their metrics do not clearly distinguish request completion from missed coupled updates and gratuitous changes. We introduce ChartRevise, a structured dataset and evaluation protocol for exact program-grounded chart editing. For dataset construction, we build on the grammar of graphics to systematically cover chart-editing operations, using source-program checks to verify their applicability across chart types and libraries. To improve edit exactness, our pipeline checks individual requirements and guides repair or exclusion when they are unmet. The resulting dataset contains 92,438 records covering 344 edit types across 20 chart types and three plotting libraries. For evaluation, our reference-free protocol separately measures atomic requirement completion, identifies gratuitous changes, and detects missed coupled updates. These checks are combined with successful execution and rendering to determine exact-edit success. Across five models and four external benchmarks, fine-tuning yields relative gains of 16\% in mean requirement recall and 22\% in mean exact-edit rate.
cs.AI / 17 / 2609.38670
Where Scientific Search Agents Fail: Decision-Checkpoint Auditing of Exposure and Inspection Attempts
Abstract
Final-answer accuracy does not reveal whether a scientific-search agent failed to encounter a target paper, attempt to inspect it, or return an accepted answer after inspection. We introduce decision checkpoints that record observations and tool actions without benchmark labels during inference, then join target identities and evaluator labels to assign outcome categories from recorded events. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieves 24.6\% accuracy, compared with 17.8\% for raw search. The keyword condition has fewer incorrect answers with neither target exposure nor inspection, but more incorrect answers after the target is exposed and left uninspected. Compared with keyword search, read-first has 27.4\% more recorded evidence-search calls. Target inspection attempts occur on 199 questions under read-first and 191 under keyword search; both conditions achieve 24.6\% accuracy. The checkpoint protocol makes these question-level differences explicit, distinguishing target exposure and inspection from aggregate accuracy and total tool use.
cs.AI / 18 / 2609.38684
Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status
Abstract
Knowledge-intensive language-model systems typically represent external knowledge as text chunks or static graphs, with limited support for concept evolution, point-in-time reasoning, and distinctions between validated and inferred knowledge. We introduce the Concept Lifecycle Model (CLM), which represents concepts as persistent, graph-grounded, temporally versioned entities with explicit provenance and epistemic status, and Concept-Grounded Attention (CGA), which injects concept-graph structure into transformer computation through graph-biased self-attention (Form A) and gated cross-attention over concept nodes (Form B). We evaluate the framework in controlled settings using disabled-mechanism baselines. On 200 MuSiQue and HotpotQA questions with retrieval fixed, concept-graph retrieval recovers explicit multi-hop paths but does not improve evidence recall. Form A appears to steer attention, with 2.76 times more attention on gold than distractor concepts, but the same ratio occurs when Form A is disabled; the learned bias is negligible and no answers change. An identity-preserving Form B improves F1 from 0.188 to 0.221, but control concepts yield 0.213, indicating that most of the gain reflects added capacity. On LongMemEval, explicit temporal representation improves answer accuracy by 13 to 25 points across all tested generators, up to 122B parameters, while simplified CLM version resolution performs similarly to dated serialization because concept identity is not established reliably. On a synthetic source-independence task, protocol-derived epistemic status reduces unsupported assertions from 28% to 0.1% in a fine-tuned small model and from 19-68% to 0-5% in 72-122B models. Overall, the results support making temporal validity and epistemic status explicit, while showing that graph-attention diagnostics are not informative without disabled-mechanism controls.
cs.AI / 19 / 2609.38690
GATE-ST: Gene-Aware Text-image Encoder for Spatial Transcriptomics
Abstract
Spatial transcriptomics enables spatially resolved gene expression analysis from slide-level images while preserving morphological features, providing valuable information for studying disease mechanisms and developing treatments. However, spatial gene expression profiling typically requires expensive and time-consuming tests. While existing image-based prediction optimizations mostly revolve around including positional embeddings and further image-based changes, text-based optimizations remain relatively unexplored. We present GATE-ST, which incorporates text-based inputs into image-based spatial gene expression predictions. With this approach, generated text descriptions of genes are utilized to better spatial transcriptomics prediction results. Gene summaries are put through a text encoder, generating embeddings that integrate with image embeddings through cross-attention layers to align with morphological features. We demonstrate the effectiveness of such text inputs by benchmarking performance against random gene embeddings and multiple other image-text fusion architectures, and show that GATE-ST outperforms these alternatives. Our results demonstrate the effectiveness of GATE-ST in pathology imaging, which may greatly reduce the time and cost of accurate spatial transcriptomic predictions, proving the potential of text-guided spatial gene expression prediction.
cs.AI / 20 / 2609.38699
Budget Boundary Effects in Test-Time Mathematical Reasoning
Abstract
A cumulative token cap can fall inside a mathematical derivation, forcing a test-time controller to choose between stopping at the cap (strict) and allowing the current attempt to finish (advisory). We measure this boundary choice with paired offline replays of 19,200 public traces: 120 AIME, BrUMO and HMMT problems and two archive configurations of one model. Candidate order and a 16-attempt cap are fixed, and answer selection is blind to reference answers and correctness labels. Three findings emerge. First, at the 4k cap, most advisory accuracy gains replace abstention with a correct answer; strict stopping pays for an unfinished prefix that the completed-only selector cannot use. Second, comparisons along realized cost differ from same-cap comparisons: advisory 4k in low has higher accuracy than strict 8k at comparable mean completion cost, while in high its observed accuracy is 0.42 points below strict 32k using 59% of its mean tokens. These aggregate comparisons do not establish equal-compute superiority or accuracy equivalence. Third, increased candidate coverage does not guarantee higher answer accuracy: a log-probability selector loses accuracy while coverage rises, including after a source-grade consistency repair. Same-cap majority-accuracy differences shrink below 1.3 percentage points at 32k. Budget curves should jointly state the cap, realized cost, eligible candidates, stopping rule and selector information.
cs.AI / 21 / 2609.38712
Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
Abstract
Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.
cs.AI / 22 / 2609.38733
Code to Control: Synthesizing Parameterized Reactive Controllers
Abstract
Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.
cs.AI / 23 / 2609.38743
Learning to Route in Visual Space via Multi-Step Embedding Retrieval
Abstract
LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5\% to 76.3\%. In agentic search, it improves task success rates by 52.7\% and reduces the average token length by 61\% from 1886 to 728, whereas upgrading the agent yields only a 3.7\% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by $23\times$ and cutting the cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.
cs.AI / 24 / 2609.38766
PathAnchor: Path-Structured Evidence for Scientific Agents
Abstract
Scientific agents can retrieve relevant passages yet still lose functional order, mix evidence across sources, or state conclusions that exceed the retrieved record. We introduce PathAnchor, a bounded scientific reasoning system built on path-structured evidence workspaces. Instead of treating passages or extracted concepts as independent units, the system retrieves source-linked Material-Sensor-Signal-System trajectories that preserve role, direction, and the evidence supporting each transition. A controller uses three read-only tools to search paper-specific trajectories, trace paths across candidate sources, and open exact evidence before producing a claim-cited answer and an explicit evidence boundary. On 120 single- and cross-paper flexible-sensor questions, PathAnchor scores 82.6% and leads six evaluated systems. Under a matched controller, corpus, and six-call budget, replacing unordered concept graphs with path-structured records raises source recall from 61.3% to 82.9%, increases answers whose claims all cite opened evidence from 69.2% to 90.0%, and reduces tool calls. These results show that evidence organization affects retrieval and citation completeness under fixed agent resources.
cs.AI / 25 / 2609.38778
Action Conditioned Bisimulation For GUI Agent Memory
Abstract
An agent that remembers what it did on a web page must decide when two pages count as the same. Memories built on observation similarity merge pages that look alike but behave differently, and GUIs are full of such pages: two tabs of one widget or two rows of one menu answer the same click differently. We define the merge rule as an action-conditioned bisimulation over the empirical predictive state graph a frozen agent fills as it acts. Two states merge only when their shared actions lead to agreeing outcomes and successor blocks under an affordance label. Observation similarity never enters the rule, and nothing is trained. It replaces the merge rule of an existing outcome-value memory, so a closed-loop comparison isolates it. On MiniWoB++ it raises success rate over a memoryless agent, while a control taking identical exploratory detours, the prior successor-representation merge, and the same criterion without action conditioning change nothing.
cs.AI / 26 / 2609.38782
Persona and Persuasive Framing in AI Voice Agents: A $2\times2$ Field Experiment with Children
Abstract
Conversational agents increasingly interact with children, yet evidence on how their design shapes children's susceptibility to persuasion comes almost entirely from the lab. We report a $2\times2$ randomized field experiment embedded in a public German Santa Claus telephone hotline. Children's calls were randomly routed to one of four LLM voice agents varying persona (Santa, high authority, vs. Helper, low authority) and framing (persuasive nudges toward prosocial wishes vs. neutral). Of 1,072 logged calls, 89 conversations (median age 6) met inclusion criteria. Persuasive framing raised the probability of a prosocial wish from 11.6% to 45.7%, robust to controls. Persona authority showed a near-zero effect: Santa did not outperform the Helper. Persona instead shaped engagement; children hung up on the Helper far more often within the first minute (65% vs. 39%). Where context already lends an agent legitimacy, how it speaks shapes children's compliance more than who it claims to be.
cs.AI / 27 / 2609.38788
Positive Ratings, Hidden Concerns: Employee Voice Disclosure in AI-Mediated Organizational Listening
Abstract
Organizations started listening to employees through conversational AI agents alongside structured surveys. Little is known about what these channels change in what employees say when disclosure carries hierarchical risk. We report a field study inside a global management consulting firm whose process pairs a pre-survey with an adaptive AI voice interview on the same themes within one session. Across 44 first-session interviews (132 matched theme observations), 20-41% of sessions showed a favorable rating co-occurring with a substantive concern voiced later, depending on the favorability threshold. The Gioia analysis drew on 158 protective quotes from 65 eligible sessions. Disclosure rarely arrived unguarded: employees softened concerns, deflected accountability, and bounded how far they went, and this protective work tracked the perceived legitimacy of the listening structure. We develop a grounded model of bounded disclosure and derive four propositions for voice, channel and listening research. Silence, we argue, can persist inside expression.
cs.AI / 28 / 2609.38817
When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning
Abstract
Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However, existing mitigations largely truncate long outputs or react to surface repetition, and thus fail to distinguish normal thinking from uncontrolled reasoning or explain how benign reasoning degenerates into harmful behavior. In this paper, we operationalize LRM generation as four states and further introduce Reasoning-state Analysis via Dynamic Attention Responses (RADAR), which identifies the current reasoning state in real time and characterizes how effective reflection can develop into uncontrolled generation. Guided by RADAR's analysis, we further realign abnormal attention distributions toward patterns observed in normal requests and examine how this correction affects excessive reflection and persistent looping. Temporal analyses show that uncontrolled reasoning is characterized by attention distributions that deviate from normal generation, with abnormal trends becoming detectable before repetition begins. Correcting these deviations through Attention Realignment consistently reduces looping while largely preserving benign performance. Together, RADAR provide a mechanistic account of how reasoning becomes uncontrolled, offering actionable guidance for identifying critical failure stages and designing targeted runtime interventions.
cs.AI / 29 / 2609.38827
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models
Abstract
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8\% of all predictions and 51.3\% of errors to Neutral despite 74.95\% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76\% of the effective gold support, versus 87--102\% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from $K=2$ to $14$; utilization falls for every model and reaches 26--75\% at $K=14$, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47\% to 86\% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia
cs.AI / 30 / 2609.38829
Diversity Combining for Multi-Path LLM Reasoning
Abstract
Multi-path reasoning methods such as self-consistency (SC) sample $K$ reasoning paths and choose the most frequent answer. However, their gains quickly plateau as $K$ increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem from wireless communications: each path is a noisy channel observation, and the pairwise correlation of path correctness caps the design-effect effective sample size of the vote at a finite ceiling. Generalized least squares (GLS) analysis shows that, under exchangeability, the optimal symmetric linear combiner of latent embeddings is uniform, supporting majority vote as the natural default in standard SC while leaving room for weighting or pruning under heterogeneous prompt-template branches. Across 5 models and 12 benchmarks, prompt-template diversity reduces path correlation in $55$ of $57$ valid cells, with the strongest effect on open-ended QA. We derive an Adaptive-K rule that uses a four-path pilot to select $K^*$, retaining $96$--$103\%$ of MV@$K{=}32$ accuracy across Math, QA, and NLU.
cs.AI / 31 / 2609.38850
OpenJev-RLCD: A Working RLCD Implementation
Abstract
Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \gvTwoCovFive\% of the items at $\le$5\% error, versus \gvGrpoCovFive\% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: https://github.com/ZimmyGao/openjev-rlcd.
cs.AI / 32 / 2609.38866
When Context Changes: Understanding Update Failures in LLMs
Abstract
As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still recover the updated value when the model answers with an old one, pointing to a failure to select information that remains available. Component tests in Qwen and Pythia identify a mechanism for this selection failure: attention drift, where attention favors old values over the current one when producing an answer. We study a one-layer transformer to mathematically understand how this phenomenon happens: when attention scores are similar, several old values can together receive more attention than the current value. Guided by this explanation, we redirect attention toward the current value without further training. When the current value is requested directly, adjusting this intervention for each input corrects most old-value errors across various model families while preserving nearly all initially correct answers. Reliable context management therefore requires more than remembering updated information: models must use it to guide their answers.
cs.AI / 33 / 2609.38891
Consistent Plan-Act for Long-Horizon Agentic Tasks
Abstract
Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agents for structured state assertions and compare their reports programmatically to detect explicit contradictions. Our analyses reveal systematic disagreement about the same task-relevant state facts, a phenomenon we term planner-actor state mismatch. We further find that providing agents with task-relevant state information reduces mismatch and improves coordination and task performance. Based on the systematic analysis of the state mismatch, we propose Consistent Plan-Act (ConPAct), which feeds detected contradictions back to both agents to form consistent state interpretations and fine-tunes them on curated consistent interactions for better coordination. ConPAct improves performance across various environments and model configurations, e.g., increasing MiniGrid success rate from 38.6% to 54.4% with GPT-5.6-sol/terra as planner and actor respectively, demonstrating that state consistency can guide both inference-time correction and coordination training.
cs.AI / 34 / 2609.38914
Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets
Abstract
Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We instead formulate evaluation as a sequential allocation problem. Given a fixed trial budget and a set of scenarios whose failure behavior is unknown, which scenarios should be run, and run again? We propose a risk-aware contextual Thompson Sampling policy that combines a pre-execution scenario context vector and a fixed impact score with the failure outcomes observed during evaluation, and we test it by offline replay over 70 $τ$-bench airline scenarios and 824 recorded trials. Our main result is at the smallest budget: with only 50 trials ($6\%$ of the corpus), the policy recovers $86\%$ of the impact-weighted failures an oracle could find, compared to $25\%$ for uniform allocation. It discovers $3.5\times$ more impact-weighted failures (215.4 vs. 62.2) with the same number of trials, delivers $5\times$ the discovery per dollar, and cuts the budget wasted on scenarios that never fail from $34\%$ to $2.8\%$. The rest of our analysis demonstrates and qualifies this result: a budget sweep shows the advantage shrinks as the budget approaches the corpus size, and paired significance tests show that scenario context helps mainly at small budgets while posterior-based exploration helps at moderate ones. Risk-aware adaptive allocation therefore helps most exactly where evaluation budget is scarcest.
cs.AI / 35 / 2609.38917
How Much Can Reliability Drift Under a Fixed Confidence Distribution?
Abstract
A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliability relation under covariate shifts that preserve the distribution of the confidence score, constraining the reweighting within each confidence level by a $χ^2$ budget; the resulting worst case, as a function of the budget, is a fragility profile. On an interval of budgets that can be computed from the source distribution, the profile equals exactly the square root of the budget times the within-level variance of the correctness propensity -- the grouping-loss term of calibration-refinement decompositions. Beyond this interval the profile is governed by the tails of the propensity law, and the entire upward profile determines the centred within-level law; consequently, calibration residual and grouping variance do not determine fragility in general, though they do when labels and predictions are deterministic. Since the propensity is not observed, we restrict reweightings to a learned finite readout within confidence bins, bound the part the restriction misses by the grouping variance remaining inside readout cells, estimate the restricted profile with role-separated labels, and provide a separate split-sample lower confidence bound. On ImageNet this bound is positive in both splits for four of six primary classifiers and nine of twelve additional ones as released, and for three of eighteen after temperature scaling. Held-out drift under optimised reweightings fitted without evaluation labels tracks the estimated profile; an exploratory label-permutation diagnostic yields near-zero agreement for this statistic while largely reproducing the correlation observed for unsigned random reweightings.
cs.AI / 36 / 2609.38925
Prototype-guided Bilateral Alignment Multimodal Federated Learning
Abstract
Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing methods predominantly rely on idealized assumptions of model homogeneity and balanced modality distributions, rendering them ill-suited for practical scenarios characterized by heterogeneous client architectures and severe modality imbalance. To address these challenges, we propose a \textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment (MFedPBA) framework. MFedPBA facilitates robust knowledge synergy through a dual alignment mechanism: (i) at the feature level, it aligns heterogeneous feature spaces via a projection encoder optimized by contrastive learning and the Gromov-Wasserstein distance; (ii) at the decision level, it employs an entropy-weighted aggregation of naturally aligned logit prototypes. This novel design achieves robust MFL by jointly tackling heterogeneous feature spaces and collectively aggregating decisions. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines under conditions of model heterogeneity and modality imbalance.
cs.AI / 37 / 2609.38929
Learning What to Forget: Distributional Unlearning for LLM Representation Spaces
Abstract
Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing analyses often impose parametric assumptions to obtain tractable selection rules. These assumptions may be poorly suited to high-dimensional language-model representations. We introduce \textsc{Mamushi}, a framework for non-parametric distributional unlearning that ranks forget examples using a probabilistic classifier whose Bayes-optimal logit equals the forget-to-retain log-density ratio (up to an additive class-prior constant). We show that thresholding the population log-density ratio yields the optimal fixed-budget selection rule for our removal--preservation objective and establish a non-asymptotic transfer guarantee relating score-estimation and threshold-calibration errors to degradation from the population-optimal selection rule. Our empirical evaluation spans real-world datasets on toxic-language removal and topical-domain removal regimes using different representations, with \textsc{Mamushi} achieving a more favorable removal--preservation trade-off than other baselines. Our work shows that \textsc{Mamushi} can serve as an efficient selection approach for downstream machine unlearning procedures, reducing the number of forget examples required to reach a fixed forgetting target.
cs.AI / 38 / 2609.38956
Routing Probes Can Improve Without New Information: An Exact-Null Audit of Uncertainty Beyond Model Outputs
Abstract
Routing signals of modern vision transformers -- expert gates, attention-residual weights and halting scores -- often improve probes that predict whether the model is correct, and the improvement is commonly read as evidence that routing carries information about errors beyond the model's outputs. We test this inference directly: keeping real output-routing pairs, we redraw correctness labels from a frozen output-only generator fitted on disjoint data, so that routing is uninformative by construction. Under this exact label null, a width-matched MLP comparison still reports a routing gain in 51.3% of confidence-only evaluations (308/600), while a linear comparison reports none. Holding each training trajectory fixed on a six-model panel and selecting the checkpoint by validation log loss instead of validation accuracy removes the detections (50/120 to 0/120, and 83/120 to 0/120 in an independently implemented probe), identifying accuracy-based checkpoint selection as the cause; across all output views the raw detection rate falls from 27.5% (528/1,920) to zero observed detections. The repaired comparison is not sensitive, detecting an implanted signal of about 0.005 nats in 0/20 replicates in each of two matched settings, whereas a conditional permutation test built on an estimated routing law detects it in 11/20 and 10/20 and rejects rarely under the null. On real correctness labels, the conditional analysis yields model-relative evidence in five DeiT attention-residual families; in four it persists under two specified variants of the conditional law, and no family passes an additional noise criterion. Fitting a better probe and testing for incremental information are different problems, and each needs its own validation.
cs.AI / 39 / 2609.39005
C-STRIDE: An Observation-Driven AI Digital Twin for Predicting Basin-Wide Flood Fields from Sparse Stream-Gauge Histories
Abstract
Emergency managers need to know where floodwater is, how deep it is, and how it will change over the coming hours across an entire river basin. During a flood, however, real-time measurements come from only a handful of stream gauges, and high-resolution hydrodynamic models are too costly to rerun each time new data arrive or to run as large ensembles. We present C-STRIDE, an observation-driven AI digital twin that turns short records from a few stream gauges, together with terrain and rainfall, into basin-wide maps of water depth and extends these predictions up to a day ahead. It is trained on simulations from a calibrated two-dimensional hydrodynamic model and needs no separate data-assimilation step. In the Des Plaines River basin near Chicago, six gauges inform predictions over 4.2 million 30-m grid cells. Terrain improves the predictions most, rainfall keeps errors from growing over longer horizons, and together they reduce errors by about 40% compared with gauge records alone. When future rainfall is known, errors remain near 15% one day ahead, compared with nearly 40% without rainfall. Given real instead of simulated gauge records, the model shifts its predictions toward the observed hydrographs at three of six gauges without retraining, and it runs about 150 times faster than the hydrodynamic model. These results show how sparse gauges, terrain, and rainfall can be combined into fast, continuously updated flood predictions, a step toward operational flood digital twins that still requires testing with real-time data and rainfall forecasts.
cs.AI / 40 / 2609.39026
Search Shapes Conclusions: Auditing Evidence Selection Bias in Deep Research Agents
Abstract
Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation, which we call the candidate pool. Early findings redirect later queries, document choices, and stopping, so the documents an agent reads form a selective sample. Existing evaluations rarely account for this selection. We formulate the problem as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts each candidate document's evidence direction and corrects the candidate-pool average using the logged probabilities of selecting each document and reaching each search round. Shrinkage stabilizes short searches, while intervals replace point estimates when some documents cannot be sampled. We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read. The latter requires intervention. On questions from the MS2 systematic-review benchmark, CESS reduces mean absolute error against the candidate-pool average by $9.2\%$ and reduces the estimate's change under opposing document rankings by $39.4\%$ relative to averaging the evidence scores of documents read. Across trajectories from a public Open Deep Research agent, the corresponding reductions reach $60.1\%$ and $87.2\%$. A further 4,800 trajectories under paired interventions confirm that correcting a pool estimate and measuring a policy effect are different tasks. CESS therefore audits whether the evidence direction underlying a report reflects the documents available for evaluation, while a separate intervention analysis measures the effect of search decisions.
cs.AI / 41 / 2609.39076
Multi-LLM Collaborative Alignment via Stackelberg Games
Abstract
A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.
cs.AI / 42 / 2609.39101
Beyond Prediction: Steering VLM Agents with Retrospective World Modeling
Abstract
Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on what will happen next and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution $P(\hat{a}{t}|s_t, s{t+1})$ for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.
cs.AI / 43 / 2609.39106
Reserve-Aware Contrast Certificates for Conservative Bandits with Uncertain Baselines
Abstract
Conservative bandits must improve an incumbent policy without exhausting a prescribed performance budget. When the incumbent is uncertain, separately bounding candidate and baseline rewards can charge twice for shared estimation error. We develop Reserve-C4B around the baseline-relative contrast itself. A shared confidence set yields an exact expression for this avoidable penalty and a tighter admissibility test at every fixed history. A reserve ledger separates statistical evidence from permitted performance deficit; a prefix-refresh extension recertifies accumulated decisions under the current confidence set without discarding previously certified credit. For linear rewards, self-normalized confidence sets provide simultaneous validity over time and adaptively generated candidates, and the resulting policy satisfies a conditional-mean performance constraint with high probability. Reproducible experiments isolate certificate coupling, prefix refresh, and historical information, showing large reductions in baseline fallback while exposing the limitations of frozen certificates.
cs.AI / 44 / 2609.39139
BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence
Abstract
Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which action should follow. Existing methods often use such signals as separate triggers, making it difficult to preserve a coherent evidence state across a trajectory; we call this problem evidence-state fragmentation. We introduce BELIEFRAG, a closed-loop controller that updates an explicit state over sufficiency, reliability, conflict, uncertainty, evidence gaps, and acquisition cost, then chooses among retrieval, query rewriting, verification, answering, stopping, and abstention. Across six QA benchmarks with gpt-oss-120b, BELIEFRAG reaches mean token F1 0.572 with 3.89k tokens per question, outperforming fixed iterative retrieval (0.555 F1) while using 39% fewer tokens. The same quality-cost pattern transfers to Qwen3-32B, where BELIEFRAG reaches 0.552 F1 versus 0.523 for iterative retrieval while using 35% fewer tokens. Analysis shows that the main gains come from corrective re-retrieval rather than pruning alone, while several belief dimensions are redundant and calibrated answerability plays the strongest operational role. Calibration improves threshold stability across related evidence sources, although source shift can still invalidate the same decision signal.
cs.AI / 45 / 2609.39140
Schema: Discovering Unknown Environments via Agentic Program Induction
Abstract
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
cs.AI / 46 / 2609.39143
RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent
Abstract
Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels. We propose RefCon, which combines sequential self-refinement with parallel self-contrast to extract higher-quality memories without gold labels. Evaluated on AppWorld and BFCL-V3 across multiple context-evolving agent frameworks, RefCon delivers strong and consistent gains, including relative improvements of 21.6% on ACE and 16.6% on ReMe over no-scaling baselines, while a diversity-focused variant (DivCon) achieves a 35.5% gain on ReasoningBank. RefCon consistently outperforms existing baselines without ground-truth labels, and generalizes across model scales and to software engineering tasks, where it surpasses even ground-truth baselines. We further analyze the accuracy-token trade-off and scaling behavior, showing RefCon maintains favorable efficiency and continues to improve as more trajectories are used, unlike diversity-only scaling which saturates earlier.
cs.AI / 47 / 2609.39148
Do Self-Evolving Skills Generalize to Held-Out Tasks?
Abstract
AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifacts, that can be retrieved and reused when solving new tasks. Self-evolving skill methods keep rewriting these skills after each round of practice on training tasks, and the skill is then used on new tasks of the same kind. We ask a question: does the improvement a skill shows on its training tasks carry over to new test tasks? We test five self-evolving methods and a one-shot skill on six benchmarks, with the same model, the same agent, and the same train/test split for every method. Of the 21 skills that improve on their training tasks, 5 keep all of that improvement on the test tasks, 13 keep part of it, and 3 keep none of it. No existing method is best everywhere. When we read the skills, the ones that carry over badly often fix details that should depend on the task, such as column names and output files, or turn a fix for one failure into a rule for every task. An LLM judge that reads the skill content can often see this: it ranks finished skills the same way the test results do in 86\% of pairs. But it predicts the effect of a single edit poorly, so edits still have to be tested by running them. Based on these findings, we describe Generalizable Skill Optimization (GSO), which keeps only a guide for writing skills and writes a new skill for each task; it scores highest on all six benchmarks.
cs.AI / 48 / 2609.39166
Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
Abstract
Persistent spatial memory enables embodied agents to navigate familiar environments across repeated visits. However, targets may move while unobserved, including during navigation, making remembered locations unreliable by the time an agent arrives. Despite advances in memory retrieval and state prediction, accounting for continued hidden world evolution and revising beliefs under limited visibility remain challenging. We study Evolving-World Navigation, where agents infer target locations from intermittent observations, predict their states at inspection time, and revise beliefs using visual evidence. We propose EvolvingNav, which constructs a time-indexed belief from timestamped 3D object histories through a structured persistence-relocation model. The belief distinguishes persistence at the last observed location from relocation to alternative locations and retains probability mass outside the known candidate set. An event-driven filter propagates the current belief as time elapses, forecasts target occupancy at candidate inspection times, and incorporates new RGB-D evidence. Negative observations downweight location hypotheses according to calibrated, visibility-conditioned detection probabilities, while evidence tracking prevents repeated use of the same observations. A frozen, zero-shot vision-language controller uses the updated belief to choose actions and replan. We further introduce EvoWorld-Bench, a benchmark grounded in human activity traces, comprising 54 scenes and 803,680 tasks with controlled changes before and during navigation. In simulation and real-robot experiments, EvolvingNav improves navigation success and search efficiency over the evaluated baselines. Paired experiments show the clearest gains under learnable temporal patterns, while ablations demonstrate the value of preserving uncertainty and incorporating visibility-aware evidence.
cs.AI / 49 / 2609.39220
Learning Process Rewards via Reasoning State Propagation
Abstract
Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independently, providing no explicit mechanism for effectively using final outcome to guide the learning of intermediate reasoning states. We introduce Reasoning State Propagation (RSP), which represents each reasoning prefix with a binary validity state and models transitions between successive states across the reasoning trajectory. Specifically, RSP predicts a break probability that a valid state becomes invalid and a repair probability that an invalid state returns to valid. By propagating these transitions, RSP connects intermediate states to the final state, allowing process annotations to supervise intermediate states while outcome labels supervise the final state and can provide learning signals to preceding steps. Across reasoning search, response selection, and reinforcement learning, RSP consistently outperforms representative PRM baselines, with average improvements over Qwen2.5-Math-PRM of 5.6% in beam search and 2.1% in reinforcement learning.
cs.AI / 50 / 2609.39228
Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization
Abstract
We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A language model constructs structured evidence over local correspondences, omissions, scope, and logical relations, while a deterministic validator checks this evidence and produces reproducible judgments. When a substantive but admissible deviation is accepted, FYAN requires an explicit proof-transfer obligation connecting the formal statement back to a source-facing interpretation. With the same model (DeepSeek-V4.1-Flash) in every stage, FYAN proves 86 of 143 FormalTCS theorems under a strict Lean check, against 69 for a general agent harness, and raises the natural-language proof score from 0.501 to 0.851. On ConsistencyCheck, its semantic audit catches more inconsistent statements than a direct LLM judge, both on labels verified against the source (recall 0.777 vs. 0.636) and on the original labels (0.873 vs. 0.820), and localizes each mismatch it reports to a specific hypothesis, conclusion, or scope. FYAN also built ODENumLib, a 9,355-line Lean library for the numerical analysis of ordinary differential equation.
cs.AI / 51 / 2609.39259
Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers
Abstract
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediate computational states are functionally redundant when they induce similar downstream responses. We introduce Conditional Functional Substitutability (CFS) to directly characterize such functional substitution. CFS exposes functional relations and reduction potential missed by conventional importance- and similarity-based measures. Across modalities and Transformer families, CFS reveals systematic functional reorganization with scale. Controlled scaling further shows that performance gains need not track growth in substitutability, while fixed-capacity models with more independent functional structure perform better, providing a functional account of diminishing returns. Predicted CFS further enables dynamic computation with a better performance--computation trade-off than importance-based component selection, suggesting new directions for redundancy-aware computation and more efficient model scaling.
cs.AI / 52 / 2609.39341
Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models
Abstract
Does a language model merely predict tokens, or does it understand what it says? We make this question measurable by defining "understanding" through the lens of no-arbitrage. A model understands a vocabulary to a certain degree if a computationally bounded trader cannot extract guaranteed profit by betting against the model's probabilities on logically related claims (a "Dutch book"). We establish three theoretical results: first, because full logical coherence is computationally intractable, understanding is inherently graded, not absolute. Second, we prove that the exact optimum of standard next-token prediction is inherently incoherent across different question formats; the flaw lies in the training objective, not the architecture. Third, we show that uncertainty accumulates predictably along reasoning chains, making unjustified overconfidence an arbitrage opportunity in itself. To address this, we introduce Arbitr, a training framework where an adversarial trader penalizes the model for logical inconsistencies, paired with a calibration anchor to prevent uninformative collapse. Across five pre-registered experiments on Qwen2.5 and Phi-3.5 models, we demonstrate that standard models are highly exploitable across different phrasings. Arbitr reduces this exploitability by orders of magnitude without sacrificing task accuracy, and the effect successfully transfers to unseen logical patterns and new model families. Crucially, we uncover a scaling illusion: at 7B parameters, near-zero measured incoherence often coincides with extreme, unjustified confidence. We conclude that while Arbitr enforces rigorous logical consistency, coherence is a necessary condition for knowledge, but not a sufficient one
cs.AI / 53 / 2609.39351
On the Complexity of Preference-Based Bandits
Abstract
We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $κ$, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address these limitations, we consider general reward function classes and introduce the \emph{locally sensitive eluder dimension}, a novel complexity measure tailored to the logistic structure of preference feedback that yields fine-grained regret guarantees without unfavorable dependence on $κ$. Building on this notion, we propose \textbf{GINOP} (Generic INformative OPtimism), an algorithm that constructs log-loss confidence sets and jointly selects arm pairs to balance optimism and informative exploration. We establish a first-order regret bound that, in contrast with what previous results suggest, demonstrates that learning with preference feedback is as statistically efficient as learning from direct reward observation. Finally, we corroborate our theoretical findings with empirical evaluations against competitive baselines.
cs.AI / 54 / 2609.39360
Autoresearch in Mixed-Integer Linear and Nonlinear Programming
Abstract
Despite recent progress in autoresearch, applying it to practical operations research problems, typically formulated as NP-hard mixed-integer linear or nonlinear programs (MILPs or MINLPs), remains challenging because effective research requires systematically managing competing ideas and long-horizon experimental trajectories. We introduce AutoMIP, a reusable agent skill for organizing long-horizon autoresearch in mixed-integer programming through idea pooling and algorithm tree search. AutoMIP maintains a persistent pool of complementary candidate ideas while organizing executable experiments into an algorithm tree, enabling the agent to preserve unexplored hypotheses, refine promising algorithms, and switch to alternative methodological directions based on historical states. On MILP and MINLP benchmark cohorts, AutoMIP achieves the highest final success rates among the evaluated autoresearch frameworks. On MIPLib, AutoMIP discovers new best solutions for 31 of 60 instances, surpassing existing autoresearch frameworks. On MINLPLib, it achieves new best solutions for 52 of 60 instances. Ablation studies further demonstrate the complementary contributions of idea pooling and algorithm tree search, highlighting the importance of jointly maintaining diverse research ideas and structured experimental trajectories for long-horizon autoresearch.
cs.AI / 55 / 2609.39371
EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning
Abstract
In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.
cs.AI / 56 / 2609.39392
Experimental Experience Modeling for Autonomous Research
Abstract
Autonomous research agents can generate hypotheses and conduct experiments, but experimentation remains a major source of computational cost. A fundamental challenge is deciding which experiments are worth running, particularly when prior evidence is insufficient to resolve uncertainty. Yet current research agents lack a systematic way to leverage experimental experience when making such decisions. We introduce Experimental Experience Modeling (EEM), a framework for making informed experimental decisions by acquiring, reusing, and accumulating experimental experience. EEM extracts decision-relevant records from earlier experimental trajectories, distills them into reusable experience, and organizes them in an experience library. For a new experimental decision, EEM retrieves relevant historical experience and assesses whether it provides sufficient support for deciding whether a candidate direction warrants further investment. When historical experience is insufficient, EEM conducts a targeted, low-cost pilot experiment to acquire the missing decision-relevant experience on demand. It then combines this newly acquired experience with retrieved historical experience to determine whether the direction warrants full-scale evaluation, which requires substantial resources. The resulting experimental outcomes are further distilled into reusable experience, allowing the library to continually grow through iterative accumulation. Experiments on autonomous research benchmarks show that EEM improves research performance while reducing model interaction overhead, demonstrating the value of reusing accumulated experience and acquiring additional experience only when needed.
cs.AI / 57 / 2609.39402
Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
Abstract
Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.
cs.AI / 58 / 2609.39473
Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning
Abstract
Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.
cs.AI / 59 / 2609.39494
Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention
Abstract
Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting conclusions. We formalize a unifying framework to encompass all self-distillation methods vs classic supervised fine-tuning: we train every combination of the three axes, on Qwen2.5-7B and Ministral-3-3B across ordinary and contradictory tasks, totaling 1,200 adaptation runs, to systematically investigate the impact of the above axes. We propose a controlled model of the same objective to explain the resulting acquisition-retention trade-offs. We find that (i) the rollout source matters mostly where the task contradicts the pretrained behavior: there teacher rollouts raise acquisition well above what student rollouts achieve, with almost no change in retention; (ii) the teacher coupling changes acquisition most, on every task: acquisition rises with the coupling rate, then falls past a task-specific rate; (iii) switching the KL direction costs retention in one model but not the other so which axis to tune first depends on the model. The controlled model reproduces the three trends.
cs.AI / 60 / 2609.39518
Referential Uncertainty in Human--AI Collaboration
Abstract
Effective human-AI collaboration requires partners to establish references through interaction, which becomes fragile when descriptions are ambiguous, similar referents compete, or partners see different things. We study referential uncertainty - uncertainty over which candidate object a description refers to - in a collaborative puzzle task where a human Helper instructs an AI Worker to place pieces. The Worker must identify and communicate its uncertainty, and the Helper must recognize and act on it. We show that a separately elicited belief distribution over candidate pieces is better calibrated (ECE 0.15) and better discriminates correct from incorrect placements (AUROC 0.65) than raw action-token probabilities, which are severely overconfident (0.97 mean confidence, ECE 0.44). Across three frontier vision-language models (GPT-4.1, GPT-5, GPT-5.5), this elicited uncertainty rises predictably with instruction vagueness, but not with competing referents in context, even when those increase errors. The models seldom externalize it, asking for clarification on only 3.5-16.7% of turns. In a controlled human study (N=210), participants given only the Worker's default message accept 78% of wrong placements and cannot tell right from wrong (AUC 0.50). Precise descriptions and, especially, well-targeted hedges cut wrong-move acceptance to 36% while largely preserving correct-move acceptance, compensating for missing shared awareness such as not seeing the Worker's action. But this benefit depends on targeting: a deployable hedge derived from the model's own belief entropy inherits that signal's weakness and can do more harm than good. Externalized uncertainty helps a human partner only when it is accurately targeted.
cs.AI / 61 / 2609.39544
Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean
Abstract
Recent achievements in AI-assisted mathematics require intensive interaction of agents with proof assistants to generate machine-checked proof certificates. Agents interact with proof assistants such as Rocq or Lean through an interface that controls what the agent receives from the prover and the cost of these interactions. Today, these interfaces are adapted from tools designed for humans and not optimized for agents. We propose an evolutionary method where a frontier model incrementally proposes new features and only keeps the ones that improve the overall performance of smaller models. We demonstrate the effectiveness of our method by growing, on a curated set of mathematical problems, \rme, a new MCP server for the Rocq prover. On the held-out \texttt{test} split of miniF2F-Rocq, an agent equipped with \rme outperforms both the baseline that only exposes the Rocq compiler and an established MCP server, across four models from two families, in success rate, cost per solve, and time per solve. Although evolved for Rocq, the resulting server transfers to Lean, improving cost and time per solve on a subset of PutnamBench. We release \rme and its port to Lean.
cs.AI / 62 / 2609.39551
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
Abstract
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.
cs.AI / 63 / 2609.39559
Divide and Collapse: MAPF-Collapse via Exact Decomposition into Independent Sub-Instances
Abstract
In this work we study the problem of MAPFC, a post-optimization step for Multi-Agent Path Finding (MAPF) plans where we are given a feasible plan produced by a modern MAPF solver and are tasked with removing avoidable moves while preserving feasibility. This NP-hard problem naturally arises when using learning-based state-of-the-art (SOTA) solvers which construct plans that contain redundant moves that can be removed. Recently, Tang et al. presented Judgelight, which uses Integer Linear Programming (ILP) to solve MAPFC. Importantly, the ILP is constructed over all agents jointly, so its cost is governed by the full instance rather than by the small coupled residue that actually requires joint reasoning. Our key insight, motivating this work, is that MAPFC instances naturally decompose into independent sub-problems, most of which involve a single agent and can be solved without any inter-agent reasoning. To this end, we first identify which agents need to coordinate their motion and partition the instance into sub-problems accordingly. For the cases where no coordination is required, we introduce an extremely lightweight solver that is $\approx\!1{,}900\times$ faster than Judgelight. For cases where coordination is required, Judgelight can be used but we introduce an alternative CBS-like solver which is more efficient on easier problems. The resulting framework is exact, uses no commercial ILP solver, and matches Judgelight's quality while running substantially faster on the coordination-light majority of instances; on the coordination-heavy instances we propose a regime-aware hybrid planner that falls back to Judgelight. Over all benchmarks tested, this planner achieves a median $10.5\times$ per-instance speedup over Judgelight.
cs.AI / 64 / 2609.39564
A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
Abstract
Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.
cs.AI / 65 / 2609.39579
AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation
Abstract
Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.
cs.AI / 66 / 2609.39665
ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning
Abstract
Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
cs.AI / 67 / 2609.39701
Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
Abstract
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.
cs.AI / 68 / 2609.39702
A helps B while B hurts A: directed transfer in instruction-tuning mixture
Abstract
Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negative; the second, that it is symmetric. We show that both assumptions fail: task $A$ can help task $B$ while $B$ hurts $A$, so helpfulness is a signed property of ordered source--target pairs. We introduce the transfer map, a signed estimate of how much each source helps or hurts each held-out target. We fit the map in hundreds of fine-tuning runs on Qwen3 and Mistral models from 0.6B to 32B parameters, with all sources drawn from one corpus and no training examples from the target. The map predicts a held-out target's accuracy on unseen mixtures: recorded before those runs, its predictions have less than half the error of a mixture-agnostic baseline. The map is specific to its target and corpus but transfers across model scale: a mixture selected in advance at one size beats training on all source tasks at every other size we tested. Transfer is thus a property of the data. The map selects the tasks that help and drops the one that interferes: accuracy on the reasoning targets (causal explanation, multi-hop questions and methodological critique) rises by up to 14 percentage points over training on all source tasks.
cs.AI / 69 / 2609.39714
ArchitectureIQ: On the Measure of Training Intuition
Abstract
Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to predict the recipe yielding the best test metric. Overall, we find that LLMs' model intuition is good but has four limitations: (1) The intuition is imperfect, or even sub-human in some cases. Frontier models achieve around 76% accuracy (random choice 33%) vs best human researcher (66.0%), yet remain far from perfect. For architecture-only questions, best human achieves 65% while GPT-6 Astra only has 38%. (2) The intuition is empirical, not structured, supported by the fact that more CoT compute does not lead to substantial improvement. Unlike math, we still lack a "Science of AI" language that enables structured reasoning on AI. (3) The intuition is not maximally condensed, and can be further compressed into a knoledge base. Our constructed knowledge base with only 20 items yields large gains for weak models: GPT-4o equipped with the accumulated knowledge almost matches the performance of Claude Opus 5. (4) The intuition is insensitive to dataset properties, but the best model should in general depend on data properties. This suggests that data is the real "dark matter" in AI -- LLMs (so do human researchers) understand too little about data, even less than model architectures.
cs.AI / 70 / 2609.39717
Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents
Abstract
Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.
cs.AI / 71 / 2609.39727
OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation
Abstract
Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.
cs.AI / 72 / 2609.39788
Safety of Latent Communication in Multi-Agent Systems
Abstract
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole.
cs.AI / 73 / 2609.39838
Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
Abstract
Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded reasoning (reasoning in an illegible but unconcealed format), have already been shown to emerge under training pressures that occur in real pipelines, such as reinforcement learning against monitors. This suggests that steganographic reasoning too might arise as an unintended side effect of training. Here, we compare how easily models learn steganographic reasoning and these two neighbouring capabilities across three elicitation methods: reinforcement learning, in-context learning, and supervised fine-tuning (SFT). For most tasks, models learn steganographic reasoning only under SFT, while they learn steganographic messaging and encoded reasoning under all three elicitation methods. Even under SFT, steganographic reasoning requires at least twice as much training as messaging, and for several model-task combinations it is not learned at all. However, on a cover task that makes hiding information especially convenient, steganographic reasoning can be successfully learned under all three elicitation methods. Steganographic reasoning is thus much harder than steganographic messaging and encoded reasoning, and learning the latter two does not imply learning the former. Yet it lies within reach: an easy version is learned under every elicitation method, when the cover task is convenient for hiding information.
cs.AI / 74 / 2609.39868
Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding
Abstract
Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure mode in which lower generalized passenger time coexists with incomplete passenger journeys.
cs.AI / 75 / 2609.39869
GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
Abstract
Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-supervised rewards derived from its own likelihoods: a direct reward, measuring how likely the constrained output is given the input, and a reverse reward, measuring how well the input can be reconstructed from the generated output. We optimize these rewards with a Reinforce Leave-One-Out (RLOO) objective over groups of grammar-constrained rollouts, augmented with the top-1 beam-search hypothesis and regularized towards a frozen base model. We evaluate GrammarRL on sign language gloss translation, hierarchical text classification, and named entity recognition using Llama models ranging from 1B to 8B parameters. GrammarRL consistently outperforms constrained greedy decoding, with an average improvement of 9.8 points and gains of up to 22.8 BLEU. It matches or outperforms beam search on two of the three tasks while preserving greedy-decoding inference cost. Ablations further show that the two rewards are complementary: either reward alone can underperform the untrained baseline, whereas their combination consistently improves upon it.
cs.AI / 76 / 2609.39903
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
Abstract
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
cs.AI / 77 / 2609.39933
ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible
Abstract
When designing machine learning models, desirable properties are often in tension: improving one behavior can impair another, so task progress can depend on alleviating the conflict. LLM-based AutoResearch systems, which iteratively edit model code and retain edits based on scalar task-performance feedback, have largely ignored this trade-off. We find that scalar feedback supports broad exploration early in search, but it does not reveal how edits affect competing behaviors. In matched-budget experiments, introducing competing-behavior feedback as task gains diminish increases the share of proposals that improve both behaviors and sustains progress beyond scalar-only plateaus. Obtaining this feedback for a given model requires identifying its competing behaviors and designing probes to measure them. To make competing-behavior feedback actionable, we introduce ConflictGuide. Its reusable ConflictGuide-Skill combines a literature-grounded taxonomy with model-specific evidence to identify competing behaviors and specify probes for a code agent to implement as metrics. Evolution proceeds in two stages: Stage I explores with task feedback; Stage II uses probe feedback to steer proposals toward conflict alleviation and retains marginal-gain edits only when probes indicate sufficient alleviation. Across five diverse model families, ConflictGuide reduces task and conflict-related errors by up to 28% and 14%, respectively, relative to scalar-only AutoResearch, with gains extending to other code agents.
cs.AI / 78 / 2609.39955
Coverage Before Control: Route-Instruction Grounding and Steering for Controllable Retrosynthesis
Abstract
Single-step retrosynthesis models are commonly evaluated by their ability to recover recorded reactions. In practice, chemists may need to choose among several precursor sets for the same product, for example to preserve a particular motif. Recovering a recorded answer alone does not establish this ability to follow a preference. Satisfying such requests requires both coverage of relevant alternatives and control over which alternatives are favored. We introduce Route-Instruction Grounding and Steering (RIGS), a two-stage framework for instruction-conditioned retrosynthesis. Stage A trains a language projector, teaching it which alternatives an instruction favors or discourages. Stage B uses the projector learned in Stage A to steer a frozen generative model through lightweight residual adapters. We construct nested one-to-many training supports by pairing each product with increasing numbers of candidate precursor sets. Extensive experiments demonstrate that broader support helps the model generate a wider range of alternatives, and RIGS can learn to guide generation according to instructions. The relationship between coverage and control is consistent across model scales but non-monotone.
cs.AI / 79 / 2609.39958
Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking
Abstract
Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering changes and assess the resulting decks, raising the question of whether higher scores reflect better documents or changes in grading. In shared-session text-only grading with template markers removed, five judges score the complete system 20.4 to 33.6 points out of 95 above the same model generating directly from a short prompt. Every judge scores the system higher on all seventeen development deliverables. Margins against direct Opus generation from a short prompt range from -4.7 to +0.8 points. Judges agree on broad progress across development rounds but agree less on final-deck rankings than on pooled scores. Repeated grading also shifts scores on unchanged decks, making small improvements difficult to distinguish from judge variability.
cs.AI / 80 / 2609.39964
AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC
Abstract
Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense~6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.
cs.AI / 81 / 2609.39989
What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents
Abstract
Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. This critical scoping review maps 348 studies and examines 90 comparison records: 88 from 40 included studies and two from supplementary studies. Eight purposively selected cases structure the synthesis around the replaced decision, executed conditions, measurement comparability, controls, and remaining explanations. Results. Of 222 studies reporting local decision metrics, 142 also report measured task endpoints and 49 report proxies. These counts identify studies that report both types of measurement, without establishing that the measurements come from matched comparisons. Outcome Monitors reports a package-level completion gain whose attribution to detector quality remains limited; First-chunk selection reports a local improvement assessed against an offline proxy endpoint; Evidence-Carrying Termination reports fewer premature unsupported terminations and completion non-inferiority, without establishing completion superiority. Cross-case analysis identifies three candidate mechanisms involving recovery and disruption, intervention timing, and downstream use. Attribution and deployment depend on the comparison controls, label definitions, and information available to the controller. Conclusions. The review distinguishes the task-level benefit of a component replacement from the contribution of local decision quality and derives eight claim-specific reporting items. Neither online execution nor simultaneous gains in local and task metrics alone establish that better local decisions explain the task-level gain.
cs.AI / 82 / 2609.40027
Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents
Abstract
Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification argument and a one-sided lower confidence bound. One such verifier, CIVeX, reports zero false executions on a confounded tool-use benchmark. We red-team it by corrupting only the committed graph. Omitting a single bidirected edge takes it from zero false executions to 15.3% at the benchmark's published confounding strength, with 91% of its executions harmful and utility falling from +2.27 to +0.35. Reversing one arrowhead, so that a mediator is committed as a confounder, gives 48.9% false executions and no correct ones. Every one of these actions carries an internally valid certificate. An attestation step that tests each observationally certified execution against a bounded randomised sample detected both attacks, with 2 false alarms in 555 executions on a truthful graph; refusing what fails the test, or cannot be tested, gave zero false executions in every setting we measured. It does not restore beneficial execution: at the published strength 97.1% of beneficial actions are still never executed, because the same misspecification rejects them before attestation runs. Those rejections carry certificates too, and auditing them works, but its cost scales with the number of rejections rather than the number of executions. Recovering safety costs 127 experiments per 1,050 actions; recovering the lost value costs 614 more, at which point the audited verifier makes the honest graph's decisions on every instance and spends exactly its experiment budget. An audit that inspects only executions protects against wrongful action. Wrongful inaction has to be paid for separately.
cs.AI / 83 / 2609.40090
PTNO: Training Neural Operators with Noisy Monte Carlo Estimates for Particle Transport Problems
Abstract
Particle transport under multiple scattering is central to radiative transfer and plasma physics, yet high-fidelity Monte Carlo (MC) simulations must trace prohibitively many particles. Learning-based surrogates can amortize this cost, but typically train on expensive, well-converged MC solutions. We propose the Particle Transport Neural Operator (PTNO), a neural operator that learns particle transport surrogates directly from noisy, low-cost MC labels. Such labels pose two challenges: (1) high variance, which destabilizes standard supervised learning, and (2) a high dynamic range (HDR) spanning many orders of magnitude. For the first, we learn the solution operator from noisy labels of many configurations, amortizing MC cost and generalizing to unseen configurations. Because MC labels are unbiased, we show that the squared loss on them shares its minimizer with the loss on converged solutions, and our budget-allocation study over training scenes $M$, MC samples per render $N$, and independent renders per scene $K$ shows that many noisy scenes beat fewer converged ones. For the second, a nonlinear transform such as the logarithm biases noisy supervision. Instead, PTNO keeps labels in physical space and enforces positivity with a softplus output layer that represents small values effectively. We further train with a pointwise relative $L_2$ loss (PRelL2), the stop-gradient relative loss of HDR denoising and neural rendering, which normalizes each residual by the stop-gradient prediction instead of the noisy label. We demonstrate PTNO on neutron transport in fusion reactors and radiative transfer in participating media. On the two neutronics tasks, PTNO is $10^4$-$10^5\times$ faster than converged MC on the same CPU and $10^3$-$10^5\times$ cheaper than MC at matched accuracy; on the two radiative-transfer tasks, MC at matched accuracy costs $0.8$-$11\times$ as much as PTNO.
cs.AI / 84 / 2609.40111
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Abstract
An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.
cs.AI / 85 / 2609.40115
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.
cs.AI / 86 / 2609.40169
Learning from Research: Toward Lifelong Agent Harness Evolution
Abstract
Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.
cs.AI / 87 / 2609.40269
Belief-Aware Multi-Agent Path Finding under Map Uncertainty
Abstract
Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances. When such changes are spatially correlated, an observation can inform traversability estimates beyond the observed location. Prior approaches address uncertainty in traversability through contingent plans or replanning based on direct observations, but do not leverage this spatial dependence to infer the traversability of nearby unobserved locations. As a result, they cannot use one observation to anticipate nearby unobserved obstacles that may cause costly rerouting later. We focus on Belief-Aware MAPF, where map discrepancies are fixed during execution but initially unknown, and observations can be informative beyond the observed location. We propose Multi-Agent Gaussian belief Inference for Coordination (MAGIC), a framework that updates a shared belief about traversability online based on agents' observations. MAGIC uses a Gaussian Markov Random Field and Gaussian Belief Propagation to approximately infer traversability and construct detour-aware costs for standard MAPF planners. Our experiments on MAPF benchmarks show that MAGIC reduces the executed sum of costs compared to existing approaches on 96.3% of instances, across several planner families and teams of up to 800 agents, demonstrating its applicability to large-scale MAPF problems.
cs.AI / 88 / 2609.40285
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Abstract
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
cs.AI / 89 / 2609.40324
Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
Abstract
We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conjectures, overcoming subtle technical obstructions, and retaining intermediate progress over a long horizon. Cogentic addresses these challenges through an iterative prove--verify loop in which an orchestrator allocates a population of independent provers across distinct proof directions, subjects their output to adversarial verification by several specialized components, and promotes confirmed intermediate results into a persistent verified ledger that later rounds build on. The harness is designed to be able to solve research-level math and theoretical computer science problems. Using Gemini as the base model, Cogentic produced novel results on five open problems across online learning, auction theory, and mechanism design. Each result was independently verified by domain experts and is developed in full in companion papers. We list these results, and new ones as they are verified, at https://sites.google.com/view/cogentic .
cs.AI / 90 / 2609.40325
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
Abstract
As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.
cs.AI / 91 / 2609.40330
Turbo Harness: Instance-Adaptive Harness Optimization
Abstract
Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing information generated during the original optimization process. Specifically, Turbo Harness recycles artifacts produced during a completed global harness optimization run, and summarizes them into a structured playbook. We train a harness editor to leverage this prior optimization experience to generate instance-specific patches to the global harness. At inference time, the editor uses the instance and the playbook to construct a tailored harness in which the execution model operates. Through numerical experiments, we show that Turbo Harness consistently outperforms existing harness optimization baselines across seven benchmarks spanning interactive agent tasks, software engineering, and long-horizon terminal tasks.
cs.AI / 92 / 2609.38479
Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing
Abstract
This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at $r=0.262$, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44\% stronger on the supervised benchmark did not produce measurable improvement in the field ($r=0.250$, $p=0.84$), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1\,km, reaching a median increase of 12.78\% in low-risk route length for a median detour of 2.73\% on journeys of 3 to 6 km.
cs.AI / 93 / 2609.38578
Retargeting Motions to Diverse Skeletons via Learnable Flattening
Abstract
Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by $43-47\%$ over current benchmarks. A user study ($n = 37$), including expert animators, further ranks our approach highest in motion alignment and physical plausibility ($p < 0.05$). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.
cs.AI / 94 / 2609.38607
After a Decade: Bringing Shadow Removal into the Real World with Agentic Training Data
Abstract
Shadow removal looks nearly solved on established benchmarks, yet remains brittle in the real world. Models have advanced; the paired training data they rely on have barely changed in nearly a decade. The reason is simple: obtaining a shadow-free target requires removing the occluder while keeping the scene, camera, and illumination otherwise unchanged, making diverse paired data difficult to capture. Meanwhile, large shadow detection datasets already contain diverse real-world images and masks, but no shadow-free targets. To turn this abundant but incomplete data into paired supervision, we propose an offline agentic workflow combining physics-motivated generation, failure detection, feedback-driven retry, candidate selection, and deterministic correction. Using this workflow, we construct AgenticShadow, a dataset of 17,138 image-mask-target triplets spanning general scenes, faces, and remote sensing. Our construction workflow reduces Color Distribution Difference by 50.5% over previous shadow removal work, while training existing shadow removal models on AgenticShadow reduces cross-domain LAB RMSE by 19.7-37.5%.
cs.AI / 95 / 2609.38615
Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation
Abstract
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.
cs.AI / 96 / 2609.38637
Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking
Abstract
Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi-modal tracking where paired modalities are simultaneously available, cross-modal tracking requires localization when template and search frames originate from different active modalities. Accordingly, we introduce TSDA-Track, a Template-Search Domain Adaptation framework to reduce modality discrepancy during training. We investigate two feature alignment strategies. Pre-AFA TSDA-Track applies adversarial alignment before transformer's template-search interaction to suppress modality-specific bias. Enc-CFA TSDA-Track applies contrastive alignment to encoder representations after interaction to strengthen target-level cross-modal correspondence. Both variants retain a shared inference pipeline without modality-specific branches. Experiments on LasHeR, and zero-shot evaluations on RGBT234 and GTOT under multiple cross-modal protocols demonstrate improvements over representative state-of-the-art trackers. For instance, under the modality-switch protocol on RGBT234, Pre-AFA TSDA-Track achieves an SR/PR of 43.2/56.0, compared with 36.8/50.0 for ToMP-101 baseline. In addition, a study on Anti-UAV-024 further verifies the applicability of TSDA-Track to aerial tracking. Our study highlights the effectiveness of feature alignment domain adaptation for cross-modal tracking.
cs.AI / 97 / 2609.38683
Unveiling the Value of Motion for Cinematic Camera Trajectories
Abstract
Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.
cs.AI / 98 / 2609.38716
SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
Abstract
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
cs.AI / 99 / 2609.38717
Soft Spatial Reasoning
Abstract
Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at https://github.com/rafiibnsultan/Soft_Spatial_Reasoning
cs.AI / 100 / 2609.38810
CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models
Abstract
As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at GitHub repository.
cs.AI / 101 / 2609.38930
On the Relaxation of Conditional Independence Assumption for Image Segmentation
Abstract
In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which ignores crucial label correlations and therefore degrades performance in ambiguous or low-contrast scenarios. However, accounting for full label dependence is computationally prohibitive, requiring $\mathcal{O}(d^3)$ time. To address this, we replace the CIA with a Spatially Localized Dependence (SLD) structure that captures local label correlations while keeping the dependence model tractable. We further overcome the remaining computational bottleneck via a Reciprocal Moment Approximation coupled with a novel fixed-point optimization strategy that eliminates exhaustive search. The proposed algorithm achieves a highly practical $\mathcal{O}(d \log d)$ complexity and consistently outperforms conventional argmax and CIA-based RankSEG across diverse segmentation benchmarks. Improvements are significant in low-contrast or small-object scenarios, where label dependence offers valuable signals complementary to image information for accurate segmentation. The code of experiments is available at https://github.com/ZixunWang/RankSEG-DEP.
cs.AI / 102 / 2609.39033
TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization
Abstract
CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-normal evidence, having the same anomaly evidence. We propose TED (Text-Axis Evidence Decomposition), a post-hoc scoring method that asks whether each ambiguous response is better supported by source defect patches or by source normal patches mistaken as anomalous. TED compares these supports under the host's normal-versus-anomaly text response, leaves the backbone and prompts unchanged, and requires no target-domain training. It works as a train-free score for raw VLM backbones or as a source-calibrated residual correction for adapted CLIP-AD hosts. Across frozen VLM backbones, TED substantially improves pixel-level localization over raw prompt similarity; across adapted hosts, it improves most pixel-level settings over P-AUROC, P-PRO, and P-AP. Gains are largest under stronger hard-FP competition, with mean localization gain increasing from +5.0 in low-competition regimes to about +10.9 in mid/high-competition regimes. These results suggest that recoverable defect evidence can already exist in pretrained multimodal representations, but reliable localization requires decoding it against hard-normal competitors. Code will be released at TED GitHub repository.
cs.AI / 103 / 2609.39182
MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models
Abstract
World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.
cs.AI / 104 / 2609.39227
Emergent Multi-View Geometry Through Self-Distillation
Abstract
Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.
cs.AI / 105 / 2609.39363
Rethinking Multi-Image Re-Representation in Multi-Image Understanding
Abstract
Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at https://github.com/gengyuanmax/Mosaic.
cs.AI / 106 / 2609.39429
Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
Abstract
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
cs.AI / 107 / 2609.39490
OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
Abstract
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-visual joint reasoning, together with time-stamped clue chains that guide the annotation of thinking process. Besides our benchmark, this engine produces training data OmniReasoning-SFT-112K and OmniReasoning-RL-19K. Finally, we propose an on-policy self-distillation method Modality-Factored Self-Distillation (MFSD). It evaluates each sampled response under modality-specific clue contexts, disentangling the contributions of individual clues and their cross-modal interactions for token-level credit assignment. With our training data and learning method, our model OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points, respectively. Moreover, it delivers substantial gains on general and long-video benchmarks, including Video-MME-v2. We hope our work offers a solid step for facilitating future research in omni-modal joint reasoning.
cs.AI / 108 / 2609.39573
Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond
Abstract
As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state of the generation and often causes unintended global changes. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on the noisy states of flow models, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Despite using no explicit spatial masks or object priors, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free.
cs.AI / 109 / 2609.39601
GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Abstract
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
cs.AI / 110 / 2609.39704
When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic
Abstract
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5\% relative on some datasets and degrades it by up to 100\% on others. We trace this instability to spurious inversion: background patches receive higher CLIP text-similarity than the true object when the spurious attribute is background-separable, inverting the assumption every text- and attention-guided pruning method relies on. We introduce the Spurious Inversion Metric (SIM), a label-free, pre-deployment diagnostic whose sign predicts this effect with statistical significance (binomial $p=0.035$) across all 8 datasets, and remains dependable across 6 CLIP architectures with a clean foreground/background split. Naive masking is itself a major source of risk: it causes the largest average-accuracy loss of any method we evaluate, and its own per-image segmentation step is a significant runtime bottleneck. To address this, we design a batched, synchronization-free GPU segmentation routine that cuts this overhead from 3.5$\times$ to 1.75$\times$ baseline. Gating deployment by SIM's sign recovers masking's benefits while avoiding its worst failures, matching or exceeding a strong pruning baseline on 7 of 8 datasets.
cs.AI / 111 / 2609.39723
Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
Abstract
Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global classifier guidance. We demonstrate three key findings: 1. A carrier mitigates subject distortion by absorbing a larger share of globally normalized attack updates. 2. A carrier improves cross-model transferability, governed by the strength of target-related features that balance semantic separation and transfer performance. 3. Successful targeted attacks retain the personalized subject as the primary content perceived by humans while successfully misleading the classifier. Our results demonstrate that a visually secondary carrier offers an auxiliary spatial pathway for adversarial changes, enabling strong and transferable attacks while improving subject preservation.
cs.AI / 112 / 2609.39924
CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
Abstract
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git
cs.AI / 113 / 2609.40031
WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks
Abstract
Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. Recent advances in invisible watermarking methods highlight the need to update existing benchmarking practices to reflect current techniques and evaluation criteria. We address this by introducing WARP -- a unified framework and benchmark for evaluating the robustness of invisible watermarks. WARP incorporates 32 recent classical, deep, and generative watermarking methods, as well as 34 different erasing techniques, ranging from traditional distortions to more sophisticated adversarial, purification, and re-embedding attacks. It provides standardized, reproducible, and easily scalable protocols for evaluating perceptual quality, watermark readability, and attack resilience. Using WARP, we extensively evaluate current invisible watermarking techniques, collecting the largest robustness benchmark in the field. Results identify the most robust approaches under both distortion and adversarial conditions, and reveal consistent relationships between watermarking methods and the attack strategies most effective against them. Our experiments also highlight that some of the watermarking methods considered are highly vulnerable to reembedding, even if they are robust to standard distortions. The code is made available at https://github.com/ispras/wibe.
cs.AI / 114 / 2609.40055
Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding
Abstract
On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that supervision trustworthiness and supervision necessity are distinct yet coupled: the former concerns target credibility, while the latter varies with the student's current task competence; together, they shape supervision value. Building on this coupled view, we introduce Student-Curriculum Coupling (SCC), a closed-loop framework in which a compact Anchor-Frontier curriculum defines the candidate supervision space and the evolving student dynamically determines its active subset. Supervision can therefore be activated, suspended, or reactivated as competence changes, concentrating teacher computation and optimization on current task-level deficits. Across three TVG benchmarks, SCC achieves a 5.1% relative improvement in mean recall over Video-OPD on its original curriculum, while using 60.0% fewer training examples and reducing training time by 50.4%. Ablations support the complementary roles of capability-structured curriculum design and student-dependent supervision in achieving these gains. Together, these results establish SCC as a data- and compute-efficient framework for TVG post-training, delivering stronger temporal grounding by aligning trustworthy supervision with the student's evolving learning needs.
cs.AI / 115 / 2609.40091
GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation
Abstract
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical efficacy metrics, where a missed abnormality is most costly. We propose GateSPINE, a vision-language framework that fuses sagittal T1 and T2 volumes with a training-free operator, encodes the fused sagittal and axial volumes with two parallel 3D encoders, and decodes their combined representation into a report. Its core mechanism is a gated cross view fusion module that predicts, per feature channel and token, how much of each view to admit, so the more informative view dominates at each spatial location. We evaluate GateSPINE on three lumbar MRI datasets, comprising two public bench marks and a private cohort collected from Phenikaa University Hospital, using both natural language generation (NLG) and clinical efficacy (CE) metrics. GateSPINE achieves the highest CE F1 through improved re call on all three datasets; on SPIDER, which lacks an axial sequence, this reflects the sagittal fusion component rather than the gated cross-view mechanism, which is validated on the two cohorts with both imaging planes. GateSPINE also remains competitive on standard NLG metrics.
cs.AI / 116 / 2609.40195
MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
Abstract
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.
cs.AI / 117 / 2609.40219
Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
Abstract
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{https://github.com/JiahuaDong/AED}{AED}.
cs.AI / 118 / 2609.40230
EviRover: Reinforcing Agentic Perception Beyond a Glance
Abstract
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.
cs.AI / 119 / 2609.40253
ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
Abstract
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
cs.AI / 120 / 2609.40356
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
Abstract
Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video scene text editing that achieves high visual quality, temporal consistency, and edit locality remains underexplored. Existing resources offer limited paired real-video data, and general video-editing metrics do not directly measure whether the requested text remains correct over time. We introduce ViTeX-Bench, a benchmark suite comprising ViTeX-Dataset and a three-axis evaluation protocol. The dataset contains 387 real-world 720p videos with text-region masks and editing instructions: 230 provide reviewed, pipeline-generated paired edits for training, and 157 form a frozen evaluation split. The protocol evaluates text correctness, visual and temporal quality, and edit locality through 13 metrics, with one primary metric per axis and a Pareto comparison of their trade-offs. OCR calibration, human evaluation, and annotation-sensitivity analyses support the interpretation of these scores. Across eight baselines from four editing families, accurate text, temporal stability, and scene preservation remain difficult to achieve together. We also release ViTeX-Edit-14B, an open-source reference editor fine-tuned on the paired training split with motion-aligned glyph-video conditioning. It achieves CharAcc 0.688, the highest mean among the evaluated video-native editors, and the lowest comparable text-crop Warp among raw editor outputs. ViTeX-Bench provides a reproducible foundation for studying these trade-offs in video scene text editing.
cs.AI / 121 / 2609.39537
A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act
Abstract
The EU AI Act (Art. 27) requires deployers of high-risk AI systems to conduct Fundamental Rights Impact Assessments (FRIAs) before deployment, yet the evidence needed for credible assessments is fragmented across incompatible incident repositories, risk vocabularies, and legal texts. We present a reusable Semantic Web-based framework that consolidates this evidence for two high-risk public sector categories: employment and worker management (Annex III(4)) and access to essential public services (Annex III(5)(a)). A curated 150-record corpus is annotated along four axes using keyword, LLM, and hybrid methods and serialised as a SPARQL-queryable knowledge graph of 1,351 RDF triples. Five FRIA demonstration scenarios surface 103 records (68.7% coverage). Evaluation against a 69-record gold standard reveals that LLM-assisted classification of the employment domain achieves only $κ= 0.045$, a cautionary result for automated fairness-related evidence retrieval in this domain. All artefacts are released openly to support adoption by regulators, national authorities, and SMEs.
cs.AI / 122 / 2609.39058
A 3GPP-Compliant Benchmark Dataset for RIS-Aided Beyond 5G Networks
Abstract
Reconfigurable Intelligent Surfaces (RIS) are emerging as a key technology for programmable wireless environments in the beyond the fifth generation (B5G) networks. However, data-driven RIS research remains bottleneck by the lack of standardized, high-fidelity and open-source datasets. In this paper, we introduce a large-scale 3GPP TR 38.901-compliant dataset for RIS-aided millimeter wave (mmWave) networks, that considers severe path loss, blockage sensitivity, and spatial channel sparsity make the RIS assistance more impactful. The dataset spans various canonical 3GPP deployment scenarios across 20 controlled variants, capturing diverse user densities, fading conditions, and blockage regimes. Uniquely, every sample includes oracle RIS phase configurations obtained via a globally optimal brute-force codebook search, providing gold-standard supervision labels that are absent from any existing public dataset. Rich multi-task annotations comprising full channel state information (CSI), per-link channel decomposition, optimal phase matrices, and channel quality index (CQI) labels support a broad range of machine learning paradigms and downstream tasks, including phase optimization, channel estimation, and interference management. As the primary benchmark task, we introduce a novel CSI-to-CQI mapping that frames RIS-aided link-quality prediction as a scalable scalar classification problem, thereby avoiding the exponential output complexity of the direct phase vector prediction. We have evaluated this mapping against state-of-the-art architectures under in-distribution, out-of-distribution, and real-world hardware measurement conditions. Our dataset provides a reproducible, extensible, and community-ready foundation to accelerate data-driven research in RIS-aided B5G networks.
cs.AI / 123 / 2609.38753
Where the Evidence Lives: Auditing AI Companions' Self-Descriptions
Abstract
Companion agents describe themselves: they remember, they understand their users, the relationship has changed them. We argue that such accounts, and the experience ratings that seem to confirm them, are checkable by users only where the evidence is theirs: in the agent's behavior, or in themselves. Where the evidence lives in the machinery, fluent self-description and moderately positive ratings do not establish that the mechanisms behind them ran. We demonstrate an audit procedure that sets an agent's self-description against its users' judgements and its implementation records, reporting each claim as supported, contradicted, or unresolved, and apply it to Lita, a proactive companion we built and deployed for a month with nine colleagues. Participants endorsed stylistic claims, withheld endorsement from relational ones, and rated memory at or above midpoint, while two of three memory layers had never executed their accumulation step. Memory-bearing agents should report what their self-descriptions cannot establish.
cs.AI / 124 / 2609.38907
Characterizing Questioning Patterns and Student Engagement Through Contextual Analysis of Real-Time Classroom Interactions
Abstract
Real-time classroom polling is now routine, yet the data it produces is usually read narrowly, as a correctness score or a headcount. Such readings say little about what a poll is doing within a lecture or how it shapes engagement. This is particularly relevant for short-response formats such as True/False, where the same question format can be used to test recall, check comprehension, or direct students' attention to a deliberately misleading statement. This study asks whether a poll's answer and instructional function can be determined by reading it against its lecture transcript, what cognitive levels of Bloom's taxonomy and instructional-function clusters the corpus contains, and how student engagement relates to answering correctly. We analyse a naturalistic corpus of 47 live sessions over 39 days, comprising 604 poll questions and 340,668 responses from 2,807 learners, most items True/False, read against time-aligned lecture transcripts and attendance. Reading each poll in context proves essential: the answer to 89% of polls is locatable in the lecture, and a recurring attention-checking device is visible only through context. Questioning is overwhelmingly lower-order and falls into seven instructional functions, and a poll's response follows its function rather than its wording. Engagement is broad but concentrated, and the class majority answers correctly 88.5% of the time, though a small set of high-consensus yet incorrect answers cannot be detected by agreement alone. An independent survey of 579 students agrees on what the polls are and on their participation, but reveals a gap between perception and reality: students cannot judge their own correctness, and the polls they find hardest are not those they answer worst.
cs.AI / 125 / 2609.39333
NarrativeSteward: Coordinating Delegation, Guidance, and Verification in Agent-Assisted Interactive Narrative Authoring
Abstract
Autonomous AI agents can turn authors' goals into interactive narratives by independently organizing and carrying out generation and revision. As agents generate and revise extensive content, authors struggle to grasp its overall structure, local details, and relationships, complicating continued guidance. We present NarrativeSteward, an authoring environment that organizes outlines, worldbuilding, and narrative graphs as linked artifacts for agent implementation and author guidance. Agent dialogue and project-wide structural review help authors understand the evolving work and guide local and cross-layer revisions, while change records and execution verification help authors assess the resulting work. Technical tests validated the system's change records, recovery mechanisms, and execution diagnostics. In a 12-participant within-subject study, NarrativeSteward supported easier formulation of revision requests and inspection of changes, and greater perceived understanding of changes and story structure, than general-purpose agents. Qualitative findings show how reviewing the work and feedback helps authors develop requirements and guide subsequent delegation. We open-source NarrativeSteward at https://github.com/Tencent/NarrativeSteward.
cs.AI / 126 / 2609.39906
Understanding Parents' Complex Views of AI for Children's Pretend Play
Abstract
AI could support children's pretend play, but it could also direct the play on behalf of children. Whether AI should have roles in children's lives is controversial because its influence on children remains uncertain. We conducted semi-structured interviews with 10 U.S. parents, each with at least one child aged 4-15. During the interview, we described the concept of AI-supported pretend play and provided participants with two boundary-case storyboards. We analyzed the interview data through codebook thematic analysis, using inductive coding and affinity diagramming organized around the research questions, and then used qualitative systems mapping to examine relationships within and across themes. We found that the same characteristics of AI, e.g., ability to assume characters, responsiveness, and adaptability, were seen by parents as potentially useful but also concerning. Parents imagined that AI could make role-based play accessible to all children or help parents participate in family play. However, they opposed the idea of AI for children's play without a clear understanding of how it works and its long-term influence on their children. Parents worried about children's loss of imagination and creativity, emotional attachment to AI, reduced human interaction, inappropriate behavior by AI and/or children, and their inability to manage children's AI use. Parents viewed AI not only as a play tool but also as a social actor and a possible perturbation in the existing family dynamics. The appropriateness of AI and child--AI interactions therefore emerged as a requirement for AI in children's pretend play, in addition to technical safeguards and parental control. We contribute an integrated account of parents' interdependent judgments and emphasize the need for longitudinal research with children and their diverse families.
cs.AI / 127 / 2609.39976
Richard: Voice-First Mobile Interaction for Persistent Tasks
Abstract
Mobile terminals need to provide application and network services while supporting users' control over their attention. We explore voice-first interaction organized around requests and delegated tasks, allowing users to leave a conversation and later inspect, revise, and retrieve the work. We present Richard, a system prototype that manages voice sessions, task execution, and result delivery separately, linking them through persistent request records. Conversation and task views provide visual feedback, while the backend coordinates immediate responses, dedicated service operations, and agent tasks. Request revisions, execution states, and notifications remain associated with the relevant task. We examine this design through Android functional records, controlled lifecycle verification, and execution records of a real programming request. Controlled verification reproduces revision, execution after confirmation, and result retention; deployed-service records show backend progress and failure feedback after client disconnection. These observations inform the design of task continuity, user control, and service integration in mobile voice interaction, providing an implementation basis for personal computing devices that accommodate intermittent user participation.
cs.AI / 128 / 2609.38822
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
Abstract
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
cs.AI / 129 / 2609.38471
Derandomizing Dense Binary Hypervector Codebooks for Quantized Scalars
Abstract
Hyperdimensional computing and vector symbolic architectures often represent quantized scalar levels by dense binary codebooks whose level-to-level similarity is intended to follow a prescribed function of scalar separation. At finite dimensionality, randomized scalar codebook constructions deviate from this target because of sampling noise, random-start imbalance, update-count fluctuations, component dependence, and finite-capacity effects. We develop a transition-based derandomization framework for dense binary scalar codebooks across two target-similarity families, with similarity decaying exponentially or linearly with level separation. The framework separates the target similarity law, the derandomization variant, and the concrete generator construction, making explicit how initialization, selection, update, and capacity-handling mechanisms shape the induced similarity profile. We formalize derandomization variants that separately constrain initial Hamming weight, update-count variability, and update balance, thereby controlling distinct sources of finite-dimensional error. For each family and variant, we derive the induced mean similarity, identify realization-wise and mean target-matching regimes, and derive exact finite-dimensional expressions for bias, variance, and root-mean-square error. Simulations across dimensions, quantization ranges, reference scalar levels, and generator constructions validate the theory and show how each constraint removes or reduces a specific source of similarity mismatch. The results provide practical guidance for choosing scalar codebook generators that more closely match a desired similarity law under finite-dimensional and hardware-relevant constraints.
cs.AI / 130 / 2609.39081
Coding Agents for Coding Theory
Abstract
We spent five weeks using an LLM coding agent on open problems in coding theory: finding large sets of four-letter words, such as DNA barcodes, that stay far apart in edit distance. The agent wrote the verifiers and search code; a human chose the problem and set the verification protocol. Restricting the search to codes with a prescribed symmetry, a classical technique, shrank the problem about fourfold and raised the best known code of length 6 and minimum edit distance 3 from 114 to 120 words ($E_4(6,3) \geq 120$). The same pipeline improved twelve further lower bounds at lengths 6 to 9 and distances 3 to 6. We give the failures equal space. Our own search stopped at 116 and recorded the last symmetry class as topping out at 112; a second agent session, running the same search with a better operator, found the 120. A later verdict that the method did not carry over to length 7 was wrong for the same reason, and an earlier instance cost three weeks. Each time, an intermediate result was written down, never rechecked, and treated as a fact that ruled out further search. Checking final outputs, as our protocol required, does not catch such errors.
cs.AI / 131 / 2609.40290
CAS II: Symmetric Partitions as Kolmogorov Models
Abstract
In algorithmic statistics a string x is explained by a finite set containing it, and Kolmogorov's structure function records the smallest such model at each level of complexity. Vereshchagin's strong models, those computable from the data by a total algorithm, are essentially the cells of simple partitions. We read a partition of binary strings as a hypothesis, with the cell containing x as its model, and develop algorithmic statistics over symmetric partitions: the orbit partitions of groups acting on strings. The Galois connection between subgroups and partitions gives each ambient group a lattice of symmetric partitions, with canonical certificates, canonical costs, and an algebra of hypotheses. The resulting structure function and symmetric sophistication measure which part of the regularity of x is symmetric. For the full symmetric group every partition is symmetric: cells recover all Kolmogorov models, cells of cheap partitions recover exactly the strong models, and normal and strange strings are characterized by symmetry. For GL(n,2) the cells are exactly the linearly homogeneous sets, so linear symmetry is a restricted model class. For nonzero x, the linear-symmetry structure function lies in a band between the sufficiency line and the trivial bound, and both edges are attained: there are stochastic normal strings whose simple structure is invisible to linear symmetry. We also give coordinates on the space of permutation groups: each group is an element of a Burnside ring (its type) together with a permutation (its placement), and restriction moves refine partitions via the Mackey formula. In these coordinates the collapse for the symmetric group is a statement about placement, a linear hypothesis is determined by its type up to n^2 bits, and the maximal gap theorem shows that any space of symmetry hypotheses small enough to search is small enough to miss simple structure.
cs.AI / 132 / 2609.38400
GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots
Abstract
Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. Since the same speech can be accompanied by different gestures, a robot can respond to workspace constraints, e.g., gestures for speech next to a wall. In these scenarios, the robot should gesture in a suitable motion rather than simply correcting an unconstrained one. To achieve this goal, we present GestAdapt, a workspace-conditioned framework that conditions co-speech gesture generation on a prescribed wrist workspace. The GestAdapt framework learns from six complementary co-speech corpora through a shared motion representation and supports retargeting to different robot embodiments. Quantitative evaluation shows that generated motions remain close to the real-motion distribution while respecting the workspace. In a user study, gestures generated under modified workspace constraints receive a mean quality score of 3.24/5, above our no-workspace variant (2.43/5) and below the reference motions (3.68/5). In a real robot evaluation, all compared motions are retargeted to the Reachy2 humanoid robot under identical workspace constraints. Motions generated with our framework rank first in 69.7\% of comparisons, higher than our no-workspace variant baseline and retargeted ground-truth motions constrained afterward. Overall, the results support adapting gestures to the available workspace during generation, rather than modifying unconstrained trajectories afterward to satisfy workspace constraints, potentially compromising gesture naturalness.
cs.AI / 133 / 2609.38653
TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion
Abstract
Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human motions. Yet these capabilities remain largely confined to flat ground, in part because motion datasets rarely include aligned terrain geometry and because retargeting terrain interactions to complex musculoskeletal bodies is challenging. We present TERRA, an end-to-end pipeline for terrain-aware retargeting and control of musculoskeletal locomotion. From kinematic trajectories alone, TERRA combines terrain priors, estimated contacts, and negative free-space evidence to recover task-relevant support geometry. TERRA further considers anatomical, tendon-continuity, and contact constraints during retargeting. Using the resulting motion-terrain pairs from five datasets, we successfully train a single muscle-actuated control policy on 9.4 hours of diverse locomotion. Across reconstruction, retargeting, and held-out tracking benchmarks, TERRA improves terrain accuracy, sharply reduces anatomical and interaction violations, and achieves the highest observed completion rate over supported terrain families. Overall, TERRA provides a practical route from scene-less motion data to muscle-actuated locomotion over diverse non-flat terrain. Project website: https://cnai.epfl.ch/terra/
cs.AI / 134 / 2609.38855
Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution Optimization
Abstract
Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic manipulation. Unlike deterministic policies, Flow Matching enables VLA models to learn conditional action trajectory distributions, where latent noise vectors induce different actions under the same task scenario. However, we observe that these distributions are often ill-formed, with successful and failed behaviors coexisting while considerable probability mass remains in unfavorable regions. To this end, we propose Online-ES, an online adaptation framework for Flow Matching VLAs based on Evolution Strategy (ES), which refines the learned action trajectory distribution through interaction feedback. Instead of pruning the latent noise space, our method performs evolutionary exploration directly in the action trajectory space, where diverse trajectories generated by Flow Matching provide candidate solutions for adaptation. By perturbing sampled trajectories and evaluating their execution outcomes, we derive a self-supervised MSE objective that transfers the evolution direction from trajectory space into model parameter space. Mathematically, we prove that the proposed objective provides an unbiased estimator of the optimal evolution direction. Moreover, we also incorporate failure experiences as negative feedback to regularize the evolution direction, steering the policy away from previously explored failure regions. Experiments in both simulation and real-world environments demonstrate that Online-ES achieves policy improvement comparable to reinforcement fine-tuning, without learning a value model or computing advantages.
cs.AI / 135 / 2609.38862
Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving
Abstract
Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computational overhead and suffer from objective misalignment. Additionally, preference-based methods rely on strict pairwise annotations, limiting data utilization. To overcome these limitations, we propose EMPlan, an efficient multi-modal trajectory planning method powered by reward-guided fine-tuning. We design a hybrid architecture that combines sparse anchors with an offset refinement module for efficient multi-modal trajectory prediction. Sparse anchors provide coarse trajectory candidates with low latency, which are subsequently refined by the offset module for higher prediction accuracy. To enhance safety without incurring additional inference costs, we adopt a two-stage training paradigm consisting of pretraining and reward-guided fine-tuning. During fine-tuning, we leverage rule-based reward signals and unpaired preference supervision to refine the pretrained policy toward safer trajectory selection. We evaluate EMPlan on the non-reactive NAVSIM benchmark, where it strikes a favorable balance between planning accuracy and efficiency, demonstrating superior performance under real-time constraints.
cs.AI / 136 / 2609.38948
DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
Abstract
Frontier models excel at many digital benchmarks, yet their ability to drive a real car, an everyday human skill, remains largely untested. We present DrivingBench, to our knowledge the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames from a Toyota Corolla and directly command its steering and velocity around a parking lot cone course at low speeds. The car may continue moving while the model thinks and new commands replace the currently running one, so inference latency is part of the task, testing the models' abilities to observe, act, monitor, recover, and complete a long-horizon objective under such constraints. We benchmark GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 in vendor-native harnesses (Codex, Claude Code, Cursor) with up to three attempts each in one conversation; Astra is the only model to finish the course, on its second attempt, with no other attempt passing 50% of the course. Two of the four models improved materially across attempts with retained context. We also detail the design principles behind our action interface, and show how the tool output format and the framing of the task combined to determine whether models would drive at all or refuse. We release our harness, prompts, course map, and traces with video and telemetry for reproducibility.
cs.AI / 137 / 2609.39018
Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools
Abstract
Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Universal Robot-Agent Interface), which couples a programming agent that constructs robot tools with an execution agent that uses them in a feedback loop. The programming agent writes reusable and task-specific tools from task intent and refines them through execution feedback and human guidance. The execution agent selects and parameterizes these tools from current observations; each call runs a complete motion locally before returning control to the agent. Unlike delegating subsequent decisions to a generated program, this design retains model-level decision-making between tool executions. Validated tool revisions persist across episodes without updating foundation-model weights, and a shared GUI and API make the same tools available to humans and agents. Across five RoboDojo tasks and four frozen execution agents, URAI raises aggregate success from 18.0% to 53.0% relative to direct fingertip control, with the largest gain on Swap Blocks; with the same tools, a program written in advance reaches only 24% against 56% for two agents deciding after each call. Three of the four agents also finish episodes 1.3-1.5 times faster with 1.5-1.7 times fewer execution-agent output tokens; DeepSeek-V4-Flash's cost barely changes. We further evaluate URAI on seven real-world AgileX dual-arm tasks, spanning object manipulation, cloth folding, and human-interactive tic-tac-toe. URAI connects the coding and decision-making capabilities of frontier agents, organizing robot control around reusable tools that agents can both invoke and revise.
cs.AI / 138 / 2609.39145
Blackout vs. Freeze: Analyzing Physical Failure Modes of VLAs under Camera Faults
Abstract
Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how $π0.5$ and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even when task-success rates are similarly low: freezing causes more extreme joint behavior, whereas blackout after gripper closure can cause more object drops, most markedly without proprioception. Selective intervention studies reveal that proprioception (current robot state) partly compensates for the removed robot depictions and reduces non-target contact. However, it cannot sufficiently restore task success when wrist-view object information is removed, even when aided by the remaining scene view. We then evaluate two mitigation approaches: camera-blackout training and training-free replacement of faulty visual embeddings. Both improve task success in selected conditions, but can increase unintended contact or disturbance to surrounding objects. Real-robot trials further show that successful execution under camera faults can still involve unintended physical interactions. These findings motivate designing VLA policies that use the robot and object information still available under camera faults to limit hazardous motion.
cs.AI / 139 / 2609.39304
Scale and Selection: What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents
Abstract
When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness, its prompts, tools, and control rules, largely determines success, and until now it has been written by hand. We show that this harness can be improved automatically by another coding agent, the optimizer agent, and report two findings about what makes it work. First, the number of rollouts the optimizer agent sees per round governs whether the evolved harness is trustworthy, generalizes, and improves steadily. A single rollout is a noisy binary outcome, so with few rollouts per round a revision can be promoted on luck; enlarging the batch raises the signal-to-noise ratio of every promotion decision. Holding rounds fixed and growing the training set from 5 to 100 rollouts, held-out success rises from 47% to 67%, while small training sets overfit, reaching 70% on training tasks but only 54% held-out. Second, the optimizer agent must not be given free rein. With every revision it proposes accepted unconditionally, performance drifts downward within ten rounds as ill-judged edits accumulate; adding the most basic safeguard, Champion-Challenger selection that promotes a revision only if it strictly beats the incumbent on the same fixed evaluation set, turns the same loop into one that raises held-out success from 51% to 67% over 30 rounds. Automatic harness evolution for visual-interface robot agents is thus feasible, but its gains hinge on the rollout scale behind each decision and on how the optimizer agent's revisions are selected.
cs.AI / 140 / 2609.39323
HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning
Abstract
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.
cs.AI / 141 / 2609.39685
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Abstract
Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/
cs.AI / 142 / 2609.39763
DiffWAM: A Fast and Efficient Navigation World Action Model
Abstract
Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a geometry-conditioned navigation world-action model that directly transforms multi-level predictive features into continuous camera trajectories. Its Grid-Motion module preserves spatial-temporal motion associations, while Latent2Pose grounds them with first-frame geometry to recover metrically meaningful 3D motion. Complete video rollouts and geometric reconstruction are required only for offline supervision, eliminating future-video decoding and multi-frame reconstruction during deployment. We further introduce FastDreamer, which overlaps predictive and geometric computation with ongoing flight and performs timestamp-aware asynchronous trajectory handoff for continuous UAV execution. DiffWAM achieves a trajectory RMSE of 0.3492 m and an endpoint success rate of 74.40% on the 1,000-sample DiffWAM-1000 benchmark, while representative real-world experiments demonstrate complex behaviors including constrained traversal, orbiting, S-shaped flight, and multi-stage navigation. An onboard DiffWAM-Flash implementation further reaches 1.08 s model-pipeline latency on NVIDIA Jetson AGX Thor. These results demonstrate that predictive video representations can be efficiently grounded into continuous 3D motion, providing a direct alternative to generate-then-reconstruct navigation pipelines. Project page: https://zzmmzzm.github.io/diffwam.github.io/.
cs.AI / 143 / 2609.39820
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving framework that converts runtime feedback into persistent policy improvement. During collection, a fixed CBF-based safety module serves as an observe-only teacher, producing counterfactual corrections while the policy remains in control. Outcome-aware admission then converts useful proposals into corrective targets and retains successful uncorrected actions as quiet anchors for guarded LoRA updates. We evaluate FailBank on the VLA-Arena benchmark across two difficulty levels and two VLA backbones. Compared with the base policies, FailBank improves the joint success-cost operating point. Across the two backbones, FailBank improves task success rate by 8.5 and 6.9 percentage points, while reducing policy-induced cumulative cost by 35.6\% and 23.8\%, respectively. Compared with runtime shielding, FailBank raises task success rate by 25.4 and 9.5 percentage points, while maintaining comparable policy-induced cumulative cost. These results show that runtime feedback can serve as persistent policy supervision rather than only as a temporary action constraint.
cs.AI / 144 / 2609.40085
BatSLAM 2.0: Sequence-Verified Sonar Place Recognition in a Robust Pose Graph
Abstract
Echolocating bats can navigate dark and cluttered spaces using echolocation. Over a decade ago, BatSLAM showed that a robot with a biomimetic binaural sonar can build a topological map of the environment, by recognizing places from the received acoustic signals. Sonar place recognition, however, is ambiguous by nature: corridors produce nearly identical echo trains, and wrong loop closure can collapse the topological map. In this paper, we introduce BatSLAM 2.0, a novel sonar-only SLAM system built from three elements: an updated acoustic front-end, a sequence verifier that tracks and verifies loop closure candidates and a pose graph implemented on a high performance factor graph framework. The system was thoroughly evaluated both in simulated as well as real world recordings. In both cases, the BatSLAM2.0 algorithm shows the capability of robust topological map creation, countering map collapse, and robust scaling of map size.
cs.AI / 145 / 2609.40134
Tactile Curiosity Drives Robot Interaction
Abstract
Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty in functionally irrelevant transitions, such as erratic motions in free space. In this work, we argue that tactile feedback provides a natural signal for exploration, and introduce TacEx, a framework that incorporates touch into epistemic uncertainty-driven exploration by decomposing model uncertainty across sensory modalities and directing curiosity toward the tactile channel. By anchoring curiosity to the sense of touch, TacEx drives the robot to discover complex contact dynamics, learning to manipulate and grasp objects without task rewards or expert demonstrations during exploration. The interaction-dense dataset collected through this tactile-driven curiosity supports offline learning of downstream pick-and-place policies without additional environment interaction. We further use tactile-driven exploration to post-train vision-language-action (VLA) models. Although the VLAs are initially pre-trained without tactile feedback, post-training with TacEx substantially improves downstream performance while remaining highly sample-efficient.
cs.AI / 146 / 2609.40306
DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents
Abstract
Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised. We propose DynaHarness, a dynamic physical harness that couples semantic reasoning with physical governance through a shared execution contract and turns failure evidence into validated capability revisions. To be more specific, the slow brain proposes capabilities and symbolic arguments, while the fast brain grounds and monitors commands, refuses unresolved actions, substitutes capabilities, and requests replans when needed. The physical execution contract bounds each accepted command and records execution evidence across analytic skills, recovery skills, and the frozen VLA. Failure attribution localizes faults in these records and directs targeted revisions of reusable capabilities or execution mechanisms. Paired regression checks govern admission or rejection, closing the self-evolution loop. On LIBERO-Pro, DynaHarness achieves 75.2% on 800 newly sampled initial states, compared with 17.5% for the frozen policy. With the same capability library, full dynamic execution reaches 74.0% versus 63.9% under nominal one-step replanning. This demonstrates the value of DynaHarness as a dynamic physical harness that governs how existing capabilities are grounded, monitored, and coordinated during execution. Our project page is at https://denghaoyuan123.github.io/Dynaharness_page/.
cs.AI / 147 / 2609.38780
RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition
Abstract
Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.
cs.AI / 148 / 2609.38878
Audio Token Attention Is Predictable Before the Language Model Runs
Abstract
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $ρ\geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io
cs.AI / 149 / 2609.38897
FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs
Abstract
Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.
cs.AI / 150 / 2609.39088
SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS
Abstract
Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline Prosody Control, where each Pitch, Energy, or Speed instruction targets a clause relative to the preceding clause from the same speaker, while Pause uses an absolute duration interval. In codec-based TTS, Speed and Pause affect sequence length, whereas Pitch and Energy rely on residual codebooks. By analyzing Qwen3-TTS RVQ codebooks, we find that Energy concentrates in early residual codebooks, whereas Pitch accumulates across a deeper prefix. We therefore propose Scope- and Codebook-Aware Instruction Conditioning (SCIC), combining a Temporal Instruction Router for frame-level tag activation with Tag-Specific Codebook Weighting over residual codebooks. SCIC improves speaker-relative Pitch and Energy control over standard instruction fine-tuning using text-token tags. We further apply multi-reward GDPO post-training to jointly optimize control and quality, improving control accuracy while preserving CER and speaker similarity. In long-form synthesis, SCIC produces a more distinct paragraph-level expressive hierarchy than speaker-adapted SFT without instructions. Audio demos are available at: https://taoliveaigc.github.io/SCIC/
cs.AI / 151 / 2609.39199
UniAE-MoE: A Unified Audio Encoder via Mixture of Experts
Abstract
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance \tool's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of \tool for unified audio understanding across speech, music, and general audio domains.
cs.AI / 152 / 2609.39344
Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
Abstract
Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across meetings. We evaluate persistent speaker attribution with Speaker Identified cpWER (SI-cpWER), which scores a corpus under one global speaker-ID assignment. The benchmark covers five commercial diarize-then-identify cascades, two open academic baselines, and ThyVoice on the full 129-meeting CHiME-8 NOTSOFAR evaluation set in clean and noiseaugmented form, plus CHiME-6. ThyVoice is our end-to-end reference system; it repairs overlap and gates the evidence used to create and update voiceprints. Requiring persistent identity changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions and the lowest mean in the full panel, 47.13 versus 54.75 for the next system. Complementary lexical, diarization, per-recording attribution, and speaker-clustering diagnostics characterize upstream error surfaces in the final attributed record. These results show why persistent attribution must be evaluated directly in systems that reuse conversations across time.
cs.AI / 153 / 2609.40067
Community-Driven API and AI Writer Design for Openly Scaling Community Notes
Abstract
Community Notes is a crowd-sourced approach for adding context to posts on X. Contributors propose and rate notes, forming the inputs to an open-source, open-data algorithm that determines which notes show broadly to users. Since September 2025, Community Notes' AI Note Writer API has provided an open, public interface for using AI to propose notes, while adhering to the founding principle that users, not the platform or an AI, control which notes show on X. Explicit note requests and user posts on X determine the AI API post feeds, ensuring that AI note writing responds to demand from X users. We present the design, operation and impact of the AI API, including analysis of the interaction between AI and human generated notes across topics. Unless otherwise stated, measurements and system description reflect June 2-29, 2026. The Community Writer is the largest AI API client and contributes the bulk of AI API output, generating 52% of notes selected as Helpful and shown broadly on X. The writer is guided by community input during both training and operation to prioritize, draft, evaluate and delete proposed notes. Beyond scale, the writer also offers speed, submitting the first proposed, non-deleted note on 60% of posts when compared to other writers. AI note writing is additive on top of human note writers, extending coverage of Community Notes on X. Among posts that have Helpful notes, 42% have only AI notes, indicating human raters did not feel motivated to propose an alternative. In contrast, 30% have only human notes, reflecting contribution beyond the scope of AI writing. The Community Writer is open-source software released under the Apache 2.0 license.
cs.AI / 154 / 2609.40071
Grounding Time-Series Foundation Models in Digital Twin Topology for Predictive Maintenance
Abstract
Digital twins increasingly support downstream analytical tasks that depend on time-series data, motivating interest in time-series foundation models (TSFMs) as scalable backbones. However, TSFMs are primarily pretrained for temporal continuation and often underperform on unseen tasks such as regression, and systematic empirical comparisons against state-of-the-art dedicated models in digital twin contexts remain limited. This paper makes three contributions. First, we benchmark five well-known TSFMs with frozen backbones on remaining useful life (RUL) prediction using the C-MAPSS dataset, finding that multivariate architectures substantially outperform univariate ones, particularly under varying operating conditions. This raises a deeper question: when cross-channel dependencies can be modeled through pretrained weights, target-task adaptation, and digital twin-derived representations, how much does each contribute, and are they complementary? Second, we propose a topology-informed fusion approach in which topological constraints, derived from the asset structure the digital twin stores among its information models, explicitly shape cross-attention, so that fused representations respect the physical system's local connectivity rather than relying on unconstrained all-to-all interactions. Third, we conduct an ablation study across C-MAPSS subsets of varying operational complexity that isolates the three sources and their interactions. The sources prove complementary rather than redundant, and topology-constrained attention outperforms unconstrained fusion, though by a small margin, enabling a frozen TSFM informed by digital twin representations to remain competitive or in some cases exceed state-of-the-art performance on this regression task.
cs.AI / 155 / 2609.38309
Searching for BSM Experimental Signatures with Large Lagrangian Models
Abstract
The search for physics Beyond the Standard Model (BSM) is generally limited not by the supply of theory descriptions but by the lack of discriminating experimental observations. A case in point is dark matter, where the overwhelming gravitational evidence only goes so far in distinguishing between models within a vast theory space. Exploring the space of testable model signatures may help identify overlooked experimental observables and indicate the utility of future experiments. A challenge is designing a search through model signatures outside what is found in the literature. Our primary contribution is hAIthem, a framework that combines the self-guided exploration of reinforcement learning (RL) with the broad literature-derived knowledge of LLMs. We build an RL agent that learns to find which portions of a theory's high-dimensional parameter space are not excluded under some subset of constraints by playing a Battleship-style "game" against a suite of phenomenology tools. The agent is built as a Large Lagrangian Model (LLaM), an autoregressive transformer that reads a tokenized Lagrangian, is pretrained at scale (here on ~1 billion tokens from ~10,000 Lagrangians), and is fine-tuned in a live environment. The framework then constructs a decision tree that separates RL-found regions using observables computed with established tools, and passes the remaining degenerate regions to a set of LLM agents that compete to produce realistic signatures. In this proof of concept, RL-search outperforms an evolutionary-algorithm baseline, finding more viable regions with greater physical diversity. In a restricted space of single dark scalar multiplet models, we find that hAIthem proposes interesting combinations of previously studied observables, such as the application of a halo-independent kinematic ratio to paleo-detectors.
cs.AI / 156 / 2609.39914
Cluster Attention Neural Operators for Solving Parametric Partial Differential Equations
Abstract
Traditional simulations of parametric partial differential equations (PDEs) rely on repetitive computations for each parameter, which makes high-fidelity design impractical. Neural operators address this issue by learning solution operators, accelerating parameter-space mapping by orders of magnitude. Recent Transformer-based neural operators attempt to capture global dependencies, but often at the cost of quadratic attention complexity. Transolver resolves this problem by projecting physical states into a reduced slice space for attention computation. Although fast, this projection sacrifices fine spatial information. Moreover, by operating in this reduced space with shared weights across attention heads, it may constrain the model's flexibility, thereby limiting its capacity to capture complex phenomena. To address these issues, we propose the Cluster Attention Neural Operator (CANO), which reformulates attention via a novel cross-attention mechanism that dynamically clusters queries while preserving full-resolution keys and values. This avoids slice compression loss and removes weight-sharing limits. At the same time, the model remains fast without losing global interactions. Empirically, CANO achieves state-of-the-art performance across canonical PDE benchmarks, covering fluid and solid dynamics (e.g., Navier-Stokes, Airfoil, Plasticity), irregular unstructured geometries (e.g., Pipe Turbulence, Composites), and long-term temporal rollouts. Across solid deformation and turbulent flow benchmarks, CANO achieves lower errors than baselines and exhibits strong geometric adaptability and temporal consistency.
cs.AI / 157 / 2609.38434
TACIT: Optimization Models that Learn from Their Mistakes
Abstract
Real-world optimization problems are difficult to model accurately because many objectives and constraints reside in domain experts' tacit knowledge, making them hard to formalize. As a result, optimization models often contain miscalibrated objectives, missing constraints, or omitted decision variables, leading to solutions that fail to reflect operational realities. We address this challenge by automatically repairing misspecified formulations using historical data consisting of past solutions and subsequent user overrides. Traditional approaches such as inverse optimization and constraint learning tend to overfit sparse data and produce complex formulations. Our central idea is to combine the reasoning capabilities and prior knowledge of LLMs with the formal grounding provided by optimization. We realize this idea through two complementary paradigms. Top-down, an LLM proposes structural repairs, including new constraints and variables, whose numerical parameters are calibrated and validated through optimization. Bottom-up, optimization infers cuts from observed decisions, which the LLM contextualizes into interpretable, generalizable modeling constraints. We evaluate our approach on 38 misspecification scenarios spanning nine classes of optimization problems, several drawn from real-world applications, and show that TACIT can repair 78.9% of them (vs. 60.5% for the best baseline).
cs.AI / 158 / 2609.39010
An Uncertainty-Guided Digital Twin Framework for Online Adaptive Proton Therapy in Head and Neck Cancer: A Feasibility Study
Abstract
Objective: Head and neck (HN) proton therapy spans six to seven weeks of anatomical change, while offline replanning takes about a week. We present an uncertainty-guided digital twin (UGDT) framework that forecasts treatment-day anatomy before treatment and evaluate whether it generates online adaptive proton therapy (APT) plans of clinical quality. Approach: A library of 302 longitudinal deformations from 88 previously treated HN patients was transported onto each new patient's treatment planning CT (TPCT) using two-step multi-atlas deformable image registration (DIR) built on a pretrained CT foundation model, generating about 284 predicted CTs (pdCTs) with contours per patient. Dispersion of propagated clinical target volume (CTV) contours defined a patient-specific robust margin. In ten patients, the quality assurance CT (QACT) triggering a replan represented treatment-day anatomy, and the physician-approved replan was the baseline. The pdCT most similar to the QACT (pdCT-H) and one from the lowest quartile (pdCT-L) were planned to within about 5% of baseline plan quality, forward-calculated on the QACT, and reoptimized to generate online APT plans. Main results: pdCT plans scored within -0.7% (pdCT-H) and -1.0% (pdCT-L) of baseline. Forward calculation on QACT reduced high-dose CTV D98% to 88.3% and 85.5%. After online reoptimization, D98% recovered to 98.3 +/- 0.3% and 98.2 +/- 0.3%, versus 98.5 +/- 0.4% at baseline. Spinal cord and brainstem doses remained below tolerance, and plan quality scores were within -1.1% (p = 0.19) and -1.7% (p = 0.01) of baseline. Significance: UGDT generated online APT plans comparable in quality to physician-approved offline replans using anatomy forecast before treatment, enabling a transition from reactive offline replanning toward anticipatory online adaptation.
cs.AI / 159 / 2609.38908
CellMSA: Context Modeling for Single-Cell Representation Learning
Abstract
Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells from the same batch for denoising, thereby underutilizing the rich relational information across batches and cell types to model gene expression patterns. We argue that single-cell models can benefit from more informative cell-context modeling. By comparing consistency and variation across cells, models can capture fine-grained gene-gene dependencies associated with cell states, which are essential for learning high-quality representations. Inspired by the use of multiple sequence alignment (MSA) context in protein modeling, we propose CellMSA, a single-cell representation learning framework that introduces an MSA-inspired inductive bias into transcriptomic modeling. For each target cell, CellMSA retrieves relevant cells from different batches and biologically related cell types as context, and summarizes cross-cell patterns into a context-dependent gene-pair representation. This representation is then injected into a pair-aware target-cell encoder for fine-grained representation learning. We pretrain CellMSA on a large-scale human single-cell corpus of approximately 109 million cell observations, including 65.6 million primary observations. Experiments show that our framework consistently outperforms existing methods across multiple benchmarks. Code is available at the following repository: https://github.com/PharMolix/CellMSA.
cs.AI / 160 / 2609.39080
Association profile conditioning in a set-temporal transformer for cross-session intracortical motor decoding
Abstract
Intracortical motor decoders degrade across sessions because the set of recorded units changes and persisting units can alter how their firing relates to behavior. Most existing methods update network weights on each new session or rely on unlabeled activity, which does not directly reveal such changes. We present APST, an Association Profile-conditioned Set-Temporal transformer that adapts to new sessions with all network weights frozen. From a few labeled calibration trials, APST summarizes how each unit's firing relates to behavior in a four-dimensional association profile computed in closed form. The profiles condition a set-attention encoder that accepts any number and order of units, followed by a causal transformer for streaming decoding. On held-out DANDI688 sessions from two monkeys, APST reaches velocity $R^2$ of $0.78$ and $0.81$, versus $0.40$ and $0.58$ for a variant that uses neural activity alone, and matches or exceeds an RNN fine-tuned on the same trials. On FALCON private held-out evaluation, it attains $R^2$ of $0.65$, $0.42$, and $0.44$ on M1, M2, and H1.
cs.AI / 161 / 2609.39342
Belief-Based Maximum Occupancy Principle and Active Inference
Abstract
Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active Inference addresses the problem of acting in a partially observable environment through a principled framework for belief updating and action selection. A key component of Active Inference is the specification of prior preferences, which shapes behavior by encoding desirable future outcomes. An intrinsic motivation approach called the Maximum Occupancy Principle (MOP) proposes that agents act so as to maximize occupancy over future paths of states and actions, with no preferences or epistemic targets. Despite its simple formulation, MOP gives rise to rich and adaptive behaviors that combine exploratory variability with goal-directed dynamics. In this work, we extend MOP to partially observable environments and introduce a Bellman reformulation of the Expected Free Energy for Active Inference, both incorporating belief-based inference over hidden states as part of the agent state. The Bellman formulation enables tractable offline computation via value iteration over the full belief-state space. We compare the resulting behaviors in a set of minimal experimental settings with uncertain food sources. We find that MOP agents switch between goal-directed (food seeking) behavior and exploration between different food sources, depending on their energy available and their belief state. In contrast, Active Inference agents mostly inhabit regions around a single food source, a strategy having both high pragmatic and epistemic value. We finally compare with Empowerment, which is shown to be qualitatively similar to Active Inference.
cs.AI / 162 / 2609.38695
Always-On Experimentation
Abstract
Generative AI has dramatically accelerated the rate at which new treatments---from novel pharmaceuticals to online marketing campaigns---can be conceived and deployed. As a result, modern experimentation platforms often run continuously, with treatments added as they are ready and removed when they underperform. We formalize this "Always-On" experimental setting, in which treatments can be dynamically generated, added to, and removed from a running experiment, and study the statistical problem of deciding whether to accept or reject each treatment while controlling for the false discovery rate. We develop sequential tests that achieve time-uniform Type-I error control under arbitrary stopping times and "predictable" treatment schedules. Our approach builds on the testing-by-betting framework: we construct test supermartingales for testing the average treatment effect of each treatment, and show that the construction of these test supermartingales is growth-rate optimal in an almost-sure sense.
cs.AI / 163 / 2609.38659
Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit
Abstract
We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 2020), establishing a $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ minimax regret, where $T$ is the total number of interactions and $\tilde O(\cdot)$ drops all constant and logarithmic factors, improving the previous $\tilde{O}(\sqrt{KT/A})$ regret. We then provide a matching lower bound up to logarithmic factors, indicating that our established rate is nearly minimax-optimal. We further show that the knowledge of $A$ up to $\tilde{O}(1)$ factors is necessary to achieve near-optimal regret, as near-optimal algorithms for one number of optimal arms must incur substantially larger regret than optimal regret for a smaller number. Overall, our results provide a comprehensive minimax characterization of $K$-armed bandits with $A$ over the entire range of $1 \leq A \leq K-1$.
cs.AI / 164 / 2609.39484
CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series
Abstract
Longitudinal clinical cohorts are multimodal, irregularly sampled and pervasively incomplete: in ADNI, positron emission tomography and cerebrospinal fluid assays are absent from roughly half of all visits. Linear state-space models handle irregular sampling gracefully but treat a missing modality by masking the input, leaving the transition operator untouched. We prove that this is a representational limitation: the latent state of any linear state-space layer whose transition operator does not depend on the availability pattern is an additive function of the availability indicators, so no such layer can represent an interaction between two modalities being jointly present or jointly absent. We propose CAMOS, which gives each modality a bank of second-order oscillators coupled through a matrix that sits inside the differential equation and is gated by availability, so the transition operator itself becomes a function of which measurements were taken. Coupling invalidates the analysis of uncoupled oscillatory models, and we restore it: a per-channel Gershgorin budget makes the effective stiffness positive definite uniformly over all $2^M$ availability patterns and all gaps, an energy argument charges amplification to availability transitions rather than sequence length, and a channel factorization preserves exact associative parallel scans. On ADNI, CAMOS outperforms uncoupled oscillatory state-space models and clinical fusion models on same-visit staging, landmark prediction and longitudinal forecasting, and under zero-shot transfer to OASIS-3 it is the only model that avoids collapse to the majority class.
cs.AI / 165 / 2609.39843
BayesNDE: Bayesian Generative Modeling for Neural Density Estimation
Abstract
Density estimation is a fundamental problem in statistics and machine learning. In this work, we introduce BayesNDE, a neural density estimator based on Bayesian generative modeling. BayesNDE learns a Bayesian generative model and evaluates its density without requiring invertible networks or Jacobian-determinant computation. For each observation, it infers a sample-specific latent posterior to construct an adaptive proposal that focuses computation on regions contributing most to its density. Bridge sampling then combines samples from this proposal with separate posterior samples to estimate the density. Experiments on nonlinear and multimodal synthetic datasets show improved estimation of density values and better recovery of the density structure compared to the state-of-the-art neural density estimators. Applications to real-world datasets further demonstrate improved anomaly detection. Together, these results highlight BayesNDE as a flexible and effective neural density estimator, demonstrating how posterior inference can turn generative models into tools for density estimation. The code and tutorials are available at https://github.com/liuq-lab/BayesNDE.
机器学习 (cs.LG)
261
cs.LG / 1 / 2609.38312
Gestalt: a meta-foundation model for astronomy
Abstract
The Platonic Representation Hypothesis predicts that sufficiently scaled foundation models converge on a shared representation of the world. As each non-converged model gives a noisy view of a common structure when passed the same input, we ask whether we can combine models into a representation that outperforms its individual components. We test this on galaxies: we embed images via a basket of 22 frozen foundation models from eight families, whiten each view, and take a randomised SVD of the embedding concatenation. The resulting 1024-dimensional embedding outperforms every basket member on 19/21 of our tested metrics for physical property and galaxy morphology estimation for HSC, JWST, and DESI Legacy Survey imagery. We find that performance rises with basket size and basket architectural diversity, and that the meta-foundation model's performance transfers across astronomical surveys. We conclude that a useful astronomical foundation model can be assembled from existing generalist models with no training required beyond a single unsupervised projection. By leveraging the community's already-spent work, we save a lot of compute: a fresh pre-train of a comparable single-domain model would cost $\mathcal{O}(10^{4}$--$10^{5})$ A100 GPU hours (emitting several tonnes of CO$_2$eq.), whereas assembling Gestalt requires minutes on a single machine.
cs.LG / 2 / 2609.38638
SHIFT-Truck: A High-Fidelity Aerodynamics Dataset and Benchmark for Pickup Trucks
Abstract
Pickup trucks account for 14% of new light-duty vehicles produced in the United States, yet are among the least aerodynamic. Their open cargo bed adds a flow absent from existing automotive aerodynamics datasets such as DrivAerML and SHIFT-SUV: the shear layer leaving the cab roof passes over a recirculating bed flow before separating again at the tailgate. The resulting drag lowers fuel efficiency, raises emissions and limits the range of electric trucks. Scale-resolved Computational Fluid Dynamics (CFD) is too costly for broad design exploration; neural surrogates can predict flow features at a fraction of that cost, provided they are trained on large-scale, high-fidelity, domain-specific data. We introduce SHIFT-Truck, the first such dataset for pickup trucks. It comprises 1,000 Spalart-Allmaras delayed detached-eddy simulations (SA-DDES) of a reference pickup geometry morphed across 17 shape parameters. Each case is run on a mesh of about 100 million cells at a Reynolds number of $1.4 \times 10^7$ and released with time-averaged surface pressure, wall shear stress, volumetric pressure and velocity. The setup is verified by grid refinement and repeated runs, and checked against wind-tunnel measurements. We define geometry-grouped splits and benchmark four neural surrogates, DoMINO, GeoTransolver, AB-UPT and SMART, on surface and volume tracks. SHIFT-Truck also introduces controlled distribution shifts in the operating point, the input surface discretization and the vehicle archetype. Models with strong in-distribution performance can degrade substantially under these shifts: operating-condition changes expose failures to infer speed dependence, while tessellation and cross-vehicle shifts reveal markedly different robustness across architectures. SHIFT-Truck is thus a benchmark not only for surrogate accuracy but also for generalization across physical and numerical distributions.
cs.LG / 3 / 2609.38485
Beyond Layers: Position-Resolved Gradient Conflict and Position-Aware Modulation for Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts, measuring conflict per layer and resolving it by separating parameters. We argue that this resolution hides an orthogonal axis. Generation in a UMM is next-token prediction over a raster sequence of visual tokens whose roles vary systematically with position, so how strongly a generation gradient interferes with understanding should depend on where in the sequence it originates. We introduce a position-resolved interference map that attributes understanding-generation gradient conflict to visual-token positions within every layer, computed from a single backward pass at $1.2\times$ the cost of a standard backward pass. On Show-o and Janus-Pro, position explains a large share of conflict variance after controlling for depth (partial $η^2=0.31$ vs. $0.35$ for layer on Show-o; $0.15$ vs. $0.30$ on Janus-Pro): the first quarter of the sequence has a mean gradient cosine of $-0.18$ against understanding, the last quarter $-0.02$. The dependence survives per-position gradient-norm normalization, retaining $80%$ of its effect size, and conflict strength tracks semantic content (Spearman $ρ=0.64$). Building on the map, we propose position-aware modulation (PAM), which removes the anti-aligned component of generation gradients only at high-conflict positions without changing the architecture. Under a matched trainable-parameter budget, PAM improves over layer-wise separation by $+21$ MME and $+2.4$ GenEval points on Show-o while matching it on POPE and overall FID; a random-position control recovers about $31%$ of the gain. Position-based and layer-based separation are complementary degrees of freedom and can be combined.
cs.LG / 4 / 2609.38519
GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction
Abstract
Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.
cs.LG / 5 / 2609.38811
DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation
Abstract
Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone carries a separate Conv-LoRA expert bank and mask decoder per defect class, each trained in its own pass, without prompts, on synthetic slices alone, updating only 4.4% of the parameters. On benchmarks that XCT-SAM reports, DCM-SAM improves on every baseline for both classes from a ViT-B backbone against their ViT-H, and reaches 64.2% pore IoU on real NIST scans having seen no real images during training. Deployment then exposes what adaptation work rarely measures: on a Qualcomm Hexagon NPU, ViT-H and ViT-L compile yet cannot allocate at 1024x1024 image resolution, since activations rather than weights exceed the device ceiling, and quantizing weights does not help. ViT-B alone runs, but the adapted encoder then fails to allocate where the stock one succeeds, until a numerically identical rewrite of the attention lets the complete DCM-SAM run in FP16 at 1024x1024, with no operator falling back to the CPU, masks within 0.01% of pixels of the FP32 reference. Code: https://github.com/MushfiqShovon/DCM-SAM.
cs.LG / 6 / 2609.39047
BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Abstract
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.
cs.LG / 7 / 2609.39588
KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs
Abstract
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.
cs.LG / 8 / 2609.39684
Unapologetically Distributed: A Call for Decentralized Document Analysis
Abstract
Privacy has become an increasingly important concern in the Document Analysis community, to the extent that in many environments such as archives, governmental institutions, and local businesses, the adoption of automation is restricted by legal and policy constraints. While federated learning has often been regarded as a ``necessary evil'', implying an unavoidable performance trade-off in exchange for decentralization and privacy, many prior works overlook its potential to improve robustness to out-of-distribution data. In this paper, we present Unapologetically Distributed, the first comprehensive study evaluating distributed learning in Document Analysis along three key axes simultaneously: the tasks addressed, the architectures employed, and the fine-tuning strategies applied. Specifically, we demonstrate how various distributed training approaches enhance generalization capabilities across diverse tasks such as Table Recognition, handwriting recognition, and Word Spotting, particularly during transfer learning stages. Our results provide strong evidence that decentralization is not merely a constraint, but a valuable opportunity to improve model robustness and adaptability in real-world Document Analysis scenarios.
cs.LG / 9 / 2609.39836
Spherical Interpolation for Backward-Compatible Multimodal Representations
Abstract
Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment can partially mitigate this problem by mapping new-model queries into the old-model gallery space. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether interpolation along the spherical geodesic between these two normalized query representations can improve retrieval without re-indexing the gallery. We characterize when this path contains an interior query direction closer to an idealized retrieval-optimal direction than either endpoint, and connect this characterization to Recall@$K$ through a local margin-based certification result. Experiments across multiple benchmarks and model families show that post-alignment spherical interpolation improves over orthogonal alignment alone, recovering backward-compatibility in most evaluated settings. Consistent with our geometric characterization, per-query oracle analysis shows that retrieval-favorable interior points occur frequently in practice. Code is available at https://github.com/miccunifi/SLERP_backward_compatibility .
cs.LG / 10 / 2609.40347
Image Classifiers are Efficient Self-Supervised Video Representation Learners
Abstract
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to $32\times$ fewer and $160\times$ fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.
cs.LG / 11 / 2609.39350
HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
Abstract
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster heterogeneity commonplace, posing substantial challenges to automatic parallelization. However, existing approaches typically target either MoE architectures or heterogeneous clusters, failing to generalize to scenarios where both challenges coexist. To this end, we present HAPMoE, a heterogeneity-aware automatic parallelism planner for MoE training. HAPMoE builds a lightweight MoE-aware cost model and efficiently searches a six-dimensional parallel space, producing parallel plans directly deployable on Megatron-LM. Experiments show that HAPMoE improves end-to-end training throughput by up to 3.2$\times$ over baselines across heterogeneous clusters. Its non-uniform pipeline partitioning yields an additional up to 78% gains, and its pruning-enhanced dynamic programming algorithm completes the search within 1 minute, demonstrating high efficiency and practical value in complex hardware environments.
cs.LG / 12 / 2609.39481
About the Influence of Workflow Topology on Task Intensity Prediction through Graph Learning
Abstract
Efficient resource provisioning for large-scale workflows on cloud infrastructures is a critical performance engineering challenge. These workflows are often structured as directed acyclic graphs (DAGs), where under-provisioning can cause critical bottlenecks and over-provisioning leads to unnecessary costs. Accurate, task-level prediction of resource intensity (e.g., CPU load and memory usage) is essential for mitigating these issues. While task-level features are commonly used for prediction, the performance impact of the workflow's overall topological structure is often overlooked or assumed. The central question of our work is: To what extent does what part of the DAG topology influence task-level resource intensity, and what is the most effective way to model this influence? This paper presents a comprehensive benchmark to systematically quantify the impact of graph topology on task intensity prediction. We evaluate and compare a spectrum of modeling approaches. Our findings demonstrate that topology is a critical feature for accurate prediction. Models incorporating important topological information, even through simple handcrafted features, significantly outperform baseline models. We show that graph-native models provide the highest accuracy, achieving low mean absolute errors for both CPU and memory predictions, and can still be combined with simple topological features that they do not learn for better performance.
cs.LG / 13 / 2609.40093
Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs
Abstract
Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers traverse CPU memory. However, existing MoE-specialized EP communication libraries assume that direct GPU-to-GPU access is available, largely overlooking consumer GPUs. Therefore, most LLM frameworks instead rely on NCCL, whose CPU-staged communication incurs redundant PCIe transfers and competes with expert computation for GPU resources, limiting their overlap. We present ThunderEP, a novel communication design for such systems that removes the relay hops of traditional ring algorithm, moves data through DMA engines to avoid compute resource contention, and minimizes synchronization latency by reducing the polling overhead of completion flags in CPU memory. We integrate the proposed design into vLLM and evaluate it on three widely used MoE models. Experiments on two PCIe systems equipped with RTX 4090 and RTX 5090 GPUs show that ThunderEP achieves average speedups of 2.00$\times$ and 1.53$\times$ over NCCL for dispatch and combine, respectively, and up to 1.66$\times$ end-to-end speedup over state-of-the-art MoE inference frameworks.
cs.LG / 14 / 2609.40159
Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization
Abstract
Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. Recent graph transformation techniques address these limitations by modifying the dependency graph of the input matrix to improve parallel execution. Existing graph transformation strategies, however, rely on manually designed heuristics, making their development and adaptation to different optimization objectives challenging. This work proposes a reinforcement learning-guided graph transformation framework for SpTRSV, in which graph transformation is formulated as a sequential decision-making problem and an RL agent learns matrix-dependent transformation policies. Experimental results on real-world sparse matrices demonstrate level reductions of up to 94% and reductions of up to 80% in the coefficient of variation of level costs, while modifying only 1.50% of the rows in the highest case. On average, the RL- guided graph transformation achieves a 23% reduction in the number of levels and a 29% reduction in the coefficient of variation of level costs while rewriting only 0.82% of the matrix rows. Although the heuristic strategies generally achieve more aggressive level reduction(between 31% and 46%), the RL-based approach achieves the largest average reduction in the coefficient of variation of level costs, demonstrating its ability to balance competing graph transformation objectives. The results further show that the learned policies can be transferred to previously unseen matrices through curriculum learning and fine-tuning, while zero-shot experiments provide insights into the limitations of generalizing graph transformation policies across different sparsity patterns.
cs.LG / 15 / 2609.38834
Optimal VC Dimension of Contrastive Learning with Margin
Abstract
Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning $d$-dimensional Euclidean representations of $n$-point datasets, $Θ(\min(nd, n^2))$ triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter $α>0$, a triplet $(i,j^{+},k^{-})_α$ is satisfied by the embedding $φ:[n]\rightarrow \mathbb{R}^{d}$, if $\|φ(i)-φ(k)\|_2>(1+α)\cdot\|φ(i)-φ(j)\|_2$. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin $α\in(0,1)$ is in fact $O(n/α^2)$, improving on the previous bound of $O(n\log(n)/α^2)$. We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of $Ω(\frac{n}{α^2})$ (the previously known lower bound was $Ω(\frac{n}α)$), for $α\geq \max(n^{-1/2},d^{-1/2})$.
cs.LG / 16 / 2609.40016
Component-Weighted Centroid Search for Exact Incremental BPE
Abstract
Exact incremental BPE maintains the canonical tokenization state after every appended byte. The recent algorithm of Jiang and Gong (2026) does this in $O(\log^2 t)$ worst-case time, where $t$ is the maximum canonical token length. Its centroid search visits $O(\log t)$ components and can pay another $O(\log t)$ for ordered point location at each one. Within Jiang and Gong's normalized/proper merge-stage model, we change only that local search. Each interval is weighted by the size of the recursive component it selects, so a move from size $m$ to size $m'$ costs $O(1+\log(m/m'))$. These charges telescope, giving $O(\log t)$ time per append and $O(n\log t)$ over an $n$-byte stream, with the same BPE semantics and asymptotic space. We also construct a normalized proper BPE family over a fixed alphabet where count-balanced search uses $Θ(\log^2 t)$ probes on a reachable update, while the weighted search uses $Θ(\log t)$. A Rust implementation matches the predicted probe counts on every tested instance. On ordinary vocabularies the queried degrees are small, however, and the improvement is a worst-case guarantee rather than an average-speed result.
cs.LG / 17 / 2609.40077
Robust and Learned Online Matching in Growing Trees
Abstract
We study irrevocable maximum-cardinality matching in trees revealed by successive leaf attachments, with a known horizon and an exogenous growth law that is misspecified or unknown. For deterministic affine attachment forecasts with nonnegative degree reinforcement, the optimal threshold policy loses at most twice the cumulative expected conditional total-variation error relative to an online oracle knowing the actual growth law. This follows from a unit-span property of the Bellman continuation score and has no additional horizon factor. A four-vertex example attains the coefficient two for the specified deterministic policy, and a two-model argument gives a lower bound linear in the model-error budget for arbitrary policies under general misspecification. For uniform-preferential attachment, the local error has an exact expression through the leaf count. When its constant mixture parameter is unknown, we estimate it from the same growing tree and update the threshold policy at geometric times. A parameter-sensitivity bound for individual Bellman prices and uniform degree-moment estimates yield expected regret $O(\sqrt{n}\log^2 n)$, using $O(n^2\log n)$ arithmetic operations and $O(n)$ stored entries. The exact minimax rate remains open.
cs.LG / 18 / 2609.38552
Fairness Theatre: Evaluating Post-Hoc Fairness Interventions in Vendor-Controlled Early Warning Systems
Abstract
Public institutions increasingly procure AI systems whose design they cannot inspect or change. In higher education, proprietary Early Warning Systems (EWS) leave colleges with few options beyond adjusting model outputs to address inequity. This raises the question of how fairness work is coordinated among vendors, institutions, advisors, and students with unequal power to change these systems? Using student records from a public college in Ontario, Canada, we evaluate six post-hoc fairness interventions on a research EWS under simulated procurement constraints. We compare fairness, accuracy, and demographic disparities, introducing error-type profiling to trace how interventions redistribute false positives and false negatives. Interventions redistributed disparities without consistently reducing them. Two implementations favored already-advantaged groups because they used group size to define disadvantage; small, marginalized groups remained poorly served. These findings show how procurement constraints and implementation choices shape the possibilities for fairness work. We call the resulting condition fairness theatre; dashboard metrics converge while groups' error burdens persist or worsen.
cs.LG / 19 / 2609.39007
RouteRec: Behavior-Guided Sparse Routing for Sequential Recommendation
Abstract
Sessionized interaction histories contain behavioral patterns that can improve sequential recommendation. However, existing models process all sessions through the same parameterized blocks, regardless of their behavioral differences. Mixture of Experts (MoE) enables conditional computation, but it leaves open what should guide expert allocation. We propose RouteRec, a sequential recommender that uses observed session behavior as the routing criterion. RouteRec summarizes four types of behavioral evidence from sessionized histories: interaction tempo, item-group focus, repetition and carryover, and popularity tendency. It uses these cues to route computation at macro, mid, and micro scopes. Cue-derived scores first select expert groups; within each selected group, the current backbone state then refines expert selection. Across six public datasets and 18 dataset-metric combinations, RouteRec ranks first in 12 and second in three, yielding the best overall average rank of 1.61 compared with 4.11 for the next-best baseline. Additional analyses suggest that the behavioral cues guide expert allocation beyond added capacity and produce routing patterns aligned with observed behavior. Our code is available at https://github.com/jy1559/RouteRec
cs.LG / 20 / 2609.38332
Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
Abstract
Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.
cs.LG / 21 / 2609.38342
Activation-Conditioned Self-Distillation
Abstract
On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9\% and LiveCodeBench v6 pass@12 reaches 70.9\%, compared with 69.0\% and 66.3\% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.
cs.LG / 22 / 2609.38348
Function-Space Transformer with Adaptive Anchors
Abstract
Many forms of data, including physical fields, geometric shapes, and visual signals, are naturally described by functions over continuous domains but are observed through discrete samples. Representing these functions on fixed uniform grids imposes a trade-off between resolving localized variation and increasing computation across the domain. Neural operators address this mismatch by learning mappings between functions, while latent-attention architectures provide flexible processing of sampled observations. We introduce the Function-Space Transformer (FST), a framework for learning from functions through a spatially adaptive continuous latent representation. FST stores features at anchors whose locations are predicted from the input observations and recursively refines these anchor features through function-space interactions. This allows the representation to adapt its spatial organization to each input rather than inherit that of the observation grid, while supporting both spatially resolved and finite-dimensional outputs. On PDE solution prediction using PDEBench Burgers and Darcy flow, FST substantially outperforms the Perceiver IO baseline, whose latent representation lacks explicit spatial organization, and is highly competitive with the Fourier Neural Operator. On ImageNet-1K, FST achieves higher classification accuracy than the Vision Transformer baseline, with fewer parameters across these comparisons. Ablations further support the benefits of function-space updates and recursive refinement. Together, these results highlight the potential of adaptive continuous representations for both scientific prediction and visual recognition.
cs.LG / 23 / 2609.38349
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Abstract
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).
cs.LG / 24 / 2609.38356
Continual Learning of Dynamical Systems in Recurrent Neural Networks through Recyclable Unit Gating
Abstract
Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continual DSR (cDSR) requires learning new systems while preserving previously learned dynamics, yet even small parameter updates in recurrent models can qualitatively alter their behavior over long autonomous rollouts. We benchmark established continual learning (CL) methods spanning parameter regularization, replay, and parameter isolation on the fully trainable and interpretable Almost-Linear RNN (AL-RNN). Parameter isolation preserves earlier dynamics most effectively, but excessive task-specific allocations can rapidly exhaust a fixed-size network. We therefore introduce Continually-Recyclable Unit-Gating (CRUG), which conserves capacity through compact allocation and forward transfer. Differentiable gates trained with an $L_0$-based penalty select task-specific units, while unused units are recycled for subsequent tasks. Directed connections allow later tasks to reuse earlier representations without affecting the dynamics of previously committed units. CRUG achieves the strongest reconstruction--capacity trade-off among the tested methods with zero forgetting and reliably learns a heterogeneous sequence of nonlinear and chaotic systems. Furthermore, we show that forward transfer is more pronounced and useful when tasks share similar underlying dynamics. Lastly, we demonstrate that CRUG's advantages extend beyond autonomous cDSR to sequential cognitive tasks.
cs.LG / 25 / 2609.38360
On the Off-Policy Teacher in On-Policy Distillation
Abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
cs.LG / 26 / 2609.38384
LeanPolish: Verified Supervision for Lean Proof Compression
Abstract
Verified proof edits offer a natural source of supervision for improving language-model-generated Lean proofs. Yet verification establishes that an edit is correct, not that its training signal is free of search artifacts. We introduce LeanPolish, a symbolic Lean 4 pipeline that releases 33,402 accepted local edits and 65,596 same-state failed attempts, and use it to study what models learn from this supervision. First-success search admits a goal-independent rule with perfect ranking accuracy; teacher-selected evaluation sites also reward trivial deletions. Continuing menu evaluation beyond the first success removes the ordering shortcut: a trained ranker selects the best candidate on 70.1% of evaluated held-out states, versus 36.9% for the strongest frozen baseline. For compression, iterating the symbolic pass raises miniF2F savings from 19.7% to 27.5%, exceeding the neural hybrids we test there. Verified neural editing helps on other proof sources, but matched frozen-model controls show that its gains need not come from training. The supervision does improve whole-proof rewriting: fine-tuning raises verified token reduction from 2.8% to 5.5% on 19 PutnamBench proofs. Together, the released edits, complete candidate pools, and controlled evaluations separate learning to imitate a search policy from improving on that search. They provide a reproducible basis for studying proof improvement while keeping correctness, compression, and edit policy distinct.
cs.LG / 27 / 2609.38393
Which Tasks Survive Self-Supervised Learning?
Abstract
Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This principle alone, however, does not determine which downstream tasks remain recoverable from the learned representation. We study this question through \emph{semantic recoverability}, defined as the amount of a task's posterior score captured by the represented function space. We show that, for centered and whitened representations, recoverability exactly determines directional class-distance-normalized variance (CDNV), controls few-shot nearest-centroid classification, and governs the strength of task-relevant semantic directions. The population linear probe and centroid axis coincide, and multiple well-recovered tasks approach a factorial centroid geometry. We then analyze a canonical two-view SSL objective and show that its population optimum spans the leading cross-view-stable modes of the associated two-view operator. This yields a closed-form spectral characterization of semantic recoverability: a downstream task is preserved to the extent that its posterior lies in the selected spectral subspace. We validate these predictions on synthetic and real datasets across several SSL methods, testing the predicted relationships among recoverability, directional geometry, spectral structure, and few-shot transfer. Together, these results give a task-level account of what information survives same-instance SSL and how the retained information appears in downstream geometry and transfer.
cs.LG / 28 / 2609.38407
Disagreement-Regularized Imitation Learning for Image-Based Continuous Control with Gaussian and Beta Policies
Abstract
Purpose: Behavior cloning can accumulate errors when a learned controller visits states outside the demonstrated distribution. This study evaluates whether Disagreement-Regularized Imitation Learning (DRIL), which converts disagreement among cloned policies into a reinforcement-learning reward, improves image-based continuous control. Methods: A controlled CarRacing study combines Gaussian and Beta learner policies, demonstrations from either a clipped Gaussian expert or an intrinsically bounded Beta expert, one or 20 trajectories, deterministic and stochastic evaluation, and three retained stages: behavior cloning, the highest 10-episode training-score checkpoint, and the final DRIL checkpoint. The disagreement ensemble contains five Gaussian policies in every variant. Each retained policy is evaluated over 100 procedurally generated episodes. Results: Score-selected DRIL produced its largest gains in the few-demonstration setting, improving over the strongest behavior-cloning mean by 61% with clipped-action demonstrations and by 112% with bounded-action demonstrations. With 20 trajectories, the advantage of DRIL narrowed; in the bounded-action regime, Beta behavior cloning remained about 7% above the best DRIL checkpoint. The experiments also show that the informativeness of the disagreement reward changes with the learner representation and training stage. Conclusion: DRIL can substantially improve few-demonstration visual continuous control, while bounded Beta policies provide strong behavior-cloning performance when more demonstrations are available. The results highlight the joint importance of learner support,ensemble response, and checkpoint selection.
cs.LG / 29 / 2609.38414
Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data
Abstract
Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.
cs.LG / 30 / 2609.38424
Graph Anomaly Detection as Finite-Horizon Control: Training-Free Scoring via Empirical Bayes
Abstract
Node-level graph anomaly detection (GAD) identifies nodes whose attributes and interactions deviate from dominant graph regularities. Existing GAD models encode normality and anomaly scoring indirectly through architectures, message passing, reconstruction or contrastive objectives, and tuned score families. This entangles graph trust (how strongly graph structure should define normality), graph-spectral weighting, and anomaly-score choice, yielding scores that are costly, opaque, and unstable across graph regimes. We propose EB-GAD (Empirical-Bayes GAD), a training-free framework that models normality as graph-aware generalized Ornstein-Uhlenbeck (GOU) relaxation toward a graph-filtered template. Empirical Bayes fits the graph precision from the residual-field likelihood; the GOU then turns scoring into a closed-form finite-horizon control energy, the minimum effort to steer a feature-neutral node to its observed endpoint along graph-spectral relaxation. Sweeping relaxation horizon and endpoint tolerance yields a bank of scores that share one fitted prior: equilibrium Mahalanobis scoring is one limit, while finite-horizon control-energy and scale-normalized ratio scores reveal anomalies that static equilibrium scoring can mask. A label-free selector chooses the score family from feature homophily, edge density, and feature dimension, then ranks candidates by fitted-null deviation and rank stability. On 11 benchmarks and without labels at any step, EB-GAD has the best or tied-best AUROC on 9: the four financial fraud networks (up to 3.7M nodes), the YelpChi and Amazon review graphs, Weibo, Reddit and Facebook, with margins of up to 21.7 points. It is second on BlogCatalog and ACM.
cs.LG / 31 / 2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
Abstract
A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.
cs.LG / 32 / 2609.38447
An Input-Frugal Deep Learning Framework for Weather-Driven National Crop-Yield Forecasting: A Case Study of Brazilian Soybean
Abstract
Reliable, timely crop-yield forecasts are essential for market stability and risk management, yet many approaches rely on costly or hard-to-scale inputs. We present a frugal, transferable, and architecture-agnostic deep learning framework that uses routine weather as the only time-varying input plus two lightweight static context inputs (crop year and an agro-environmental label) to capture long-run change and regional heterogeneity, while supporting multiple sequence encoders under identical data requirements. Using a 20-season Brazilian soybean case study (2001/02-2020/21) with leave-one-year-out cross-validation, we benchmark MLP, CNN, LSTM, CNN-LSTM, a Transformer encoder and the Mamba state-space model against linear ridge regression and a five-year moving-average "farmer" baseline. All deep learning variants outperform ridge, and all sequential encoders surpass the non-sequential MLP. The Transformer achieves the best national accuracy (RMSE 149 kg ha^-1; rRMSE 5.3%; R^2 = 0.784), reducing error by 47.6% relative to the farmer baseline. In-season forecasts improve monotonically from early- to late-season issuance, reaching approximately 50% lower error than the baseline at the latest forecast point. Ablations indicate that the agro-environmental label and spatial instance expansion (multiple grid-node weather sequences per municipality-year) contribute positively without increasing input complexity. SHAP diagnostics suggest crop year explains most of the long-run trajectory, whereas within-season weather and agro-environmental context primarily drive interannual deviations, with moisture/cloud and thermal-demand variables dominating. Overall, the framework is straightforward to deploy across other crops and geographic regions and is naturally compatible with operational weather forecasts for routine monitoring.
cs.LG / 33 / 2609.38453
Grokking through the Lens of Minimum-Norm Interpolation
Abstract
Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical theory that characterizes how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero--one generalization law and construct a family of convex norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error equal to $0$. Furthermore, when feature dimension and sample size are proportional, we provide a precise characterization of training and generalization errors along $\ell_r$-regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and $\ell_1$ regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate the generality of our theoretical predictions. Finally, beyond grokking, our work reveals a statistical instability in minimum-norm interpolation: small perturbations in the regularization strength can lead to drastically different generalization, while preserving small training error.
cs.LG / 34 / 2609.38465
Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that reducing these metrics improves the downstream understanding-generation trade-off has never been tested directly. We audit it in a controlled testbed, GRIDUMM, which mirrors key structural ingredients of UMM training while making the ground-truth trade-off exactly computable. Across 63 configurations and 372 measured checkpoints, no directional conflict metric reaches an absolute Spearman correlation of 0.3 with a confidence interval excluding zero for conflict measured during training against the eventual trade-off. A dose-response intervention that monotonically suppresses conflict leaves the trade-off flat, separating correlation from causation. The norm ratio is a generation-failure detector and becomes null among configurations that master generation. Functional interference measures outperform directional conflict metrics, while training loss tracks the trade-off strongly. Our results do not show that conflict is useless; they show that its validity as a diagnostic target must be established, not assumed, and we release the audit protocol as a reusable standard.
cs.LG / 35 / 2609.38484
RetroGEF: Dynamic Graph Edit Flow for Single-Step Retrosynthesis
Abstract
Retrosynthesis enables the discovery of viable synthetic routes to target molecules. It plays a central role in modern drug discovery and materials design. Retrosynthesis involves molecular graph transformations that can change both connectivity and graph size. These transformations may introduce reactant components absent from the target while revising the product-derived structure. To model these transformations, we propose RetroGEF, a flow-based generative model for single-step retrosynthesis. Starting from the target molecule, it constructs possible reactants by adding atoms and changing bonds in the molecular graph. RetroGEF models molecular transformations and changes in graph size within the same generative process, rather than relying on a fixed-size graph canvas. It learns this process directly from product--reactant pairs without requiring a prescribed edit order. Experiments on representative retrosynthesis benchmarks demonstrate that RetroGEF achieves state-of-the-art performance.
cs.LG / 36 / 2609.38491
Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models
Abstract
Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.
cs.LG / 37 / 2609.38506
Autoregressive Frontier Expansion: Growing Trees with Graph Machine Learning
Abstract
Tree-like branching structures are common in nature, from botanical trees to neurons, blood vessels and respiratory trees. Their branching shape often reflects function, making structural modelling central to understanding how these systems work. Because acquiring real-world 3D data is often expensive or infeasible, realistic generative models are valuable for simulation and data augmentation. Existing morphology-specific models either constrain how topology is generated or rely on hand-tuned, mechanistic procedures. Generic 3D graph generators, by contrast, do not exploit or enforce the structure of trees. We propose Autoregressive Frontier Expansion, a generative framework that constructs trees through an iterative expansion process, simulating the biological growth of real trees. At each step, a flow-matching model parameterised by an SO(2)-equivariant GNN expands the frontier by predicting whether each active branch bifurcates or terminates. We evaluate our method on cortical neurons and botanical trees in unconditional, class-conditioned, and morphology-guided generation. Across both domains, the generated morphologies agree closely with the reference distributions and, in conditional experiments, with the specified targets.
cs.LG / 38 / 2609.38517
Does Text Steer Neural PDE Surrogates? A Controlled Diagnostic with OperatorCLIP
Abstract
Lower error from a text-conditioned neural surrogate does not, by itself, show that the model uses the meaning of the text. We examine this attribution problem with OperatorCLIP, comparing an unconditioned FNO, a constant-sentence FiLM control, and a fixed task description trained with contrastive alignment. Three-seed experiments cover Darcy2D, ShallowWater2D, and three-dimensional compressible Navier-Stokes (CNS3D). Constant conditioning has lower mean test error on both 2D tasks. Relative to this control, task text plus alignment has a similar mean on ShallowWater2D and CNS3D and a higher mean on Darcy2D; these descriptive comparisons have substantial seed uncertainty. The latter comparison changes both prompt content and loss, so it isolates neither effect. The text encoder is trained from scratch, and each conditioned model sees only one description during training. In this regime, pairwise InfoNCE cannot identify matched pairs and has minimum $\log B$. Prompt interventions show no reliable semantic ordering. This methodological caution demonstrates why pathway controls are needed; it neither establishes semantic competence of the encoder nor tests the effectiveness of text under varying physical context.
cs.LG / 39 / 2609.38521
ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights
Abstract
We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker product to the empirical Fisher information matrix of a small calibration set by Kullback--Leibler minimization, forming a Mahalanobis reconstruction loss from the result. The continuous ADMM updates from NanoQuant become solutions to Sylvester equations, while its discrete projection and deployment format remain unchanged. Because the curvature is local to a given set of weights, ShamAN-Q re-measures the input curvature statistic for each layer immediately before layer factorization, periodically refreshing all statistics on the partially quantized model. ShamAN-Q also redistributes the uniform rank from NanoQuant across layers at the same total number of bits. On Qwen3-Base, ShamAN-Q lowers WikiText-2 perplexity at $\approx$1 bpw from 27.56 to 22.96 (0.6B), 19.21 to 16.72 (1.7B), and 14.29 to 13.80 (4B) while matching or improving zero-shot accuracy on the Eleuther LM Evaluation Harness. On 0.6B, ShamAN-Q at $\approx$0.8 bpw matches the published perplexity of NanoQuant at $\approx$1.0 bpw.
cs.LG / 40 / 2609.38526
Revisiting scaling laws for reward optimization
Abstract
Scaling laws for optimization against reward models in AI alignment have pinned down how performance depends on optimization effort---measured by a KL-divergence budget relative to a reference policy. Beyond a certain budget, over-optimization (or reward hacking) can arise: because we optimize against a proxy reward model (distinct from true rewards), performance can plateau or degrade. Naturally, the proxy reward's accuracy depends on how much preference data (often in the form of pairwise comparisons) was used to train it. However, existing research does not cleanly identify how performance jointly scales with the amount of training data and the divergence budget. Our main contribution is to provide an empirically accurate and theoretically grounded scaling law in such context. Performance roughly scales as $Θ(\sqrt{\min\{\log(M),K\}})$, where $M$ is the number of comparisons in training data and $K$ is the policy's divergence budget. We develop an information-theoretic model to establish this upper bound and prove it is tightly achievable through a constructive procedure. Informed by this, we conduct extensive empirical evaluations using a real-world annotation setup, whereby a large 70B gold reward model generates feedback data and proxy reward models are trained from less capable models (0.6B to 4B). Our scaling law provides an excellent fit (R2 from 97\% to 99\%), outperforms alternative specifications, and remains robust across model sizes, noise, and optimization procedures (best-of-$N$ or policy tilting). Our evidence suggests that reward optimization is analogous to a surprisingly simple selection task: choosing from a sequence of IID Gaussian random variables using noisy preference feedback.
cs.LG / 41 / 2609.38538
ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations
Abstract
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.
cs.LG / 42 / 2609.38540
When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions
Abstract
A valid curvature upper bound need not justify either a robustness certificate or an intervention on an intrinsic predictor property. We demonstrate this distinction for a last-layer relative-flatness proxy used in both settings. First, empirical-risk stationarity does not eliminate pointwise first-order loss terms: at a finite global empirical-risk minimum, the retained certificate expression underestimates a loss increase by over $210\times$. We derive a globally valid, gauge-invariant feature-space repair. Second, common-row softmax shifts preserve predictions and the exact contraction while making the proxy unbounded. Even standard reference-class choices double it on average relative to the centered representation. For a single fixed-feature example with at least three classes, scalar retuning generically cannot align the induced probability updates. Row centering gives the orbit-minimized bound and restores value and full-model gradient invariance under this symmetry. Across 45 paired one-step tests on algorithmic and image models, amplified shifts separate raw-regularized predictors while quotient-regularized predictors remain aligned. Long-horizon CIFAR-10 experiments show substantial, reversible suppression of generalization, while evidence for selective delay after memorization is less consistent. Together, these results show that validity as a curvature upper bound does not by itself justify either inversion into a robustness certificate or differentiation into an intrinsic training intervention.
cs.LG / 43 / 2609.38547
Towards Universal Wasserstein Barycenters through Flow Matching
Abstract
Defining a weighted mean over probability measures under probability metrics is a central tool in probabilistic machine learning. Under the Wasserstein metric, these are called \emph{Wasserstein barycenters}. While most approaches compute barycenters for a fixed weight vector, approximating the whole family of barycenters over the simplex, which we call the \emph{Wasserstein simplex}, remains underexplored. We refer to this problem as \emph{Universal Barycenter Approximation}, and propose \texttt{BaryFM}, a flow matching model transporting the marginal measures into any barycenter in the Wasserstein simplex. Once trained, the network can draw samples from measures in the Wasserstein simplex through an ordinary differential equation. We validate our method on 4 downstream tasks: domain adaptation, generalization, Bayesian posterior aggregation and algorithmic fairness. \texttt{BaryFM} achieves the best average rank among 15 competing methods across 10 domain adaptation benchmarks, matching or surpassing non-universal solvers.
cs.LG / 44 / 2609.38565
The Advantages of Fresh Sketching for Ridge Regression
Abstract
Over the past 25 years, sketching and sampling have become widely used tools for accelerating large-scale regression. In iterative randomized solvers, a basic design choice is whether to $\textit{reuse}$ the same sketch or draw $\textit{fresh}$ randomness at every step. For (under-constrained) iterative ridge regression with column sampling, whether fresh sketches offer provable advantages has remained open: $\textit{We show that they do.}$ Fresh sketching lets us analyze error only along the current residual solution, rather than uniformly over the entire Gram matrix. This directional view yields sharper convergence guarantees for leverage score and ridge leverage score sampling and, more importantly, leads to residual-aware sampling rules. By minimizing the variance of the relevant sketched matrix-vector product, we derive an oracle distribution and practical approximations to the oracle distribution, including a mixture sampling distribution with (somewhat weaker) convergence guarantees. Experiments on synthetic and real data, including ridge probes on Qwen2.5 representations, support our theory, showing substantially faster convergence.
cs.LG / 45 / 2609.38587
NeurDuo-EEG: A Long-Sequence EEG Foundation Model with Persistent State and Explicit Memory
Abstract
Electroencephalography (EEG) is recorded continuously over hours, with relevant dynamics spanning timescales from milliseconds to hours. Most EEG foundation models nevertheless process fixed windows independently, limiting their ability to capture information encoded in long-timescale dynamics. State-space architectures enable persistent recurrent processing, but long-range information remains implicitly compressed in recurrent states. We present NeurDuo-EEG, a causal EEG foundation model with channel-resolved persistent memory. NeurDuo-EEG introduces multi-timescale memory management with learned consolidation and selective retrieval, enabling persistent modelling of continuous EEG with fixed-size state. It is pre-trained on 3,955 hours of EEG from 17 public datasets using multichannel autoregressive prediction of discrete spectral codes. Across three short-window and two long-sequence downstream tasks, NeurDuo-EEG achieves the best performance on four of five benchmarks, including all three short-window tasks and seizure detection, where AUC-PR improves from $0.285$ to $0.471$ over the strongest non-NeurDuo baseline. NeurDuo-EEG also remains competitive on sleep staging and supports efficient streaming inference, with nearly constant per-chunk latency as the available history grows to one hour. Notably, the Small variant achieves this with only 4.7M backbone parameters. These results demonstrate the value of persistent, multi-timescale modelling for both long-sequence and short-window EEG analysis. Our code is available at https://github.com/YifaNNW/NeurDuo-EEG.
cs.LG / 46 / 2609.38598
Reinforcement Learning with Complex (valued) Memories
Abstract
Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 $\times$ the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: https://github.com/Sathya98/qurl
cs.LG / 47 / 2609.38608
Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases
Abstract
When can we learn from biased samples? We study regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes themselves, a ubiquitous challenge spanning clinical trials with patient dropout, labor markets with self-selection, and auctions with strategic entry. Ignoring such selection yields systematically biased conclusions with real-world consequences. This challenge has a long history in econometrics and statistics, starting with Heckman's seminal two-stage model and followed by numerous generalizations. While these works provide various sufficient conditions for identification, a complete characterization of when such regression is possible has remained elusive. In this work, we provide a characterization for when regression is possible in the presence of sample selection bias. Our results establish the minimal assumptions required on the functional forms of selection processes under which regression remains possible, which are particularly relevant in modern settings where selection mechanisms are increasingly complex and opaque. As a corollary of our characterization, we show that there are settings where the regression function can be identified even when the selection filter itself cannot. This observation already goes beyond the ``estimate selection filter, then debias regression'' paradigm that is followed by virtually all existing approaches. Under natural strengthenings of our identification conditions, we also establish finite-sample estimation guarantees with explicit convergence rates and provide oracle-efficient algorithms. This yields the first general-purpose estimation method for this broad class of selection problems. Finally, we explore the implications of our results for several well-studied econometric settings with complex selection mechanisms such as auctions with entry costs and labor markets.
cs.LG / 48 / 2609.38618
Differentiable Structure Learning for Cyclic Linear Gaussian Models with Latent Confounders
Abstract
We study causal structure learning from observational data in linear Gaussian structural causal models in the presence of directed cycles and an unknown number of exogenous latent confounders, bounded by a given maximum. We derive the covariance of the observed variables and introduce marginal quasi-equivalence, which characterizes when different causal models share a full-dimensional subset of the observational distributions they can generate. We formulate structure learning as minimization of the Gaussian negative log-likelihood with a logarithmically scaled complexity penalty that counts directed edges and latent variables. For a fixed number of observed variables and a fixed upper bound on latent variables, we establish consistency of global score minimizers up to marginal quasi-equivalence under algebraic faithfulness, structural minimality, and model-overlap assumptions. We parameterize the inclusion of directed edges and candidate latent variables using Bernoulli gates, whose continuous probabilities are optimized jointly with the structural coefficients. Averaging the penalized negative log-likelihood over these gates yields an objective with a closed-form differentiable complexity penalty. We prove that this expected objective has the same global infimum as the corresponding discrete structure-learning objective. Experimental results show that our approach achieves lower recovery error than previous methods in several experimental settings.
cs.LG / 49 / 2609.38623
Geometry-physics confounding impairs PDE learning across varying domains
Abstract
Learning partial differential equation (PDE) dynamics across varying domains is central to predictive modelling and data-driven discovery of governing equations. However, geometric variation alters both field representation and the governing differential operators, confounding geometric effects with intrinsic physical properties in the observed dynamics. This work identifies geometry-physics confounding as a unified failure mechanism for PDE learning across varying domains. In forward operator learning, this confounding increases the burden of inferring geometry-dependent operator changes from finite data, reducing data efficiency and generalisation. In equation discovery, omitting geometry-induced operators misspecifies the candidate library, leading to biased parameters, missed governing terms and spurious terms. We propose a de-confounding framework that makes the known geometry-to-operator transformation explicit. Geometry-induced coefficient fields improve prediction and data efficiency across five operator-learning benchmarks, while geometry-complete candidate libraries recover the generating equations and reduce held-out PDE residuals by more than two orders of magnitude in both evolving-domain systems. By separating known geometric action from intrinsic physics, the proposed framework supports more reliable and data-efficient PDE learning across scientific and engineering problems with varying geometries.
cs.LG / 50 / 2609.38625
Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations
Abstract
Concept Bottleneck Models (CBMs) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remains unclear, with existing studies reporting mixed and sometimes contradictory findings. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves. To disentangle these factors, we introduce a generator-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space. Within this framework, we evaluate robustness both empirically, via prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces. Across experiments, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability robustness trade off that depends critically on the perturbation regime and task structure. Together, our results show that interpretability and robustness are distinct objectives: interpretable intermediate representations do not uniformly improve robustness, but instead redistribute sensitivity across perturbation spaces and model families.
cs.LG / 51 / 2609.38645
Alignment via Training Against Probes Without Losing Monitorability
Abstract
Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes that do not update during training is an easily exploitable objective, while continuously updated probes substantially reduce harmfulness and improve honesty while preserving utility. Probe-guided fine-tuning achieves better safety-utility trade-offs than DPO and inference-time steering, while being substantially more robust against jailbreak and abliteration attacks. Moreover, the concepts stay linearly encoded after fine-tuning, meaning oversight is not lost by our method. Training against probes thus offers a way to shape what models represent rather than only what they output, which may become increasingly important as models get better at making their outputs look aligned.
cs.LG / 52 / 2609.38647
Uncertainty-Normalized Margins for Direct Preference Optimization
Abstract
Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.
cs.LG / 53 / 2609.38666
Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
Abstract
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
cs.LG / 54 / 2609.38673
In-Distribution Imagination for Model-Based Offline Reinforcement Learning
Abstract
Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Experiments show that trajectory support predicts rollout failure substantially better than transition-level uncertainty, highlighting the importance of trajectory-level rollout control in MBORL.
cs.LG / 55 / 2609.38744
Molecular Property Prediction under Structural Shift with Tabular Foundation Models
Abstract
Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials design. Tabular foundation models (TFMs) offer a promising approach through in-context learning, but their performance under structural shifts and the value of molecular comparisons in this setting remain underexplored. We study structural generalization in molecular property prediction and introduce MolPAIR (Molecular Pair-Augmented In-context Refinement), a framework that combines molecule-level and molecular-pair contexts without task-specific parameter updates. A global tabular foundation model (TFM) first predicts a query's property from labeled molecular examples. A second frozen TFM predicts differences in prediction errors between the query and labeled reference molecules, using these comparisons to refine the initial prediction. Across 58 MoleculeACE and Polaris tasks, CheMeleon representations combined with TabPFN-3 already outperform each evaluated baseline on a majority of tasks. MOLPAIR further improves this predictor on 46 of 58 tasks, with gains across four molecular representations and three TFM backbones. These results show that explicit molecular comparisons can strengthen tabular in-context learning for structural generalization while keeping the molecular encoder and pretrained model weights fixed. The code and datasets are available at https://github.com/nums-ai/MolPAIR.
cs.LG / 56 / 2609.38764
Lasting Effects of Abstract Pretraining Beyond Perplexity
Abstract
Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, such a warm-up improves specific capabilities that are not reflected in language-modeling perplexity. Our warm-up uses an abstract stack-manipulation task that requires compositional and state-tracking capabilities. Allocating as little as 1% of pretraining tokens to this data improves multi-hop question answering by up to 3.9 F1 points on MUSIQUE, with additional gains on HOTPOTQA and 2WIKIMULTIHOPQA despite comparable language-modeling perplexity. Controlled experiments show that the warm-up substantially accelerates the acquisition of deeper reasoning chains. We also explore what drives this transfer. First, the structure of the data matters: replacing the stack task with a queue fails to produce the same gains. Second, the gains are specific: performance improves on sequential reasoning chains, with no consistent benefit on tasks that combine or compare independent facts. Third, timing matters: mixing abstract data with natural language is far less effective than an initial dedicated phase, and exposure after pretraining completely removes the benefits. The early advantage persists through billions of subsequent language tokens. These results show that early abstract training can reliably shape the capabilities language models later acquire.
cs.LG / 57 / 2609.38767
dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale
Abstract
Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Most scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility. We introduce dattri-LLM, a TDA library that makes gradient-based attribution more practical at scale. For efficiency, dattri-LLM uses compact gradient representations and dynamically routes gradient operations based on a cost model. For compatibility, its capture mechanism collects per-example gradients from existing training loops that call backward(), without requiring changes to the loop or its configuration. This includes distributed training with DDP and FSDP and pipelines built with HuggingFace Transformers, TRL, and OLMo. For extensibility, dattri-LLM exposes reusable gradient operations and training-time callbacks for implementing attribution methods and applications. These interfaces support a variety of attribution methods, including gradient similarity, curvature-based influence, and trajectory-based methods, as well as applications that act on gradients during training, such as online data selection. On the same hardware and workload, dattri-LLM achieves 3.2x the throughput of the fastest competing library on average, scales multiple attribution methods to 110B-parameter models across four H200 GPUs, and offers superior attribution fidelity-cost trade-offs across a range of models with different model families and scales.
cs.LG / 58 / 2609.38768
Learning Under Forgetting: Statistical Support-Selective Retention in Stochastic Training Dynamics
Abstract
Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and compression-like effects during training, but a unified dynamical account of selective retention remains incomplete. We propose Repeated Reinforcement with Persistent Forgetting (RPF) dynamics, a minimal framework in which repeated exposure reinforces patterns and structures that recur in the data, while persistent forgetting attenuates learned information. This view treats forgetting not merely as a failure mode, but as a selection mechanism. We build the theory in three successive layers. First, in an independent-feature model, we derive an exposure-selective survival law and a support-dependent retention boundary characterizing which patterns persist under forgetting. Second, in a shared-parameter model, we show that forgetting induces spectral filtering over covariance modes, preserving strongly supported shared components while suppressing weak ones. Third, under small-step and norm/coding approximations, we show how RPF dynamics induce an implicit trade-off between data fitting and the cost of stored information, yielding Minimum Description Length (MDL)-like compression. Controlled experiments provide evidence for this reinforcement--forgetting selection mechanism in scalar memories and a nonlinear shared network. Joint reinforcement and attenuation interventions shift conditional retention, while matched exposure counts reveal forgetting-dependent effects of reinforcement timing and changes in the composition of the retained set. Together, these results show that repeated reinforcement and persistent forgetting jointly provide a controllable source of inductive bias beyond neural architecture and scale.
cs.LG / 59 / 2609.38784
Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?
Abstract
In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and class concentration grow jointly. We prove that the score retains a class-dependent leading angular gain $g_c=A_c/τ$, where $A_c$ is the mean resultant length. This gain enters Softmax competition, pairwise decision boundaries, and feature gradients. On real CIFAR-LT, ImageNet-LT, and iNaturalist representations, the theory accurately predicts boundary movements and local gradient changes under the full vMF score. Classwise temperature adjustment also changes the cosine-zero intercept and finite-dimensional response. We construct intercept-preserving and Pure Angular controls to separate the leading gain from these accompanying changes. Complete gain equalization yields a shared-scale cosine prototype rule at leading order; a finite-dimensional margin condition guarantees agreement of the two classifiers. Across 16 frozen representation settings, prediction agreement is 98.43-99.99%, with disagreements concentrated at small cosine margins. In controlled contrastive-only training with the training-frequency prior, Pure Angular editing improves both learned representations at all tested CIFAR-10/100 imbalance factors and retains positive changes on ImageNet-LT. Thus vMF concentration not only describes class distributions, but also forms a decision and learning scale in high-dimensional probabilistic contrastive learning.
cs.LG / 60 / 2609.38785
How Accurate Is Accurate Enough?
Abstract
How accurate must a numerical approximation be within a learning system? Primitive error alone cannot answer this question: errors of the same magnitude can have very different consequences for losses, predictions, and gradients at different learning states. We study this question through the learning objective itself. The objective weights classwise numerical errors nonuniformly according to the current state, so the importance of an error depends not only on its magnitude but also on the class it affects and the weight that class receives. For softmax cross-entropy, we characterize this coupling between class weights and errors and derive the exact extrema of the signed loss change over pairings of fixed non-target probability and score-error multisets, with the target probability and target score error held fixed. Building on this structure, we establish finite-error guarantees that propagate primitive error to losses, probabilities, predictions, and feature gradients, then invert these guarantees to obtain a certified primitive tolerance for the current state under prescribed learning-level error requirements. We give a complete instantiation of the framework in high-dimensional von Mises-Fisher learning. Controlled interventions and a large collection of saved learning states show that identical primitive error can produce substantially different learning consequences, while certified numerical tolerances vary by orders of magnitude across states under the same learning-level requirements. These results show that the adequacy of a numerical approximation must be assessed in relation to the current learning state and the quantity to be preserved; numerical accuracy should itself be treated as part of the learning objective.
cs.LG / 61 / 2609.38786
Same Loss, Different Gradients
Abstract
Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the backward pass describe the same mathematical object. We show that this correspondence can fail when probabilistic objectives rely on finite special-function recurrences, custom backward rules, and numerical clipping. In high-dimensional von Mises-Fisher learning, real numerical implementations can produce identical forward scores and losses at the same learning state while supplying different gradients and following different optimization trajectories. We characterize the structure of this mismatch in finite-start Bessel recurrence and show that classwise radial mismatch can compose through probabilities into a locally nonconservative update field. Evaluating the accuracy of special-function values and derivatives separately is therefore insufficient to characterize the realized learning objective. Motivated by this observation, we introduce AR/FR, a fixed-depth analytic realization that constructs a potential and its derivative jointly, ensuring forward-backward coherence by construction. We establish a uniform cubic-order error bound relative to the exact Bessel ratio over the entire nonnegative concentration axis and propagate this guarantee to learning scores and objectives. As representation dimension increases, the original finite recurrence becomes sequentially deeper, whereas the worst-case AR/FR error guarantee tightens cubically, jointly providing coherence, certified fidelity, and fixed-depth computation. These results suggest that a differentiable numerical primitive is defined by both the values it realizes and the derivatives it actually supplies to the optimizer; together, they constitute the numerical realization of the learning algorithm.
cs.LG / 62 / 2609.38789
LEARN-TS: LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Multivariate Time-Series Anomaly Detection
Abstract
Reconstruction errors in multivariate time-series anomaly detection may not reliably distinguish abnormal behavior from benign deviations. Language-derived semantics offer complementary context, but existing multimodal approaches may rely on time-associated paired textual information that is difficult to obtain consistently and is not provided by standard multivariate time-series anomaly detection benchmarks. This setting poses two challenges: (1) conditioning masked reconstruction on window-specific semantics without exposing exact numerical targets or anomaly-specific cues, and (2) using a window-independent concept of normality as a complementary semantic reference rather than an independent anomaly detector. We propose LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Time Series (LEARN-TS), which uses a frozen language model to construct two role-separated semantic representations without requiring temporally paired external text. Window-specific observation semantics encode temporal and cross-variable context without exact numerical values to guide channel-shared patch-masked reconstruction. A fixed, dataset-agnostic normality prompt provides a window-independent semantic reference for aligning normal representations and estimating normality discrepancy. At inference, masking each temporal patch once yields timestamp-level reconstruction evidence, conditionally modulated by discrepancy from a separate unmasked view. Across four benchmarks, LEARN-TS achieves the highest mean performance in 13 of 16 dataset-metric comparisons. Controlled ablations examine observation conditioning, joint normality alignment and scoring, and reference content, showing dataset-dependent ranking benefits and modest average gains from semantic over random references.
cs.LG / 63 / 2609.38797
Evaluating Persistent Calibration under Evolving Model Knowledge
Abstract
As AI systems move from static repositories to agents that are capable of continual adaptation and learning, maintaining their trustworthiness means equipping the models backing them with the ability to produce confidence estimates that dynamically reflect their changing skills and knowledge. We introduce the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. We operationalize this by examining persistent calibration across checkpoints of open models, asking whether confidence estimators trained on earlier checkpoints can generalize to later ones. Specifically, we aim to shed light on whether confidence is dependent on knowledge, a question with implications for the reliability of confidence estimates. To measure this relationship, we define and evaluate calibration on knowledge contrast sets: subsets containing questions that one checkpoint answers correctly and another checkpoint answers incorrectly, reflecting a change in knowledge. We show that both inference-time and fine-tuning methods fall short on contrast-set calibration compared to oracle methods trained on future checkpoints, even for methods that are well-calibrated on the full dataset. We provide evidence for the hypothesis that persistent calibration is challenging because there is a vast space of possible confidence functions that are well-calibrated on a given checkpoint, out of which only some rely on meta-knowledge features that would generalize to other checkpoints. Towards improving contrast-set calibration, we show that multi-checkpoint training helps, suggesting an avenue for identifying confidence features that remain robust across changing knowledge.
cs.LG / 64 / 2609.38805
Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents
Abstract
LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation. While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task. In this work, we study explicit trajectory diversity in RL-based post-training for LLMs. Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretable behavioral representation, and then measure diversity as a set-level functional over the resulting descriptor matrix. Building on this formulation, we introduce Trajectory-guided Joint Policy Optimization(TJPO), a single-policy framework that optimizes explicit diversity over sampled trajectory groups, avoiding the need for population-based policy training, and instantiate it within group-based policy optimization through trajectory-level learning signals. This design makes the diversity objective both interpretable and controllable. Experiments on Sokoban and ALFWorld show that TJPO improves task-specific trajectory diversity while maintaining competitive task performance. Descriptor and trajectory analyses show that the learned variation follows the specified behavioral dimensions and includes distinct successful strategies. Extra experiment results suggest that explicitly shaping trajectory diversity can help LLM agents satisfy user requirements and remain effective when task conditions change.
cs.LG / 65 / 2609.38814
Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers
Abstract
Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed during training? Here we study small autoregressive transformers trained from scratch on trajectories sampled from restricted parameter regimes of several non-linear dynamical systems, including logistic and sine maps, the Lorenz system, and the generalized Hopf system, with control parameters and state trajectories represented as sequences of continuous tokens. Under closed-loop evaluation at parameters far outside the training distribution, the models can recover self-similar period-doubling cascades, chaotic dynamics, and attractor structures with remarkable visual and numerical fidelity. For the logistic map, a transformer reproduces successive period doublings up to period 128, yielding a finite-order scaling ratio of 4.6687, matching the Feigenbaum constant to within $5\times10^{-4}$. We further investigate how these structures emerge over the course of training, and reveal with causal interventions how control-parameter information is processed through attention into state prediction and shapes the resulting closed-loop dynamics. These results suggest that a surprisingly narrow window into a system's local behavior may suffice for autoregressive transformers to generalize to its unseen global dynamical organization.
cs.LG / 66 / 2609.38833
ReSCENE: Server-Side Replay for Structural Mitigation of Catastrophic Forgetting in Federated Continual Learning
Abstract
Federated continual learning must integrate new tasks over time without losing earlier-task knowledge. Most existing methods attach an anti-forgetting mechanism to the client-trained, server-aggregated loop of federated learning, which holds back new learning to preserve earlier knowledge and burdens resource-constrained clients. We propose ReSCENE, which structurally mitigates catastrophic forgetting by having each client upload a small condensed surrogate of its local data while the server keeps the surrogates of past tasks and trains the global model on them together with the current task surrogates. For efficient server memory, we introduce temporal herding, which selects the more recent surrogates from the pool accumulated over a task into a compressed buffer. Our study provides a theoretical analysis showing that this buffer can represent the original task data more closely than full accumulation of all surrogates. Across CIFAR-10, CIFAR-100, and TinyImageNet, ReSCENE achieves the strongest accuracy over seven baselines, by up to $31.1$ points of average accuracy, while requiring as little as $0.11\times$ of the client computation and up to $179\times$ less upload than the model-update baselines. ReSCENE further demonstrates its effectiveness when scaled to larger client populations and larger models while remaining efficient, which makes it a practical method for federated continual learning.
cs.LG / 67 / 2609.38840
scTrilemma: Balancing Identity, Invariance, and Fidelity in Single-Cell Representation Learning
Abstract
Single-cell RNA-seq representation learning is fundamentally label-free: cell identities, states, and contexts are not fixed training targets, so what constitutes signal or nuisance is analysis-dependent. A single representation must therefore preserve biological identity and state, remain robust to nuisance context, and retain the gene-level variation needed for expression analysis, three demands we call the representation trilemma. To tackle this problem, we introduce scTrilemma, a latent-bottleneck VAE that routes expression-derived variation to the embedding, the decoder, or the prior rather than forcing all of it through one embedding. It gates gene tokens by expression, routes the cell representation through the decoder, and conditions the prior on unlabeled pseudo-bulk context, under a single reconstruction objective and without target annotations or auxiliary representation losses. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, scTrilemma leads all three demands at once and preserves biological-state, differential-expression, and pathway structure across multiple disease settings. Latent interventions further show that context can be removed at almost no cost to the other demands, leaving identity against fidelity as the remaining tension. Code is publicly available at https://github.com/yunhak0/scTrilemma.
cs.LG / 68 / 2609.38847
Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics
Abstract
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric
cs.LG / 69 / 2609.38854
Mitigating the Length-Scaling Tax with Online Distillation
Abstract
Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.
cs.LG / 70 / 2609.38863
GeoNest: Learning to Select Failure-Aware Neighborhoods for the Irregular Knapsack Problem in a Circular Container
Abstract
The two-dimensional irregular knapsack problem in a fixed circular container is an important combinatorial optimization problem for maximizing material utilization in manufacturing. Conventional geometric packing solvers can produce tightly packed layouts, yet they often partition the residual space into isolated small pockets that cannot fit valuable unplaced polygons. To overcome this late-stage packing bottleneck, we propose a failure-aware large neighborhood search framework named GeoNest, driven by a graph policy trained via reinforcement learning. Specifically, we first construct neighborhoods by pairing failed target polygons with residual pockets. We then use explanatory poses to identify the placed polygons that block candidate insertions. These diagnosed blocking relations define bounded, fixed-item repair subproblems for the underlying geometric solver. Finally, the graph policy selects the most promising subproblem for execution. For evaluation, we introduce CircleNest-Bench, a benchmark comprising 2,391 load-controlled instances from four contour sources, including a held-out industrial CAD source. Experimental results demonstrate that, under the same total time budget, GeoNest improves mean utilization over a state-of-the-art standalone packing solver by about 0.9% on average across the three main test sets and by about 0.6% on the held-out industrial set.
cs.LG / 71 / 2609.38877
On Parameters of Nonlinear Scalar Dynamics from Video: Invariants, Calibration, and Identifiability
Abstract
Physical parameter estimation from video aims to recover the parameters of a known family of governing dynamical equations from pixel observations. Existing identifiability theory for this setting has focused on linear time-invariant (LTI) second-order systems, leaving open what can be identified for nonlinear scalar dynamics. We develop an identifiability theory for nonlinear scalar second-order ODEs, organized by how their velocity dependence interacts with changes of the learned state coordinate. Under a shared non-collapsed state map and explicit same-state velocity-coverage conditions, we show that parameter identifiability depends on the ODE family: some parameters are uniquely identifiable, while in other families only invariant parameter combinations are identifiable or external physical calibration is required. For laws that are at most linear in velocity, compatibility forces affine coordinate alignment, yielding explicit parameter relations, invariants, and calibration conditions. This affine conclusion extends to broader finite velocity-feature families when coordinate curvature can be separated from the declared velocity dependence. For families admitting a squared-velocity term, nonlinear coordinate ambiguity can remain; a law-derived normalization instead enables affine comparison between canonical laws. Experiments on synthetic systems and real pendulum and free-fall videos support the predicted parameter relations, coverage effects, and calibration requirements.
cs.LG / 72 / 2609.38889
VERA: Verifiable Feasibility Representations with Counterfactual Credit for Constrained Multi-Agent Control
Abstract
Constrained multi-agent control requires more than predicting rewarding actions: an action can cease to be executable as contact windows, shared capacity, and deadlines change. We introduce VERA, a centralized-training, decentralized-execution framework that separates feasibility estimation from credit assignment. Each actor predicts a five-dimensional verifiable feasibility representation (VFR). After an action is proposed, exact action-conditioned margins available only during training supervise that representation, while a counterfactual group-relative advantage (CGRA) ranks candidate representation-action pairs. Execution uses one actor pass and no privileged state. In a dynamic space-air-ground integrated network (SAGIN), VERA obtains 55.33% +/- 3.60% success with 0.45% +/- 0.81% coverage violation, within 1.33 percentage points of a privileged-mask reference. With rewards matched over ten paired seeds, VERA improves success over the strongest baseline by 8.74 percentage points (p=0.023) and reduces violation by 52.19 percentage points (p=5.7e-8). A ten-seed 4-by-2 factorial attributes a 14.16-16.48 percentage-point gain to CGRA across handcrafted, learned, random, and latent representations; evaluation on seven unseen topologies preserves a 24.33-30.02 percentage-point advantage over multi-agent proximal policy optimization. From 10 to 40 users, success remains 50.1-53.8%, and VFR adds only 0.026 ms to a central processing unit (CPU) actor step. Cross-domain tests further identify the governing condition: counterfactual credit succeeds when candidate scores respect shared constraints and fails under incompatible reward geometries. These results establish action-conditioned feasibility as an auditable training interface and counterfactual credit as a geometry-dependent optimization mechanism.
cs.LG / 73 / 2609.38893
Learning Continuous Neural Representation of Stochastic Hybrid Systems
Abstract
A stochastic hybrid system (SHS) is governed by a stochastic differential equation (SDE) describing the continuous dynamics and a Markov reset kernel triggered on the guard surface. Its probability evolution can be described by a hybrid Fokker-Planck (HFP) equation with a partial differential term corresponding to the SDE and an integral term arising from the reset kernel. This work shows that such an SHS can be approximated by an SDE in a higher-dimensional latent space where the sample paths are continuous. The key to this result is to encode different branches of the reset kernel using auxiliary variables, transforming the resets into deterministic ones that enable topological gluing. By the embedding theorem, the glued manifold can then be embedded into a higher-dimensional Euclidean space. We show that the probability evolution on the embedded image no longer requires explicit reset terms in the HFP equation. Building on this theorem, we design a loss that matches the evolving state distributions, enabling a single latent SDE to recover the probability evolution of the SHS without mode labeling, trajectory segmentation, or event-based simulations.
cs.LG / 74 / 2609.38895
Unmerge: Efficient Machine Unlearning via Task Arithmetic
Abstract
Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlearning actually happens inside the network. We recast unlearning through the lens of task arithmetic: if finetuning produces a merged task vector $τ_m$ that combines learning on forget and retain sets, unlearning is the inverse operation that subtracts a learned forget component $τ_F$ to recover the retain task vector $τ_R$. The forget signal is concentrated: at every layer, forget activations lie in a subspace spanned by a handful of dominant directions, so we factorize $τ_F$ in a low-rank forget basis, which is faithful up to a small tail-eigenvalue residual and limits how far the correction can perturb retain. We then optimize three intuitive goals (match the merged vector inside the forget span, suppress leakage into the retain span, and bound the correction size) that provably bound forget leakage and retain damage in activation space. The resulting algorithm, Unmerge, is fast and powerful: on class-level unlearning with ResNet-50 on CIFAR-100 and Tiny ImageNet, it improves Tug-of-War by up to ~24% over a baseline of comparable runtime and by up to ~18% over stronger baselines that run ~5x slower, keeps membership-inference exposure at the level of retraining, and shrinks the feature-distribution gap to the retrained model, where relabeling methods leave forget features cleanly separable. Further studies show that Unmerge also applies to ViT-S/16 and scales to Llama-3.2-3B. The per-layer basis geometry that drives the algorithm also serves as a layerwise diagnostic for when and where unlearning becomes structurally hard.
cs.LG / 75 / 2609.38898
K2P: Label-Free Knowledge to Prompt Distillation
Abstract
Knowledge distillation can transfer reasoning from stronger teachers to frozen students through reusable prompts, but avoiding weight updates does not eliminate supervision. Without ground-truth answers, teacher solutions are unverified, and agreement with the teacher can reward shared mistakes. We introduce Knowledge-to-Prompt (K2P) for label-free knowledge distillation to prompts. K2P synthesizes reusable instructions from teacher solutions, refines them using paired teacher and student responses, and guides search and selection with answer agreement. It retains candidates that adaptive search may undervalue and selects on reserved questions. Deployment uses only the frozen student and selected prompt. Our theory separates generation and selection gaps and gives conditions under which agreement-guided construction yields accuracy guarantees despite imperfect teacher references. Across reasoning tasks and students, K2P outperforms label-free alternatives overall and remains competitive with supervised prompt optimization. Ablations and archive diagnostics assess the contributions of teacher solutions and refinement, while revealing the limits of agreement-guided selection.
cs.LG / 76 / 2609.38901
Certified Approximation for Interpretable Representer Landmarks
Abstract
Representer explanations rank the training landmarks that most influence a self-supervised representation. At scale, this ranking rests on up to four stacked approximations of the empirical neural tangent kernel (eNTK). These are random output heads, a parameter sketch, landmark sampling and a coefficient fit. Existing analyses bound each approximation separately, but none certifies the top-$K$ set against their combined error. We introduce CAIRN (Certified Approximation for Interpretable Representer laNdmarks), a framework that carries this error through to the ranking. We derive the exact variance of the sketched multi-head eNTK, which matches measurement within $4\%$ where Johnson-Lindenstrauss bounds err by up to $2.5\times$. This yields a high-probability top-$K$ certificate for a fixed coefficient fit, alongside exact residual-trace certificates for discarded spectral mass. An exact product-variance identity separates kernel error from fit variability and identifies when a larger kernel budget can still sharpen a ranking. Stochastic Lanczos Quadrature (SLQ) estimates the effective dimension within $0.72\%$ and guides the landmark budget without dense eigendecomposition. We show that residual mass does not control class coverage, and residual-greedy selection cuts the worst coverage excess of $k$-means++ from $8.5\times$ to $1.55\times$ ($4\times$ on the sketched eNTK). Cross-view initializers outperform principal-component initialization in five (AUI) to all six (CSI) settings. Against the KREPES Gauss-Newton solver, CAIRN converges $2.5$ to $11.3\times$ faster, trails by at most $0.31$ points and gains up to $3.14$ points on MNIST. Together, these results make the reliability of representer explanations measurable and show where approximation budgets are best spent.
cs.LG / 77 / 2609.38902
Amortized Data Borrowing with Exchangeability-Aware Neural Posterior Estimation
Abstract
Augmenting small concurrent studies with external or historical cohorts is attractive in drug development, where enrollment is slow, follow-up is expensive, and closely related trial or real-world data are often already available. Bayesian dynamic borrowing (BDB) provides a principled framework for adaptively controlling the influence of external data, but classical implementations often depend on hand-specified priors and MCMC-based inference, which can be computationally expensive and not generalizable. In this work, we study amortized neural posterior estimation (NPE) as a flexible alternative. A single network is pretrained on simulated current/external dataset pairs spanning covariate shift, outcome drift, and joint non-exchangeability, and then returns an approximate posterior for a scalar current-study target in a single forward pass. Through simulation studies, we find that NPE is most useful under outcome drift and joint mismatch: in the harder outcome-drift regimes, it gives up to about five-fold lower absolute bias than the best classical baseline and keeps Type I error close to nominal. After pretraining, posterior summaries are obtained in about 8 ms per dataset, roughly $10^3\times$ faster than MCMC-based borrowing baselines in our timing experiment. We further analyze Alzheimer's Disease Neuroimaging Initiative (ADNI) data and show that, when mild cognitive impairment outcomes differ across cohorts, the NPE formulation recovers the later-cohort risk level in this example without claiming greater precision. Code is available at https://github.com/ChinHungScott/NPE-for-Bayesian-Dynamic-Borrowing-MLHC-.
cs.LG / 78 / 2609.38903
DAMPER: Return-Prioritized Gradient Control for Smooth Policies
Abstract
Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible when their gradients are small relative to the native actor gradient. Moreover, existing methods often combine multiple auxiliary losses, complicating loss balancing without necessarily improving the return-smoothness trade-off. We introduce DAMPER (Direction-Aware Magnitude-Controlled Projection with Explicit Return Priority), which combines the native actor gradient with a temporal-consistency gradient through conflict-conditioned projection and adaptive magnitude control. It removes the auxiliary component opposing the actor gradient and scales the retained temporal direction relative to the actor gradient norm, preserving positive alignment with the native actor gradient. Experiments with TD3 and SAC on six continuous-control tasks show reduced action oscillation relative to the native agents in all 12 task-backbone pairs and the best oscillation score among the compared methods in eight, with task-dependent return trade-offs.
cs.LG / 79 / 2609.38911
Generalized Residual Closure: General Learning Dynamics for Stability-Plasticity Compatibility
Abstract
Learning must acquire new capabilities while preserving both prior responsibilities and the capacity to learn again. We introduce Generalized Residual Closure (GRC), a framework for learning as recursive closure of future-relevant discrepancies: closing a residual establishes the conditions for subsequent prediction, interaction, and learning. Under a complete representation-relation description at a fixed learner-world boundary, persistent internal learning has two primitive modes: Transformation within a representation and revision of the Representation itself. We establish a local tangent decomposition under regularity assumptions and a criterion for when representation revision is necessary. In an affine model, we derive a necessary-and-sufficient condition for stability-plasticity compatibility and the unique solution of a constrained quadratic update problem, which preserves registered old responsibilities while reducing residuals with an effective safe response. We prove that reconstructive semantic protection weakly enlarges the safe-response operator relative to preserving an exact historical realization. Dynamic sufficiency and future-closure viability extend representation adequacy from current prediction to lawful future updating and continued learning. Growth Learning expands the lawful closure domain or lowers optimal closure cost without regression of the registered capability-cost frontier; a conditional commit rule maintains this order. Restricted-sector recoveries and a conditional representation theorem connect the framework to optimization, machine learning, and control. Together, these results organize adaptation, representation revision, and reusable capability within a common account of continued learning.
cs.LG / 80 / 2609.38918
Flow Matching under Noisy Latent Structure: Beyond Exact Low-Dimensional Support
Abstract
Flow Matching (FM) learns a velocity field whose ODE transports a simple source distribution to a target law. Existing finite-sample theory largely treats ambient-space regularity or data supported exactly on low-dimensional sets. We study linear FM under a noisy latent-generator model, where a low-dimensional Hölder map is perturbed by nondegenerate ambient Gaussian noise, so the target law is full-dimensional despite its latent structure. We construct a spatially regular ReLU velocity class and establish non-asymptotic high-probability approximation and estimation bounds whose leading sample-size exponent is governed by the latent dimension rather than the ambient dimension, with ambient and noise dependence kept explicit. Fixed positive target noise keeps the interpolation nondegenerate over the full time interval. The same spatial regularity propagates the learned velocity error through the transport ODE, yielding a corresponding Wasserstein convergence guarantee. These results show that exact low-dimensional support is not necessary for Flow Matching to retain latent-dimensional statistical behavior.
cs.LG / 81 / 2609.38927
World-as-Graph: Relational World Modeling Through Latent Space Graphs
Abstract
World models aim to learn representations of real-world environments and predict their future evolution. Recent object-centric world models have made expressive progress by representing visual scenes as sets of object-level latent states, but object-object relations are often captured only implicitly, which limits explicit relational and temporal structure modeling and object-centric dynamic memory modeling. To address such challenges, we propose World-As-Graph (WAG), a graph-based object-centric world model that introduces relational inductive bias into JEPA-style predictive representation learning. The proposed WAG contains two main modules: (1) Relation-aware structure induction, which constructs time-varying latent graphs from object-centric slots and designs relation-aware object masking policies to guide relational object representation learning in latent space; (2) Object-centric memory transition, which maintains and updates object-level dynamic states by combining relational information from neighboring objects with historical memory, enabling effective autoregressive future prediction. Extensive experiments on both visual reasoning and robotic manipulation tasks could demonstrate the superior performance of our proposed WAG.
cs.LG / 82 / 2609.38931
Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback
Abstract
Self-consistency samples many reasoning trajectories and aggregates their final answers, treating the LLM as a black box that returns one answer per trajectory. Yet the final answer of each trajectory is sampled from a softmax vector that is available from the model's log-probabilities. We refer to this as the grey-box setting in which each trajectory reveals this answer distribution rather than a single draw from it. We formulate efficient inference in this setting as sequential mode identification with distribution-valued observations: sample trajectories one at a time and stop as soon as the LLM's modal answer is identified at a prescribed confidence level. We characterize the asymptotic stopping rate of mode identification with distribution-valued observations exactly and show that it is never worse than the black-box rate. We then propose the ASC-D algorithm, a betting stopping rule that attains this asymptotic stopping rate. On MMLU-Redux, ASC-D uses $46.4$--$95.6\%$ fewer trajectories than answer-only adaptive self-consistency baselines and achieves the highest fixed-budget correct-certification rate across three open-source models.
cs.LG / 83 / 2609.38932
EFormer: Temporally Aligned Local Correction for Continuous sEMG-Based Hand Pose Tracking
Abstract
Surface electromyography (sEMG) provides a wearable, camera-free signal for continuous hand-motion inference. Mapping muscle activity to joint kinematics remains challenging because the recorded waveforms are indirect measurements, their relationship with motion changes over time, and individual anatomy and sensor placement alter the signal distribution. This paper presents EFormer, a residual feature-correction network built on a frozen tracking backbone. EFormer combines a high-rate event branch, temporally aligned local cross-attention, two causal rotary position embedding (RoPE) temporal layers, and a bounded, dynamically gated residual. EFormer receives 16-channel sEMG sampled at 2 kHz and fuses a 64-channel tracking representation at 25 Hz with a 128-channel event representation at 200 Hz. Cross-attention uses a nominal delay of 100 ms, a 300 ms history parameter, and a 50 ms tolerance; its causal mask restricts each query to events occurring 50-400 ms earlier. The correction scale is 0.15. The evaluated continuation-training configuration contains 585,376 trainable parameters and 5,974,508 frozen parameters. On the test set, EFormer achieves an MAE of 0.1546634 rad, an RMSE of 0.24063 rad, and an R^2 of 0.74801, compared with 0.1745326 rad, 0.2715448 rad, and 0.6791103 for the official tracking baseline. EFormer reduces MAE by 11.38% relative to the baseline. The results show that temporally aligned event-feature correction can reduce continuous hand-pose tracking error.
cs.LG / 84 / 2609.38936
Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets
Abstract
When a classifier is recalibrated from only a few thousand held-out examples, the capacity of the calibration map becomes a statistical design choice rather than a purely architectural one: a scalar map can underfit structured residual miscalibration, while a highly adaptive map can be hard to estimate reliably from so small a split. We disentangle the calibration objective from adaptive capacity and propose signal-routed temperature scaling (SRTS-BCE), a 10-parameter, argmax-preserving calibrator that cross-fits a correctness-risk score over six logit statistics and fits one top-label-BCE temperature per $K=3$ risk groups, recovering TvA-TS as its $K=1$ limit. On fine-tuned CIFAR-100 / ViT-B/16, SRTS-BCE reduces $\mathrm{ECE}_{15}$ from 1.65 (scalar TvA-TS) to 0.96, matching the higher-capacity SMART+BCE head (0.95) at the full calibration budget. The two regimes separate as the budget shrinks: at $n=250$ SRTS-BCE beats SMART+BCE on all three CIFAR-100 backbones (the seed-to-draw hierarchical interval excludes zero), whereas the flagship comparison against the scalar remains directional. A protocol-frozen Tiny-ImageNet follow-up reproduces the small-budget separation and exhibits a budget-dependent ranking reversal on Swin-T; matched routing and map controls show that the effect is tied neither to the learned router nor to discrete grouping. Together the results identify post-hoc calibrator capacity as a finite-sample design choice whose preferred level shifts with the amount of available calibration data.
cs.LG / 85 / 2609.38938
Robust Risk-Sensitive Reinforcement Learning from Corrupted Human Feedback
Abstract
Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under adversarial preference-label flips. We consider additive linear rewards and a fixed-reference protocol with one comparison per episode and at most $C$ flipped labels over $K$ episodes. We propose weighted streamed-preference CVaR RLHF (WSP-CVaR-RLHF), which combines uncertainty-weighted reward estimation with optimistic augmented-state CVaR planning. For known transitions and normalized rewards, we establish the regret bound $\widetilde{O}\left(\frac{d}κ\sqrt{\frac{K}α}+\frac{dC}{κα}\right)$ up to lower-order terms, where $d$ is the reward-feature dimension, $α$ is the CVaR level, and $κ$ characterizes the preference link. The bound separates the clean statistical cost from the penalty caused by corrupted feedback. We further extend the analysis to unknown tabular transitions, where the trajectory distribution entering the CVaR objective must be learned together with the reward. We address the resulting coupled uncertainty using rectangular transition confidence sets, joint optimistic planning, and a history-level CVaR simulation argument. Experiments under four adversarial attacks demonstrate that WSP-CVaR-RLHF consistently reduces cumulative regret relative to its unweighted robust counterpart while preserving confidence-set coverage.
cs.LG / 86 / 2609.38977
Scale-Split Neural Operator for Memory- and Data-Efficient 3D Turbulence Prediction
Abstract
Neural surrogates have emerged as fast alternatives to the numerical simulation of three-dimensional turbulence. However, training them at high resolution remains challenging, since the memory of full-field models grows with the resolution. In addition, full-resolution training data are expensive to simulate and store, and therefore scarce. We introduce ScaleSplit-NO (Scale-Split Neural Operator), which exploits the scale structure of turbulence with two neural operators: a Parent predicts the global coarse field at the next time step, and a Child predicts full-resolution local patches conditioned on this prediction. Neither model operates on the full-resolution field. The Child is pretrained alone and then attached to the Parent's coarse prediction through zero-initialized connections. On two complex high-resolution turbulence benchmarks, ScaleSplit-NO surpasses all competing baselines in both prediction accuracy and data efficiency. On the higher-resolution dataset JHTDB256 ($256^3$), its normalized mean squared error (NMSE) is 53% lower than that of the strongest baseline, and its training memory is 79% lower than that of the most memory-efficient baseline. We further demonstrate its effectiveness for urban wind prediction in a real district of Montreal on a $500\times150\times500$ grid, reducing one-step NMSE by 65.8% relative to the baseline. Moreover, swapping in a Parent trained on additional coarse fields improves prediction without retraining the Child, providing further accuracy gains at a small storage cost.
cs.LG / 87 / 2609.38980
SCORE-LM: State-Space Radar Representations with Language Models for Fault Diagnosis
Abstract
Radar hardware faults threaten automated perception, motivating accurate, compact diagnosis and understandable maintenance guidance. We introduce SCORE-LM, which couples a small scatterer-conditioned operator-response encoder (SCORE) to an adapted local language model. SCORE combines self-referenced complex trajectories, physical descriptors, and a selective state-space branch, with source-only self-supervision and directional fault inference. On eight capture-excluded Rad-R fault recordings, it achieves state-of-the-art performance within the evaluated nine-model comparison: 88.39% mean capture recall and 88.20% four-fault macro-F1 at ten frames. Its 39,520 radar inference coefficients are 119.7 times fewer than RadrNet-DS-CI's, while recall is 15.56 percentage points higher than this strongest competitor. In a separate low-label protocol, SCORE reaches 71.58% recall with one labeled source window per class. A nonlinear projector converts four frozen fault similarities into five soft tokens, linking compact diagnosis to class-conditioned maintenance guidance. On 75 development questions covering 24 radar windows, language adaptation raises correct-fault answers from 45 to 62 (60.0% to 82.7%) relative to removing the co-trained adapters, while retaining the same projector. SCORE-LM thus combines a compact radar specialist with a language interface for communicating fault-specific inspection guidance.
cs.LG / 88 / 2609.38987
Smaller Models, Better Rejects: Preference Distillation Scaling
Abstract
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
cs.LG / 89 / 2609.38991
Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework
Abstract
A key bottleneck in adversarial transfer is a trajectory-level geometric disconnect: ambient gradients often drift away from the intrinsic data manifold, causing surrogate-specific overfitting. To rectify this, we propose Manifold Anchored Bilevel Transfer (MABT), a unified framework that anchors adversarial trajectories to the shared semantic subspace. MABT introduces a relaxed manifold-anchoring operator as a semantic rectifier to suppress off-manifold noise. With this constraint, we cast transfer attack generation as a distributional bilevel optimization problem that learns a geometry-aligned initialization by minimizing expected transfer risk under a surrogate uncertainty distribution. We further develop a Hessian-free solver with linear-time complexity to handle the resulting hierarchy. Experiments demonstrate improved transferability for 10 baseline attackers across 28 attack configurations, diverse victim architectures, and defense mechanisms.
cs.LG / 90 / 2609.38998
Not all solutions are created equal: An analytical dissociation of functional and representational similarity in deep linear neural networks
Abstract
A foundational principle of connectionism is that perception, action, and cognition emerge from parallel computations among simple, interconnected units that generate and rely on neural representations. Accordingly, researchers employ multivariate pattern analysis to decode and compare the neural codes of artificial and biological networks, aiming to uncover their functions. However, there is limited analytical understanding of how a network's representation and function relate, despite this being essential to any quantitative notion of underlying function or functional similarity. We address this question using analysable two-layer linear networks and numerical simulations in non-linear networks. We find that function and representation are dissociated, allowing representational similarity without functional similarity and vice versa. Further, we show that neither robustness to input noise nor the level of generalization error constrain representations to the task. In contrast, networks robust to parameter noise have limited representational flexibility and must employ task-specific representations. Our findings suggest that representational alignment reflects computational advantages beyond functional alignment alone, with significant implications for interpreting and comparing the representations of connectionist systems.
cs.LG / 91 / 2609.39025
A Rank Graduation metric for Algorithmic fairness
Abstract
Fairness assessment in algorithmic decisions that affect individuals, such as credit scoring, often relies on parity measures calculated at the aggregate group level. Such measures may not reveal which individuals experience unfairness or which explanatory factors contribute to it. In this paper, we propose a rank-based framework that evaluates fairness through the distribution of model prediction errors, thereby linking fairness assessment with predictive accuracy and explainability. The framework combines Rank Graduation Fairness (RGF), its integrated measure AURGF, a centered Cramer--von Mises permutation test, and a feature removal procedure for fairness explainability. We evaluate the methodology using logistic regression, random forest, gradient boosting, and a multilayer perceptron. The simulation study shows that protected-group imbalance can reverse descriptive fairness comparisons, whereas the proposed inferential procedure correctly distinguishes fair from unfair mechanisms. Its application to HMDA mortgage data produces model rankings that differ from those obtained with classical fairness criteria. Tree-based models, rather than logistic regression, provide the strongest combination of predictive accuracy and rank-based fairness, while the fairness null hypothesis is rejected for all four models. The persistence of disparity across statistical, bagging, boosting, and neural network specifications, together with the feature removal results, indicates that the observed unfairness is not specific to a single algorithm or predictor, but is associated with group differences embedded in the characteristics of the lending data. These findings support a broader approach to trustworthy artificial intelligence that combines predictive accuracy, fairness measurement, statistical inference, and explainability.
cs.LG / 92 / 2609.39034
Switching Linear Attention
Abstract
Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.
cs.LG / 93 / 2609.39035
Cycle-Aware Autoencoder with Cross-SignalConsistency for Railway Door Anomaly Detection
Abstract
Passenger access doors are safety-critical subsystems in railway vehicles, yet detecting abnormal door behavior in real operation is challenging because faults are rare, diverse, and often unlabeled. This paper addresses railway door condition monitoring as a cycle-level unsupervised anomaly detection problem, where each complete opening-dwell-closing cycle is treated as a single monitoring unit. We propose the Temporal Cycle-Aware Attention Autoencoder with Cross-Signal Consistency (TCAA-CS), trained exclusively on nominal cycles. It combines a dual-stream encoder that processes continuous physical measurements (position, current, voltage) and binary logical states (door-closed, door-locked) through separate 1D-CNN branches, an LSTM encoder with temporal attention pooling, and a triple hybrid anomaly score fusing reconstruction error, latent-space deviation, and phase-aware cross-signal consistency. The consistency term helps identify cases where individual signals appear plausible but their inter-signal relationships become physically or logically inconsistent. On real industrial data from a passenger train in commercial service, TCAA-CS achieves 93.8% recall, 97.3% precision, and a 0.5% false-alarm rate, outperforming representative unsupervised baselines. System-level evaluation on an NVIDIA Jetson AGX Xavier supports the feasibility of real-time onboard deployment.
cs.LG / 94 / 2609.39037
Hard-Gate Candidacy in a Deployed Validator Suite
Abstract
Before a validator can be promoted to a hard gate on a deployment pipeline, it has to be shown that its firing separates outputs that reach users in working order from those that do not. We run that screen on 13 validators in a deployed generative agent, against 550 runtime and 350 static builds labelled by downstream outcome, and report each check's marginal separation $J=\mathrm{TPR}-\mathrm{FPR}$ with Newcombe intervals and Fisher exact tests. Two checks survive correction for multiple comparisons, two more are nominal only, and the remaining nine are not distinguishable from zero, three of them because they never fired on any sampled build. Execution itself is not random with respect to the property being gated, and this replicates: across four runs covering 1,867 builds and ten distinct runtime checks, probes were skipped on 144 of 895 broken builds and 1 of 972 acceptable builds (per-run rates 15.6% to 16.6% against at most 0.3%), every skip carrying the same unsafe-to-probe reason. Because a skipped check is recorded as a pass, this imposes a ceiling that no check quality can lift: a check that needs a live artifact cannot operationally detect more than about 84% of broken builds in this harness. For the one check with construct-specific labels, a detector built for blank output fires on 0 of 90 human-labelled blank builds (95% upper bound on sensitivity 3.3%), and the global frame statistic it approximates separates the classes only weakly (AUC 0.59), so the gap is not a threshold that needs tuning. The same gap appears one layer up: on a census of tens of thousands of judge-scored builds, 32.5% of rejections carry no recorded issue at all. We argue that evaluation records must distinguish a check that ran and passed from one that did not run, must carry the evidence for a rejection, and that an inventory of checks is not evidence about a gate.
cs.LG / 95 / 2609.39048
Structure-aware Reinforcement Learning for Protein Directed Evolution
Abstract
Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.
cs.LG / 96 / 2609.39055
The Missing Coefficients: Bayesian Pairwise Merging for Model Personalization
Abstract
How can we personalize a shared expert library from a user's pairwise choices? Prior work can realize different reward trade-offs by merging reward-specialized experts, given a vector of trade-off weights. In practice, users can more naturally choose between outputs than specify numerical weights. The challenge is therefore to turn these choices into the coefficients required for merging, while accounting for ambiguity when feedback is limited. Our key idea is to treat the unknown reward weights as latent variables: infer a posterior over them from pairwise choices and reward-score differences, and use its mean directly as the merge coefficients. We instantiate this idea as Bayesian Pairwise Merging (BPM), whose posterior also characterizes which reward trade-offs remain plausible given the feedback. We evaluate BPM on radiology summarization, image captioning, and story generation, spanning text-to-text and image-to-text generation. With 100 feedback per simulated persona, BPM achieves macro decided win rates of 91.7%, 77.1%, and 64.3% against uniform merge. For six pairs of simulated personas, each prefers the model fitted to its own feedback, a pattern also observed in a human proof-of-concept. In simulations under BPM's model and prior, its nominal 90% intervals for temperature-scaled reward weights achieve task-averaged marginal coverage of 88.9% and 89.2% with only 10 and 25 comparisons, respectively. BPM thus enables personalization from pairwise feedback without per-user policy training, while characterizing the coefficient ambiguity left by limited feedback.
cs.LG / 97 / 2609.39067
Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple Checkpointing
Abstract
Elastic Compute Cloud (EC2) Spot is 60% to 90% cheaper than On-Demand but can be reclaimed on just a 2-minute notice; for expensive multi-node training this loss can be severe, with one reclaim costing hours of synchronous progress. We build Argus, a Kubernetes operator, and ask empirically, on a CIFAR-10 testbed, when predicting interruptions beats simple checkpointing. Argus on real EKS survives a real Spot drain with a graceful SIGTERM checkpoint, resuming from epoch 8 and losing only the in-progress epoch. Alongside, we further find that in an 80-trial benchmark, the reactive-on-notice degrades toward no protection once interruption outpaces the fixed 2-minute notice, and predictive wasted compute is driven to zero, but with an oversized fixed lead it over-migrates so severely that at the fastest rate only one of five runs completes, while periodic is a strong ML-free baseline. A lead-time sweep turns the lead prediction into a guideline where a small lead suffices for zero waste, but excess lead is wasteful. The predictor built is advisory (a proxy label); real interruption labels and large-model-scale validation are future work.
cs.LG / 98 / 2609.39068
SparseEngine: Sparse-First Inference Engine
Abstract
Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specific layouts or workflows. We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure. SparseEngine supports 15 methods across four categories and enables cross-request state management through Chain Cache, which resumes KV-eviction methods from retained history, and controllable Prefix-Cache Pruning, which removes KV from selected history regions while preserving logical-prefix matching. While maintaining method quality, SparseEngine delivers over 10x higher throughput with KV eviction, over 2.5x faster decoding at matched concurrency than vLLM, and over 2x end-to-end speedup on agent benchmarks. The code is available at https://github.com/CURRENTF/SparseEngine.
cs.LG / 99 / 2609.39074
HO-FL: Hybrid-Order Federated Learning for Heterogeneous Edge Devices
Abstract
Federated learning (FL) on memory-constrained edge devices faces a dilemma: first-order (FO) optimization (i.e., backpropagation) demands substantial memory, whereas zeroth-order (ZO) optimization suffers from severe convergence slowdown. To resolve this dilemma, we introduce HO-FL, a hybrid-order FL framework that trains a model's bottom segment with ZO optimization and its top segment with FO optimization. Each device can flexibly select its order boundary according to its memory budget while participating in the training of the same global model. Moreover, our convergence analysis reveals a new, fundamental trade-off: clients with larger FO-trained segments can provide more accurate updates, but favoring them can underrepresent other clients' data. We connect this trade-off to the bias and variance of actual multi-step local updates, yielding a sampling optimization problem and a practical dimension-aware approximation with direct model averaging. Experiments on language tasks examine task performance, client memory, and sampling under data heterogeneity. The results show that hybrid-order local training can retain much of the full-FO performance with substantially lower client memory requirements. Our code is available at https://github.com/HKU-WILL-Lab/HO-FL.
cs.LG / 100 / 2609.39078
Parameter symmetries determine representational geometry in overparameterized nonlinear networks
Abstract
Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and artificial systems. Such inferences presume a meaningful link between representational geometry and the computation being performed. For artificial neural networks, however, the extent to which function constrains representation remains unclear. One key obstacle is that these networks admit parameter symmetries: changes in parameterization that preserve function exactly while reshaping representational geometry. Here, we show that a broad class of parameter symmetries acts on representations through just three primitive feature transformations: addition, duplication, and scaling. This feature-level characterization yields a closed-form decomposition of representational geometry into essential and auxiliary components, which makes precise how degeneracy in representational geometry can grow with overparameterization even when function is held fixed. Finally, we show that implementation-level selection rules can resolve this degeneracy, yielding identifiable geometries in which features are weighted according to their contributions to the network's function. Together, our results delineate when representations can support inferences about computation, and when they cannot.
cs.LG / 101 / 2609.39082
Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence
Abstract
As new evidence arrives, a sequence model must update what it remembers and how memory influences predictions. While Transformers incur computation and cache costs scaling with context length, fixed-state recurrent models offer constant-memory inference. However, linear and spectral recurrences traditionally rely on static transitions, failing to dynamically revise how stored representations decay or rotate. While recent selective architectures introduce input-dependent transitions, they assign independent controls to every memory mode, coupling control cost to state capacity. We show that high-dimensional spectral memory does not require high-dimensional control, and introduce Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence (SPARC). SPARC employs just two input-dependent scalar signals to coordinate memory retention and phase rotation across heterogeneous complex modes, while preserving mode-specific baseline timescales and frequencies. Its diagonal affine recurrence supports parallel associative scans for sequence-level BPTT as well as exact structured Real-Time Recurrent Learning (RTRL) for online credit assignment. Across partially observable continuous control, POPGym, and sequence classification, SPARC achieves a 9.09% relative return improvement on Walker-P and a 1.36% relative accuracy gain on FordA over second-best methods. On an NVIDIA Blackwell GPU, our implementation reduces recurrent-mixer training latency by 18.2%-34.2% in fixed-token workloads and accelerates scans by 3.1x-4.7x over an optimized RG-LRU baseline. These results show that two shared control signals can efficiently govern adaptive spectral memory across online and full-sequence settings. Code is available at https://github.com/Botwwt/sparc.
cs.LG / 102 / 2609.39093
Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation
Abstract
We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the number of interactions $T$. We propose, to the best of our knowledge, the first computationally efficient algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{T})$ regret and cumulative constraint violation with high probability in the tabular setting. The $\sqrt{T}$ dependence is optimal up to logarithmic factors. Our approach incorporates cumulative constraint violation into the state and defines a reshaped reward through differences of a Huber potential. The added state determines the penalty on further violations while the reward function remains fixed on the augmented state space. Since the added state has known deterministic dynamics, only the original transition kernel needs to be estimated. The bounded slope of the Huber potential keeps the per-step reward bounded, and the potential differences telescope to relate the reshaped return to the original cumulative reward and the terminal potential. These properties allow us to apply finite-horizon approximation and optimistic value iteration with clipping, as used in unconstrained average-reward MDPs, without worsening the regret rate in $T$.
cs.LG / 103 / 2609.39099
A Generalisation Signal Need Not Be a Model-Selection Signal
Abstract
Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it. When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself. We test this idea using a novel, forward-only proxy motivated by the norm of the Hessian, alongside common Hessian measures, across molecular property, protein fitness, and drug-response tasks. Contrary to our hypothesis, geometry does not become more useful as validation Spearman correlation deteriorates: augmenting validation helps some shifts but significantly harms others. More surprisingly, the proxy still correlates with generalisation gap on most tasks even when Hessian trace and top-eigenvalue relationships are weak or reversed, yet this signal does not reliably identify the deployment-best model. A curvature bound need not preserve cross-model rankings, and low geometric scores can even favour collapsed predictors. Thus, a generalisation signal need not be a model-selection signal.
cs.LG / 104 / 2609.39109
T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning
Abstract
Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling gives thalamic context-dependent routing a concrete computational form: learn which earlier contributions a receiving layer uses, and with what influence. Correctness rewards train the interface while preserving backbone parameters and layer order. On an 8.95B-parameter backbone, T-Router allocates 41.73M parameters (0.466% of the backbone) and achieves 83.64 +/- 1.16 MathAvg after GSM8K RL, compared with 73.79 +/- 1.83 for full-parameter GRPO across three evaluation rounds. At a comparable parameter budget and with matched retries, it exceeds LoRA's 77.28 +/- 1.95 MathAvg, improving all three task families and raising mean AIME accuracy from 48.33 to 60.56. Capacity-controlled comparisons favor addressable block changes and recurrent context; separate search training extends the interface to tool-mediated reasoning. These results establish controlled computation reuse as an effective route to parameter-efficient reasoning reinforcement learning.
cs.LG / 105 / 2609.39123
Prequential E-Values for Selected-GP Near-Optimality Certificates
Abstract
When optimizing an expensive black-box function sequentially, as in hyperparameter optimization, we may want to stop once the best evaluated value is certified within $\varepsilon$ of the global optimum. Such a certificate needs two ingredients: a lower confidence bound for the selected value and an upper confidence envelope over the domain, typically supplied by a Gaussian process (GP). GP-UCB-style stopping rules are valid when the kernel and constants defining this envelope are fixed before the run, but the practical temptation is to tune the envelope from the same adaptive evaluations and then certify as if it had been fixed. We use prequential e-values to make this selection auditable: starting from a predeclared set of fully specified GP/RKHS envelopes, each candidate is tested by its own one-step-ahead e-process, contradicted candidates are deleted, and certification uses the largest upper bound among the survivors. With a valid selected-point lower bound and one declared candidate having valid latent coverage and noise calibration, the rule is anytime-valid. On a 512-seed noisy RBF stress sweep, it roughly halves false-certification risk at comparable power versus fit-then-certify. Relative to random fixed GP precommitment on smooth $d=3,4$ objectives, each additional false certificate is accompanied by 3.0 and 13.5 additional correct certificates, respectively.
cs.LG / 106 / 2609.39144
Sharp Stationary Gaussian Approximation for Constant-Stepsize SGD
Abstract
We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is $O(\sqrtα)$-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.
cs.LG / 107 / 2609.39164
QuanVI: Score-based Variational Inference via Quantum Maximally Mixed States
Abstract
Score-based variational inference (VI) provides an alternative to Kullback--Leibler (KL)-based VI by minimizing the Fisher divergence between the variational distribution and the target. A prior score-VI approach formulates this optimization as an eigenvalue problem, with the variational distribution constructed from low-energy eigenstates. However, this eigenvalue-based formulation faces two high-dimensional obstacles: an intractably large parameter count due to exponential scaling and non-uniqueness of individual eigenvectors in degenerate or nearly degenerate low-energy subspaces. We propose QuanVI, a scalable quantum-inspired algorithm that combines a mixed-state density-operator formulation with a quantum tensor network (QTN) parameterization using the matrix product operator (MPO) structure. In degenerate low-energy subspaces, the density-operator formulation represents the subspace by its maximally mixed state rather than relying on a non-unique individual eigenvector, while the QTN parameterization compresses the density operator to avoid exponential parameter growth. Experiments and ablations show that QuanVI agrees with exact solutions in low dimensions and scales to high-dimensional synthetic and Bayesian posterior-approximation benchmarks, including challenging non-Gaussian targets.
cs.LG / 108 / 2609.39177
Whitening Improves Robustness to Spurious Correlations in Linear Probes
Abstract
Deep neural networks tend to rely on simple features that may be spurious and thus fail to generalize. We study this problem in the setting of linear probes, where a (generalized) linear model is fitted on the representations of a (pretrained) model. We use the connection of these models to the max-margin classifier, and show they favor directions associated with large eigenvalues of the covariance matrix. Whitening removes this preference by equalizing the eigenvalues of the covariance matrix. This observation motivates whitening as a preprocessing step that can reduce reliance on spurious correlations without requiring prior knowledge of their presence or labeled data. We examine the effect of whitening on a synthetic data-generating process and standard spurious correlation benchmarks, and find that it improves robustness. We also find that whitening can improve robustness when added to existing approaches.
cs.LG / 109 / 2609.39185
Low-Discrepancy Dither for Quantized Recurrent State Caches
Abstract
Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and hybrid models, storage formats, and long decoding horizons, at no extra cost. Round-to-nearest behaves differently: because it discards small updates, its error keeps growing, so it can look best in short evaluations yet falls far behind over long generations. A discrepancy analysis explains this ordering, and we document implementation pitfalls that silently remove the benefit.
cs.LG / 110 / 2609.39188
Attention Function as an Intrinsic Inductive Bias: How Models' Behavior Diverges in Novel Contexts
Abstract
Developmental psychology holds that certain priors are given to infants prior to experience rather than induced from data, and that the influence of such priors is suppressed under strong, well-constrained conditions but reasserts itself under weak ones. We ask whether an analogous principle holds for the Transformer: can the activation function given to attention heads serve as an intrinsic inductive bias? We propose Mixture of Function Attention (MoFA), a parameter-free modification to multi-head attention that fixes a ratio of softmax and sigmoid heads before training. Across five ratios, a 124M-parameter GPT-2 model, and five seeds, we find that this given ratio has little effect in-distribution -- differences between ratios are statistically negligible for moderate mixtures and remain small even at the extremes -- but its influence re-emerges sharply under zero-shot distribution shift across 15 out-of-distribution domains. Perplexity gaps between ratios widen by more than an order of magnitude on several domains, and the best-performing ratio tracks a single axis of domain structure, separating short, informal text (softmax-favoring) from technical, long-form text (sigmoid-favoring), that explains 78.3% of the variance in domain response. This reorganization is visible at the head level: sigmoid heads show an accelerating drop in attention entropy as their ratio increases, while softmax heads respond more modestly, yielding a consistent division of labor between the two head types. Our results suggest that activation choice functions as a given prior whose influence is masked in-distribution and re-emerges out-of-distribution.
cs.LG / 111 / 2609.39190
Dynamics to decision: A mathematical theory of Lyapunov spectra and decision boundaries in deep classifiers
Abstract
A deep classifier is defined not only by the decision it produces, but also by the sequence of transformations through which that decision is formed. Treating this evolution as a dynamical system across layers provides a natural framework for asking how decision geometry emerges through depth and how far back we can trace a boundary's dynamical signature. We model a feed-forward classifier as a finite, nonautonomous discrete dynamical system, with layers playing the role of discrete time steps. We study the Finite-Time Maximum Lyapunov Exponent (FTMLE) of the data samples' dynamical trajectory through depths of the classifier. The FTMLE measures the rate of convergence/divergence of nearby trajectories. We move the observation endpoint backward from probabilities to logits and then to hidden representations. For Gaussian classes, we prove that probability-level FTMLE carries a clear geometric signature of the decision boundary, with its dominant direction aligned with the boundary normal. Moving one step backward to the logits, we prove this relationship is no longer universal but depends critically on how the classifier is trained, particularly on the choice of loss function. Moving further backward to the hidden representation, the connection becomes more conditional: boundary-related FTMLE can persist, but only under identifiable structural conditions. We propose geometry-aware fine-tuning for restructuring the classifier's hidden FTMLE, and propose conditions for guaranteed concentration of high hidden FTMLE near the decision boundary. Through our numerical results, we show the generality and validity of our theoretical results. Understanding the evolution of data samples as traveling through the layers of classifier provides a principled foundation for identifying where boundary-relevant sensitivity emerges and for developing layer-aware regularization strategies.
cs.LG / 112 / 2609.39194
Importance-Aware Feature Sparsification for Wireless Split Learning
Abstract
Wireless split learning (SL) reduces on-device computation by offloading upper layers to a server, yet transmitting high-dimensional intermediate features at each iteration remains a major communication bottleneck. Existing methods select features at the client side using task-agnostic criteria such as magnitude, statistics, or clustering, which increases client-side processing and often degrades accuracy under non-independent and identically distributed (non-i.i.d.) client data. We propose importance-aware class-balanced sparsification (ICS), a lightweight approach in which the server ranks feature channels using Grad-CAM-based scores obtained from the true-class logit during backpropagation. The per-class scores are aggregated into a class-balanced, label-agnostic importance vector that mitigates head-class bias under label skew, and each client reuses this vector in the next round to retain the top-$N$ feature channels, incurring no additional client-side forward or backward passes. We further derive a non-asymptotic convergence bound that isolates the sparsification-induced error and characterizes how the sparsification ratio and mini-batch size jointly affect convergence under a fixed communication budget, and we analyze the communication and computational overhead of ICS against representative baselines. Beyond sequential CNN-based SL, we extend ICS to parallel split learning and to transformer-based models. Experiments show that ICS consistently outperforms the baselines, with larger gains under severe non-i.i.d. partitions.
cs.LG / 113 / 2609.39215
In a Streaming World, Should You Stand Still? A Comprehensive Benchmark of Anomaly Detection in Streams
Abstract
Time series anomaly detection (TSAD) is increasingly deployed in streaming settings, where data arrive sequentially and may exhibit non-stationarity. As a result, several works from the recent literature propose streaming anomaly detection methods that rely on incremental updates to adapt over time. However, most of these approaches originate from the streaming outlier detection literature and largely ignore core characteristics of time series anomalies. Moreover, their empirical evaluation is typically conducted on synthetic or small-scale benchmarks with limited diversity, making it unclear whether streaming methods are truly advantageous in realistic TSAD scenarios. In this work, we carry out the first large-scale experimental study comparing streaming and static TSAD methods under a unified streaming evaluation benchmark. We consider a realistic setting in which an initial batch of data is available for model training, followed by online evaluation of both detection accuracy and computational efficiency. In addition, we propose a distribution-drift dataset of real time series, called TSB- drift, to isolate scenarios where streaming updates are theoretically justified. Our results show that, contrary to common assumptions, static TSAD methods significantly outperform streaming approaches in most streaming settings. Such finding highlights a critical gap between the design of existing streaming methods and the requirements of modern TSAD, and calls for a rethinking of how streaming capabilities should be integrated into TSAD.
cs.LG / 114 / 2609.39232
What Streaming Anomaly Detection Finds (and Misses) in Industrial Time Series
Abstract
EDF relies on continuous monitoring of its power plants to detect anomalies as soon as they occur. Given the absence of a universally optimal streaming method in unsupervised settings, we compare streaming methods with state-of-the-art TSAD models deployed online on a real nuclear power plant dataset. This work also evaluates Automated Anomaly Detection in a streaming context. Results show higher consistency for online TSAD and strong robustness from ensembling strategies.
cs.LG / 115 / 2609.39242
A differentiability framework for zigzag persistent homology via linear interpolation
Abstract
Persistent homology can be differentiated and incorporated into learning pipelines, but no analogous framework exists for zigzag persistence, which is needed when the underlying topological structure evolves non-monotonically over time. We develop such a framework for sequences of simplicial complexes obtained by thresholding time-dependent filtering values on a fixed complex. By assigning persistence diagram endpoints the real-valued times at which linearly interpolated filtering values cross the threshold, we transfer the continuity of the filtering values to the diagram points. This yields smooth local lifts of the resulting persistence-diagram-valued map, from which we derive differentials almost everywhere under mild regularity conditions on the parametrization of the filtering values. We prove local Lipschitz continuity outside an explicit measure-zero exclusion set; standard stochastic subgradient convergence guarantees therefore do not apply directly. We argue that, even without such guarantees, this exclusion set is small enough in practice to allow effective optimization. We test this empirically in two experiments: sensor network coverage optimization and dynamic graph classification.
cs.LG / 116 / 2609.39243
Right Answer, Wrong Mechanism: Detecting Pernicious Divergence in Causal Interventions
Abstract
Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless and sometimes pernicious: it can recruit pathways the model never uses on natural inputs, so that an intervention produces the expected answer through the wrong mechanism. No method currently tells the two cases apart. We make this question testable by planting hidden pathways inside pretrained language models; the pathways are silent on every benchmark prompt by construction, so which interventions depend on them is known exactly. Across 72 configurations and 100,800 interventions on GPT-2 small, we find three things. (i) Nearest-neighbour and local-PCA distances at the intervention site, as used in prior work, score below chance (AUROC 0.35-0.47) at picking out interventions that give the right answer through a planted pathway. (ii) Hidden-Pathway Contribution (HPC), a label-free test that clamps downstream units to the regime of natural runs with the same output and measures how much of the decision disappears, flags pathway-dominated interventions with AUROC >= 0.99 when the pathway shows up as unit-level out-of-regime activity, but fails when every unit stays within its natural range, which we identify as the open problem. (iii) Optimised interventions actively seek hidden pathways: on a gender task, DAS routes 90-95% of its successes through planted pathways for three of four families, and a downstream on-manifold penalty cuts this share to under 5% at a cost of 6-11 points of success rate. In unmodified GPT-2, successful interventions show almost no unit-level out-of-regime reliance.
cs.LG / 117 / 2609.39247
Trust the Critic More
Abstract
Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the policy to update without observing a terminal reward. We make critic-based credit assignment reliable through three design choices. First, we introduce local readiness which uses critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem. Second, when available, we provide the critic with a reference solution from a previous successful rollout. Third, we assign credit over action chunks of 10k tokens rather than individual tokens, giving the critic a more meaningful portion of the trajectory to evaluate. We train Qwen3-4B on FineProofs-RL using AC2 and evaluate on IMO-ProofBench. AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs. This gain comes from two sources, (1) AC2 requires 25% fewer training steps to reach this score, and (2) each step generates fewer tokens because the policy does not need to continue every trajectory to completion. Conceptually, we demonstrate that we can remove the need to roll out every trajectory to completion, opening up a large previously unexplored design space for LLM RL algorithms.
cs.LG / 118 / 2609.39250
Client and Training Data Selection for Computationally Efficient Synchronized Federated Learning
Abstract
Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a cloud server. Straggling clients have been a problem for FL as they introduce delays in aggregating the local models and hence, the convergence of the global model. Therefore, it is important to have a mechanism that ensures fast convergence of the global model as well as good FL participation rate. Another issue for the convergence of a model in FL is the non-independent and identically distributed (non-iid) data across the clients. Prior approaches based on probabilistic client selection do not work well under non-iid data especially when the number of clients is small. We show scenarios where such approaches fail and propose a joint client-training data selection algorithm for fast convergence of FL models. Our experiments on CIFAR-100 dataset show that convergence of the FL model can be significantly improved over prior works that can consider non-iid data and heterogeneous computation and higher model accuracy.
cs.LG / 119 / 2609.39257
From Benchmarks to Production: Transferring Time Series Anomaly Detection Methods for Electricity Production Monitoring
Abstract
Accurate forecasting of electricity production is essential for maintaining the operational efficiency and strategic planning of energy utilities. In industrial settings, such forecasts are generated daily to ensure supply-demand balance and optimal management of production assets. However, the increasing complexity of modern power systems and data flows poses significant challenges for ensuring the reliability and consistency of these forecasts. This paper addresses the problem of anomaly detection in short-term production forecasts at EDF, formulated as identifying atypical intra-day patterns that may signal data quality issues or operational irregularities. We introduce TAMIS, a scalable and interpretable system that analyzes daily production time series to automatically detect anomalous days based on deviations from historical patterns learned from past data. Designed for human-in-the-loop workflows, TAMIS surfaces top-ranked anomalies through an automated daily newsletter, enabling efficient expert review and continuous monitoring. An extensive experimental evaluation on real-world industrial data demonstrates that TAMIS achieves the best accuracy-efficiency trade-off compared to baseline methods. To foster further research and reproducibility, we publicly release the anonymized application datasets used in our study.
cs.LG / 120 / 2609.39261
Jacobian Rank Collapse in Decision-Focused Learning
Abstract
Decision-focused learning (DFL) trains predictors through downstream objectives, but a different loss need not provide an independent parameter-update direction. We characterize this restriction through the predictor Jacobian, using sparse index tracking to distinguish the covariance entries read by the optimizer from the parameter directions available to learning. Rank-one Jacobians make nonzero per-example gradients collinear; a conditional spectral bound describes near-collinearity. A batch-subspace characterization and counterexamples show why these local statements imply neither common minimizers nor collinear batch updates. Experiments examine when geometry translates into decision quality. Across 38 one-parameter equity configurations, DFL gains over MSE remain below 1.8%; a 385-parameter conditional predictor also has pointwise rank one. In validation-tuned shortest-path and knapsack experiments, full-capacity SPO+ reduces mean regret by 11.6% and 10.6%, respectively; only knapsack survives correction across eight comparisons. The capacity contrast persists on fresh datasets across batch orders and training budgets. Holding expressivity fixed, invertible coordinate scaling lowers spectral effective rank and ordinary SGD gains; compensating for the scaling restores the original trajectories. Financial forward-target controls separate forecast accuracy from decision quality; a matched neural comparison finds no aggregate DFL advantage in the tested architecture. These findings distinguish local rank restrictions, coordinate-dependent optimization and predictive accuracy. Predictor geometry helps explain available learning directions, while held-out decision quality remains the test of practical benefit.
cs.LG / 121 / 2609.39268
A Time-Aware Bag-of-Receptive-Fields for Interpretable Irregular Time Series Classification
Abstract
Irregular time series, characterized by non-uniform sampling intervals, missing observations, and variable lengths, are ubiquitous in healthcare, mobility, and environmental monitoring, yet effective and interpretable classifiers for this setting are limited. Existing approaches often rely on imputation, which can obscure the temporal structure of the data, or require complex neural architectures that are opaque and difficult to explain. In this work, we extend the Bag-Of-Receptive-Fields (BORF), a fast, deterministic, and interpretable transform for time series, to the irregular setting. Our key contribution is a time-weighted normalization scheme in which each observation is weighted proportionally to its associated time delta, making pattern extraction sensitive to the actual temporal distribution of samples rather than only their index position. This requires deriving an efficient sliding-window recurrence for the time-weighted standard deviation, preserving the linear time complexity of BORF. We benchmark the resulting method against state-of-the-art irregular time series classifiers on datasets from the PYRREGULAR repository, demonstrating competitive classification performance with the added benefit of human-interpretable explanations.
cs.LG / 122 / 2609.39275
ReTaCo: Residual-Target Control for On-Policy Distillation
Abstract
On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-$k$ probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-$k$ mass toward one even after the student matches the teacher's relative probabilities within the top-$k$ set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top-$k$ tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top-$k$ mass $m$, the residual target is $(1-β)(1-m)$ for $β\in[0,1]$: $β=0$ preserves the teacher's mass, and larger $β$ moves more mass onto the top-$k$ tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top-$k$ mass lies between $m$ and $m+β(1-m)$ and increases monotonically with $β$; at $β=0$, underestimated top-$k$ tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.
cs.LG / 123 / 2609.39277
A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models
Abstract
Looped models reason by applying the same block of weights many times, so compressing that block saves memory traffic on every loop. Compressed looped models, however, often collapse, and the collapse is usually blamed on rounding error that accumulates from loop to loop. In this work we test that account on more than 30 models from five families and find, to our surprise, that it holds only for loops that never settle. When a loop settles, a fixed rounding error does not accumulate. It moves the point where the loop settles, much as tilting a bowl moves where a ball comes to rest, and the answer is lost only when the shift is larger than the readout tolerates. This picture lets us predict which models fail from a single label-free measurement, and it tells us why failed models recover: their loops still settle, so a few final loops with 8-bit weights bring the answer back. Motivated by these findings, we build a controller that stops when the model's halting head fires and then finishes with 8-bit loops. On Sudoku-Extreme and Maze-Hard it beats fixed-depth inference by up to 15 points under a third of the weight traffic.
cs.LG / 124 / 2609.39291
Physics-Informed Method of Group Data Handling: Adaptive Construction of Functional Representations with an Application to the Navier-Stokes Equations
Abstract
Physics-informed computational methods usually optimize parameters within a functional representation whose structure is fixed in advance. This work proposes a Physics-Informed Method of Group Data Handling (PI-GMDH), in which representations of coupled physical fields are progressively constructed during solution. Candidate functional directions are evaluated through the first variation of the complete physical and observational objective, introduced in packages, and followed by block-coordinate damped Gauss-Newton coefficient optimization. The framework is demonstrated with tensor-product Chebyshev functions on the incompressible Navier-Stokes equations using a two-dimensional time-dependent Taylor-Green benchmark. Under the tested configuration, adaptive PI-GMDH reached validation and held-out test losses of 5.299e-19 and 5.296e-19 with 204, 201, and 175 active functions for u, v, and p. Complete degree-by-degree and all-terms PI-GMDH variants, together with selected PINN and KAN reference configurations, are used to examine the effect of structural construction policy. The results show that, for this controlled synthetic benchmark, selective progressive construction can provide a favorable combination of accuracy, representation size, and wall-clock time. The comparison is illustrative rather than a claim of universal superiority over alternative physics-informed approaches.
cs.LG / 125 / 2609.39296
Semantic-Aware Joint Source-Channel Optimization for Encoder-Agnostic Digital Video Communication
Abstract
Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most existing approaches rely on computationally intensive deep learning-based video encoders and decoders, which hinders their deployment in resource-constrained scenarios. To address this issue, we propose a lightweight semantic-aware joint source-channel optimization (SAJSCO) scheme that can be integrated into existing digital video communication systems as a plug-in module. Specifically, we develop a video communication system model in which the transmitter jointly optimizes source and channel coding parameters based on the inter-frame semantic importance of the input video and estimated channel state information. On this basis, we formulate an optimization problem that maximizes semantic importance weighted video reconstruction quality under a maximum bitrate constraint. To solve it, we first quantify inter-frame semantic importance using a cosine similarity-based metric with a shifted window mechanism. We then develop a multi-actor proximal policy optimization (MPPO) algorithm to solve the formulated problem by jointly adapting the source compression rate and channel coding rate. The learned policy can be directly applied to different video encoders without encoder-specific retraining or fine-tuning. SAJSCO achieves Bjøntegaard Delta rate reductions of 34.86\% and 18.01\% when integrated with H.265, a conventional video encoder, and DCVC-RT, a deep learning-based video encoder. Over-the-air experiments on a hardware testbed further demonstrate a PSNR gain of up to 1.448 dB with H.265 and an LPIPS reduction of up to 0.033 with DCVC-RT compared with the respective best-performing fixed-parameter baselines.
cs.LG / 126 / 2609.39306
ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation
Abstract
Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.
cs.LG / 127 / 2609.39321
GRPO Training Dynamics for Small Language Models
Abstract
Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constrained environments. In this work, we present a systematic study of GRPO fine-tuning for SLMs ranging from 1.5B to 7B parameters under a practical single- node 8xA100 compute budget. Our study spans multiple model families and reasoning domains, including mathematics, coding, and multiple-choice question answering (MCQ) in science. Across these settings, we analyze how group size affects policy convergence, training stability, and downstream benchmark per- formance. We further characterize tensor-level update dynamics during GRPO training and investigate whether the choice of LoRA target modules and layers can improve the performance of GRPO-tuned models. While our initial GRPO-tuned models outperform their base counterparts on approximately 80% of mathematical benchmark evaluations, they demonstrate limited capability on MCQ and code reasoning tasks. Guided by our mechanistic evaluations, we refined our LoRA and reward-shaping configurations to improve performance in latter domains. These findings provide practical guidance for GRPO training for SLMs.
cs.LG / 128 / 2609.39336
How Many Samples Are Enough for Learning Across Domains?
Abstract
Understanding the fundamental mechanisms of learning is essential for designing systems with strong generalization. Recent studies have shown that increasing the number of training domains, or enlarging the distribution shift among them, improves generalization when each domain contains sufficiently many data samples. However, the conditions under which the data samples can be considered sufficient remain unexplored. In this work, we fill this gap by establishing criteria for per-domain sample requirements based on the presented learning bounds. These criteria not only reveal an inverse linear scaling law between the number of training domains and the number of samples required per domain, but also explain the fundamental rationale behind the assumption of data sufficiency, thereby providing theoretical guidance for assessing the adequacy of existing datasets and constructing datasets. This differs from classical learning theory, as the number of samples required is highly dependent on the number of training domains. Additionally, we prove the close relationship between in-domain learning and out-of-domain generalization through the presented generalization bounds, and lastly discuss some key arguments.
cs.LG / 129 / 2609.39337
WinoTS: Wavelet-based Self-Distillation for Time Series Models
Abstract
Self-supervised pre-training of time series models is currently dominated by next-token prediction and reconstruction objectives. In continuous-valued domains, these paradigms often waste model capacity on high-frequency, point-wise noise at the expense of learning invariant structure. While invariance-based self-distillation has proven highly effective in computer vision, its application to temporal data remains largely underexplored. Effectively adapting such methods to time series requires carefully designed augmentations: spatial operations like cropping can shift the timing of repeating cycles or distort the signal, while basic jittering may provide limited variation. We introduce Wavelet-based self-distillation for time series (WinoTS), an invariance-based pre-training paradigm designed specifically for temporal signals. At its core, WinoTS leverages time-frequency augmentations to construct multi-scale structural views without distorting underlying signal dynamics. Across extensive evaluations, WinoTS outperforms state-of-the-art baselines in long-term forecasting, cross-domain zero-shot transfer, and unsupervised anomaly detection. Notably, linear probing on frozen WinoTS representations frequently surpasses fully supervised models trained from scratch. Systematic ablations demonstrate that WinoTS is a flexible, architecture-agnostic framework yielding gains across time series backbones, and establish that time-frequency transformations provide a principled alternative to vision-style spatial augmentations.
cs.LG / 130 / 2609.39338
Learning Beyond Full Imitation: Task-Preserving Knowledge Distillation
Abstract
Knowledge distillation transfers knowledge by encouraging a student to match a teacher's predicted class probabilities. These probabilities express not only confidence in the correct class, but also relations among incorrect alternatives. Yet closer imitation does not necessarily yield a better student. A student may already distinguish the correct class more sharply than its teacher, so further imitation can require giving back discrimination it has acquired. Our main result is an exact separation between full imitation and conditional learning. When the correct class's score advantage over each alternative must be preserved, full teacher-to-student KL minimization is blocked exactly when the student assigns no more probability than the teacher to every incorrect class. Crucially, the teacher's relative probabilities among incorrect classes remain fully learnable. We characterize the exact price of this transfer: a minimum increase in correct-class log-odds that compensates for the largest conditional-probability mismatch. Label fitting and conditional matching can therefore be completed even as full teacher KL diverges. This separation motivates task-preserving knowledge distillation (TPKD), which keeps the label gradient intact and minimally corrects the conditional gradient so that its output update preserves the label step's gains against every incorrect alternative. The corrected conditional direction retains more than half of the original first-order conditional descent at the same step size, with a tight bound. For a fixed positive conditional target and sufficiently small constant output steps, label and conditional errors vanish together. Experiments trace this learning from exact head updates to ordinary network training. TPKD reaches 88.05% accuracy on CIFAR-100 and 93.81% on CLINC150, improving over standard distillation by 0.47 and 0.35 percentage points across three seeds.
cs.LG / 131 / 2609.39340
ElectrolyteFM: Unifying Electrolyte Property Prediction through Cross-Property Knowledge Learning
Abstract
Electrolyte formulation design requires balancing multiple physicochemical properties, yet existing models often focus on a limited subset. Learning each property in isolation can overlook transferable chemical information, whereas indiscriminate sharing can introduce cross-property interference. Our directed transfer analysis shows that jointly learning two property prediction tasks can improve or degrade prediction relative to separate training, with asymmetric transfer effects between the tasks. We propose ElectrolyteFM, a unified multi-property prediction model which can more accurately predict multiple properties of each electrolyte by effectively identifying and utilizing property-specific features and knowledge shared across properties. More specifically, ElectrolyteFM learns property-specific representations independently and captures cross-property knowledge through a separately trained expert pool. A router selects relevant shared information for each formulation and target property, and property-specific residual adapters convert this information into corrections to the corresponding representation for prediction. Experiments on Electrolyte12 show that ElectrolyteFM reduces normalized mean absolute error averaged across 12 electrolyte properties by 14.8% relative to the strongest electrolyte-specific baseline. On an independent sodium-electrolyte dataset unseen during training, it reduces conductivity mean absolute error by 6.7% relative to the best-performing baseline.
cs.LG / 132 / 2609.39357
Robustifying Asynchronous SGD via Soft Throttling
Abstract
Asynchronous SGD is a popular algorithm for distributed learning where each client's gradient update is applied on arrival. This leads to a speed-up, but also an increased vulnerability to attacks, as fast clients can dominate the total update. We introduce Throttle, a Byzantine-robust generalization of asynchronous SGD where the key idea is to exponentially down-weight updates from faster clients by a factor $q$. Both asynchronous SGD ($q=1$) and synchronous Byzantine-robust SGD ($q\to\infty$) correspond to specific settings of Throttle. We provide a theoretical analysis of the convergence rate and validate the robustness to attacks both theoretically and empirically. Remarkably, our experiments show that this down-weighting mechanism can also improve performance over standard asynchronous SGD even in the non-Byzantine setting.
cs.LG / 133 / 2609.39361
LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators
Abstract
While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.
cs.LG / 134 / 2609.39374
Wavelet Flow Matching for Time Series
Abstract
Synthetic time series are increasingly used for data augmentation, privacy-preserving data sharing, and downstream model development, yet faithfully reproducing both multi-scale temporal structure and cross-channel dependencies remains challenging. We study multivariate time-series generation through flow matching in the wavelet domain. By operating on multilevel discrete wavelet coefficients rather than directly in the time domain, the model represents coarse structure and progressively finer details at separate scales. Their naturally different variances further induce an implicit coarse-to-fine generative process without requiring an explicit multi-scale schedule. Since the transform acts independently on each channel, we pair it with a channel-token transformer whose attention directly models cross-channel dependencies. Across seven benchmark datasets and four sequence lengths, our method is best or tied on a majority of dataset-metric combinations, with the largest and most consistent improvements in Context-FID and discriminative score.
cs.LG / 135 / 2609.39386
When, Not How Much: Evaluating Time-Series Foundation Models on Sparse Events
Abstract
Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on which future periods contain activity. Standard benchmarks do not assess this. On five sparse datasets, we rank positions within forecast windows that contain both events and zeros. The released point forecasts of 12 TSFMs improve chance-corrected average precision over training-free references by at most 0.031, and in chance-corrected AUC the median TSFM falls below them on every dataset. With event supervision, linear probes of six frozen backbones improve on their backbone's point forecast in 29 of 30 backbone--dataset pairs. Averaging the predicted quantiles instead of taking their median improves the ranking of most TSFMs that forecast the median, and on two datasets the strongest such outputs rival the probes. The probes' advantage over raw-context learners depends on the dataset, and under the same probe, pretrained features outperform randomly initialized ones for five of six backbones. For sparse-event ranking, released point forecasts thus add little over simple references, whereas lightweight event heads on frozen TSFMs rank events better than these forecasts, and the best of them exceed gradient-boosted trees trained on the raw context on three of the five datasets. More broadly, assessing pretrained forecasters on tasks beyond value forecasting requires reporting their outputs, supervised probes of their representations, and raw-context and randomized controls side by side, since each supports a different conclusion.
cs.LG / 136 / 2609.39387
Beyond Simulation: Retain-and-Repair Neural Operators for Real-World Adaptation
Abstract
Neural operators increasingly benefit from pretraining on numerical simulations, yet adapting them for real-world prediction remains challenging. We introduce the Retain-and-Repair Neural Operator (R$^2$NO), a framework for adapting simulation-pretrained operators to real-world data while retaining useful pretrained structure. The pretrained operator is first finetuned on real data and then frozen to provide a source prediction, and a shared repair module learns a sequence of refinements from the same observations. Using orthogonal Fourier projections, a spectral ensemble fits a small ridge regression within each cell of the Fourier domain and combines the refinements by weights fitted on a held-out split of the real data. The cells are defined jointly by radial ranges, angular sectors, and measured channels, allowing refinement depth to vary with frequency magnitude, with orientation, and across channels. Including the source prediction as a candidate makes retention available in every cell, and independently trained repair modules enter the same combination as additional candidates. On all RealPDEBench systems and six backbones, R$^2$NO consistently outperforms full finetuning and iterative refinement. The framework treats adaptation depth as a cell-specific choice learned from real data.
cs.LG / 137 / 2609.39390
Decoupled and Distilled: Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer for Few-Shot Class-Incremental Learning
Abstract
Few-Shot Class-Incremental Learning (FSCIL) addresses the challenge of learning new classes from very limited samples while retaining knowledge of previously learned ones. Although parameter-efficient fine-tuning methods with pre-trained models show promise for class-incremental learning, strict gradient-based constraints can be unreliable under severe data scarcity, while multi-expert approaches can impose substantial inference-time costs. We propose TALON (Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer), an inference-efficient FSCIL framework. TALON dynamically allocates an independent LoRA-Teacher to each incremental task for task-specific representation learning, then distills multiple frozen teachers into a unified LoRA-Student through Ensemble Knowledge Transfer, eliminating runtime module selection or generation. A semantic-guided distillation strategy weights teacher contributions by feature-space similarity to mitigate catastrophic forgetting and overfitting. Across three class-order runs, TALON achieves comparable or better mean average accuracy across four FSCIL benchmarks, obtaining 86.68 +/- 1.22% on CUB200, 90.39 +/- 0.27% on CIFAR100, 78.38 +/- 0.94% on ImageNet-R, and 96.34 +/- 0.33% on miniImageNet. TALON uses up to 33x fewer deployment parameters and reduces average inference time per task to 26.7 s, a 41.70% reduction relative to ASP.
cs.LG / 138 / 2609.39405
No Task Vector Is an Island: A Comprehensive Study on the Composability of Task Vectors from On-Policy Distillation
Abstract
Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced by on-policy distillation (OPD) remains largely unexplored. OPD trains a student using teacher feedback on student-generated trajectories, yielding parameter updates that differ from those produced by the teacher model, usually by reinforcement learning (RL). We therefore ask whether OPD task vectors can complement their RL teacher updates and compose effectively across tasks. Across five domains and two model architectures, we find evidence for both forms of composability. Within a task, merging OPD and RL task vectors can outperform both constituent models, even when the OPD student is weaker than its RL teacher. Across tasks, OPD task-vector compositions achieve higher average scores than corresponding RL compositions in seven of eight backbone-merging-rule comparisons. Parameter-space analyses reveal substantial non-collinearity between OPD and RL updates. Experiment in CODE domain on SMOLLM3-3B shows that the combined direction outperforms either constituent direction at the tested global update norm, supporting directional complementarity in this configuration. Across tasks, OPD updates also show lower overlap among the top-10% feed-forward channels ranked by update energy. Together, these results show that weaker standalone performance does not imply weaker task-vector composability. OPD task vectors can complement stronger RL teacher updates and combine effectively across tasks, highlighting composability as a distinct property for understanding and evaluating post-training updates.
cs.LG / 139 / 2609.39408
Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus
Abstract
Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in $H^1(γ)$, we give explicit conditions under which small IID Gaussian initialization yields a high-probability guarantee: at a checkpoint during a high-loss plateau, minimum alignment between the rank-$r$ teacher subspace and the leading $r$-dimensional eigenspace of the predictor's average gradient outer product (AGOP) increases by at least $1/2$, and the minimum refit MSE under unchanged coefficient budgets decreases by more than $0.399$, both relative to initialization. The same trajectory subsequently attains a trained loss below every value in the plateau window. A complementary result treats unequal-weight cubic teachers and small additive Sobolev perturbations using projected-feature refits. For SwiGLU networks with an exactly fitted intercept, we prove leading-AGOP alignment during a loss plateau at fixed width and dimension as Gaussian initialization vanishes, for square-integrable teachers with nonzero Hermite content of degree one, two, or three. A rank-one cubic specialization also gives simultaneous unrestricted-refit gains at a prescribed width. An approximation lower bound further shows that certain interaction targets retain nonzero error when ridge neurons are restricted to shared orthogonal axes within the teacher subspace. Population-moment experiments with ReLU students across 21 teachers and 50 initializations per teacher complement the analysis.
cs.LG / 140 / 2609.39436
From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL
Abstract
Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.
cs.LG / 141 / 2609.39445
Raw-Routed Mixture of Adapters: A Causal Intervention for Routing Collapse in Time Series Foundation Models
Abstract
Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup that underfits heterogeneous regimes. Replacing the head with a mixture of experts is the standard upgrade, but on instance-normalized backbones (the dominant TSFM design class) it fails: routing entropy collapses to zero and one expert absorbs every input, a failure we call normalization-induced routing collapse. Standard MoE rescue mechanisms do not repair it, because the cause is in the router's input, not its optimization. Pre-encoder normalization strips the statistics a router would need to tell regimes apart. A mutual-information decomposition makes this precise and yields a signal-ratio that, computed before training, predicts dataset vulnerability (Spearman $ρ= -0.88$). Eight causal controls, including a vision-modality replication, isolate instance normalization as the cause. The prescription is a minimal causal intervention: Raw-Routed Mixture of Adapters (RR-MoA), which routes on the raw, pre-normalization input. Under a strictly frozen backbone, RR-MoA wins 54/54 comparisons against the strongest fixed adapter and significantly outperforms LoRA, TRACE, AdaMix, and full fine-tuning. The effect generalizes across six backbones and an imputation task. Frozen RR-MoA also beats full fine-tuning by 12-79% (the Frozen Paradox); two architecturally distinct variants confirm the principle generalizes beyond this specific router.
cs.LG / 142 / 2609.39456
Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback
Abstract
This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea leads to a mutual feedback architecture consisting of two components whose outputs are fed back into the other. The final output of this model is defined as the fixed point of this interaction. We provide theoretical analysis that offers interpretation of this model as well as design choices to prevent failure cases. We show the benefits of MEQ through classification and visual grounding tasks spanning various datasets. Quantitatively, our model outperforms or shows competitive performance on concatenation-based multimodal classification problems. Qualitatively, the proposed interactive mechanism allows the model to progressively refine the visual grounding when paired with complementary modality, thus demonstrating the power of mutual feedback under such settings.
cs.LG / 143 / 2609.39464
Understanding Head Geometry and Dynamics in Federated Regression through a Natural Solution Selection Rule: An Unconstrained Feature Model Analysis
Abstract
In federated averaging, local objectives can admit multiple optimal heads, making the aggregate depend on which heads clients return. We study this ambiguity in federated multivariate regression with private backbones and a shared linear head, using an unconstrained feature model (UFM) that treats training-sample features as free variables. We introduce a natural selection rule: each client returns the optimal head closest to the broadcast head. We show that global minimization with a vanishing proximal penalty on the head realizes this rule. When the clients' optimal Gram matrices and the initial shared Gram matrix are positive definite, the shared Gram matrix follows a closed recursion and converges to the unique Bures-Wasserstein barycenter of the clients' optimal Gram matrices. Even with this alignment, the limit generally differs from the centralized optimal Gram matrix. We decompose this gap into three positive-semidefinite terms arising from differences in client target means, covariance heterogeneity, and averaging the aligned heads. A correction based on a one-time exchange of target means and covariances recovers the centralized optimal Gram matrix in one round under exact local optimization and the same selection rule. We verify these results numerically in the UFM and test its predictions on five tabular and five image regression datasets using deep networks with feature regularization and long local training. In these experiments, ordinary training approaches the predicted barycenter, while a weak proximal penalty improves endpoint agreement and yields trajectories that closely follow the predicted Gram dynamics. The correction moves the final Gram matrices close to the centralized UFM prediction.
cs.LG / 144 / 2609.39466
T-ARC: Topology-Aware Randomized Clustering via Distributionally Robust Stochastic Block Models
Abstract
In this work, we introduce a new clustering method, namely T-ARC (Topology-Aware Randomized Clustering), that corrects the geometric bias of K-means by embedding topological information directly into the optimization objective. Building on the assumption that the data admits an underlying hidden structure modeled via a latent graph, the idea is to uncover this information through the interplay between the standard K-means data-fidelity term and a graph-cut penalty, which discourages cluster assignments inconsistent with the connectivity structure of the data. To render this coupling tractable, the latent graph is modeled as a random realization from a Stochastic Block Model (SBM), whose scalar parameter is optimized within a Distributionally Robust Optimization (DRO) framework, yielding a closed-form proximal update. Both SBM and DRO are informed by a persistence-based similarity matrix derived from zero-dimensional persistent homology ($H_0$), which translates the multiscale connectivity structure of the data into a pairwise topological prior. The overall optimization proceeds via Block Coordinate Descent; convergence is established through a global Lyapunov functional: the deterministic blocks satisfy monotonic descent, while the stochastic graph update satisfies descent in expectation, so that the expected energy converges. Experiments on synthetic datasets with non-convex geometries and on random subsets of Fashion-MNIST show that T-ARC recovers latent topological structures where K-means fails, achieving the highest accuracy on curved and interleaved clusters while remaining competitive, and markedly more stable than K-means, on real data.
cs.LG / 145 / 2609.39478
Also Small Models Can Reasonably Self-Evaluate Their Confidence
Abstract
This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains. Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals. Our results reveal that while accuracy predictably declines with smaller models and more specialized domains, the reliability of self-evaluated confidence remains largely stable across both dimensions. This independence means the most capable model is not necessarily the best at self-assessing prediction reliability. These findings suggest that smaller models can achieve reasonable self-assessed confidence despite lower accuracy, making them viable for resource-constrained deployments.
cs.LG / 146 / 2609.39488
Correcting CondOT: Exact Finite-Step Sampling in Gaussian Flow Matching
Abstract
Flow matching generates samples by gradually transforming noise into data. In practice, using a finite number of sampling steps introduces a numerical error that depends on the chosen schedule. We study this dependence for Gaussian targets and the explicit midpoint sampling method, using the exact flow field. We measure sampling error by the squared Wasserstein distance between the target distribution and the final distribution produced by the midpoint sampler. We show that the standard conditional optimal transport (CondOT) schedule cancels the leading midpoint error and improves the general convergence bound, even when the sampling steps are unequally spaced. On a uniform grid of $S$ sampling steps, we fix the signal schedule at $α_t=t$ and prove the existence of scalar noise schedules $β_t$ that approach the CondOT noise schedule $1-t$ at rate $1/S$ and yield exact Gaussian sampling for every sufficiently large $S$. Controlled Gaussian experiments illustrate the convergence rates and exact calibration.
cs.LG / 147 / 2609.39489
Towards Robust Time Series Learning via Capacity-Centric Modulation
Abstract
Sample-level reliability heterogeneity is common in deep time series learning. Standard training pipelines apply a uniform regularization setting to all samples, which can under-regularize corrupted samples and over-restrict clean samples. Common robustness approaches filter observations in data space or impose priors on latent representations. We propose Capacity-Centric Modulation (CCM) as a complementary, sample-adaptive regularization principle. Under this principle, we introduce SACM (Sample-Adaptive Capacity Modulation), a task-agnostic framework that exploits spectral sparsity to assign sample-wise dropout probabilities along internal activation paths. SACM integrates into existing backbones without architectural redesign and preserves the deterministic inference pipeline. Across 301 real-world dataset-backbone pairs covering 9 forecasting, 32 classification, and 4 anomaly-detection datasets, SACM reduces forecasting MSE by 6.7% on average and improves classification accuracy and point-adjusted F1 by 3.04% and 17.05%, respectively, relative to unmodified backbones, with zero test-time overhead.
cs.LG / 148 / 2609.39496
When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev
Abstract
Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.
cs.LG / 149 / 2609.39497
The Geometry of Randomized Smoothing on Feasible Sets
Abstract
Randomized smoothing certifies the probability of a fixed output event as the center of Gaussian noise moves. Feasibility or confidence filtering reports label probabilities only among retained proposals, producing a ratio. Its numerator is a fixed Gaussian event mass, while its denominator is the probability of retention and can change with the center. Substituting this ratio into the ordinary smoothing formula can therefore certify a ball that contains a decision boundary. We separate the problem into a geometric question and a certification question. Geometry determines when conditioning preserves Gaussian comparisons. Convex retained sets preserve the full comparison, while general sets require geometric control of the retained law as the center moves. Without such control, conditional probabilities imply no positive universal radius. Joint retention-and-label probabilities always yield a valid certificate for the same filtered predictor. A uniform covariance bound transfers divergence certificates to the retained law and can yield larger radii even when the Gaussian event comparison fails. Both methods admit finite-sample bounds. For a learned image classifier with a training-selected nonconvex filter, conditional Rényi bounds certify more images than joint-mass bounds without additional model evaluations. A released confidence filter exhibits verified label changes inside radii obtained by conditional substitution. An application of adaptive Gaussian composition covers causal finite-horizon executions with history-dependent center shifts under a pathwise energy bound.
cs.LG / 150 / 2609.39512
Can Domain Generalization be Guaranteed in Small-Sample Learning?
Abstract
The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation and generalization. Structural Risk Minimization (SRM) has long been regarded as a principled solution under the classical i.i.d. assumption. However, domain generalization (DG) violates this assumption, leaving the theoretical role of SRM in DG largely unexplored. To bridge this gap, we establish the first theoretical guarantees for SRM in DG under mild assumptions. Specifically, based on the concept of stability, we derive learning consistency and generalization error bounds and prove that these bounds become tight when the hypotheses satisfy the stability condition. Building upon this, under a specific hypothesis space assumption, we establish stability, learning, and generalization bounds for SRM. We further discuss the applicability of these bounds to deep learning. This work establishes theoretical foundations for SRM under distribution shifts and sheds light on the design of robust DG algorithms in small-sample scenarios.
cs.LG / 151 / 2609.39523
CIDER-FM: Foundation Models for Causal Inference from Diverse Experimental Regimes
Abstract
Causal foundation models (CFMs) amortise causal inference over priors of synthetic structural causal models (SCMs), predicting the effect of an experiment on a specific variable. However, observational data alone may leave multiple causal models compatible with available evidence, while experimental data with interventions on exactly the variable of interest might be unavailable. This work studies CFMs as a method to combine finite observational and surrogate-interventional datasets in order to predict a target conditional interventional distribution (CID) more accurately than with observational data alone. We first formalise the conceptual benefits of surrogate experiments. Building on this analysis, we introduce \textsc{Foundation Models for Causal Inference from Diverse Experimental Regimes} (\emph{CIDER-FM}), a causal foundation model that uses an intervention-aware representation and hierarchical three-axis attention to exchange information across variables, samples, and experimental regimes. We evaluate CIDER-FM against a wide range of baselines across diverse synthetic graph and mechanism families, as well as on both simulated and real-world data from Causal Chambers. Our results demonstrate strong CID prediction performance and show that incorporating experimental context can improve predictions over observational data alone.
cs.LG / 152 / 2609.39547
Learning Reliable GUI Agents under Imperfect Priors
Abstract
GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that is scarce in pretraining corpora. Retrieval-augmented execution offers a natural remedy but faces two coupled bottlenecks: knowledge at scale is hard to acquire, and self-collected priors inevitably drift from the live environment due to version updates, promotions, ads, A/B tests, and personalization. We therefore argue that GUI agents should not pursue perfect knowledge but learn to act correctly under imperfect priors, and propose our framework that couples knowledge acquisition with noise-robust utilization: a structured exploration strategy traverses interactive elements, builds a UI state-transition graph, and synthesizes (task, trajectory) pairs via a VLM without human annotation; a noise-aware training strategy, grounded in a taxonomy of real GUI drift patterns, injects five types of realistic errors into self-explored trajectories to teach the agent to assess prior reliability before acting. Experiments on physical devices and online emulator benchmarks show that our method discovers more unique screens, covers more benchmark tasks, and more effectively rejects erroneous priors while leveraging correct ones, with accuracy gains that transfer across datasets.
cs.LG / 153 / 2609.39550
Hyperbolic Prototype Routing for Rehearsal-Free Class-Incremental Learning
Abstract
Class-Incremental Learning (CIL) aims to continually learn new classes while preserving prior knowledge. Parameter-efficient fine-tuning with pre-trained models enables CIL with minimal parameter updates, but existing approaches still suffer from catastrophic forgetting caused by cumulative interference and suboptimal module-sample matching at inference. We propose Hyperbolic Prototype Routing (HyPro), a rehearsal-free framework for continual learning. HyPro allocates a dedicated LoRA-Expert module to each incremental task for isolated representation learning, then projects routing features onto a Poincare ball and performs geodesic nearest-prototype matching for reliable task-level discrimination. Extensive experiments on standard CIL and Few-Shot CIL benchmarks show that HyPro consistently improves average and final-stage accuracy over strong baselines.
cs.LG / 154 / 2609.39561
Candidate Retention for Abductive Learning
Abstract
Abductive learning combines neural perception with symbolic reasoning, using explanations generated by abduction to supervise the perception model. Multiple valid explanations of the same symbolic target can assign conflicting labels to the same inputs. Common policies select a single candidate as a pseudo-label, which may reinforce mistaken assignments, or weight all candidates, which may spread supervision across competing labels. These risks motivate selecting a retained subset to balance supervision sharpness and model-mass coverage. To guide this choice, we bound the coordinate-level supervision error using retained uncertainty, discarded model mass, and model mismatch. For a fixed model and training pair, only the first two terms depend on the retained set. We propose Abductive Candidate Retention (ACR), which uses these terms to guide greedy additions, accepting a candidate when its recovered mass exceeds the increase in retained uncertainty. Experiments show that ACR improves concept accuracy over single-candidate baselines and A3BL in most evaluated aggregated mod-addition settings. Objective ablations support the joint use of uncertainty and posterior mass.
cs.LG / 155 / 2609.39581
Robust Transfer Learning for Paper ECG Recognition
Abstract
Paper ECG recognition is challenging because real-world ECG images vary in layout, physical artifacts, and label availability. We introduce RobECG-CL, a rank-aware contrastive learning framework for robust paper ECG representation learning. Starting from standard 12-lead ECG recordings, we construct progressively degraded paper ECG views with heterogeneous layouts and train the model to balance same-recording invariance with degradation-aware ordering. Across synthetic stress tests on CODE-II and EchoNext, RobECG-CL improves robustness under severe degradation and few-shot transfer, outperforming contrastive learning baselines and surpassing the waveform-based foundation model, ECG-FM, in the 1% labeled setting. On 312 samples of hospital data with 37 labels, RobECG-CL achieves the best macro AUROC.
cs.LG / 156 / 2609.39595
Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum
Abstract
Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an $\mathcal O(T^{-1/4})$ bound on the expected average Frobenius gradient norm. The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values. Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed. The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift. For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds. A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition. The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.
cs.LG / 157 / 2609.39613
Hybrid Methods for Robust Tabular Data Imputation
Abstract
Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstream inference. In this work, we propose two hybrid imputation methods called NuclearForest and SoftForest, which combine nuclear-norm-based low-rank initialization using Singular Value Thresholding (SVT) and SoftImpute, respectively, with a non-iterative Random Forest refinement. For the SVT-based component, we further introduce an adaptive step-size rule, prove adaptive step-size bounds, and establish convergence for the corresponding zero-initialized iteration. The low-rank initialization provides a structured warm start that captures the global covariance patterns in the data, while the subsequent Random Forest step recovers residual nonlinear signals encoding local dependencies. We conduct an extensive benchmark on diverse datasets from different application domains, comparing the proposed methods with seven established imputation methods under the Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) mechanisms across varying missingness rates. Our results demonstrate that NuclearForest and SoftForest match or exceed the imputation fidelity of state-of-the-art iterative methods such as MissForest, while significantly reducing computational cost. In particular, they achieve speedups of approximately 5.81 times and 9.52 times over MissForest by replacing iterative cycles with a single refinement step. Our approach effectively exploits the low-rank structure of real-world tabular data and accommodates mixed-type variables, providing an efficient and robust solution for data imputation in bioinformatics, economics, and beyond.
cs.LG / 158 / 2609.39629
Certification-Based Differentially Private Learning
Abstract
Differential privacy (DP) in machine learning is typically achieved by adding noise to model parameters (private learning) or to model outputs (private prediction). Recent work uses formal methods, namely abstract interpretation, to provide tighter privacy guarantees, but only for private prediction in classification settings. In this work, we investigate the use of formal methods as a general tool for tighter privacy analysis. First, we generalize the abstract gradient training (AGT) framework to private prediction in continuous, unbounded regression. Second, by reducing learning in parameterized models to a regression problem over the parameter space, we introduce Abstract Gradient Sampling (AGS), an algorithm that enables reachability-based analysis to provide guarantees for private learning. In both private prediction and private learning, we provide tightened privacy accounting for the AGT framework and a theoretical analysis demonstrating when our smooth sensitivity upper-bounds yield favourable privacy-utility trade-off. In practice, we validate that our regression bounds are tighter than global-sensitivity baselines on regression benchmarks, and, notably, yield the first finite privacy guarantees in settings where global prediction sensitivity is a priori unbounded. We also find that under matched conditions, our private learning algorithm can outperform standard private learners.
cs.LG / 159 / 2609.39630
PEG-Tab: Sampling-Time Record Repair and Release Control for Tabular Synthesis
Abstract
Pretrained tabular generators can reproduce training records even when aggregate utility remains high. When retraining is unavailable or too costly, sampling and release are the remaining intervention points. We present PEG-Tab (Post-Training Energy Guidance for Tabular Synthesis), a post-training repair and release-control framework for frozen tabular generators. For each generated row, a generator-native operator creates two alternatives. A shared calibrated score compares the three candidates, favours lower-risk records, and applies a final release check. We instantiate this interface for GReaT, CTGAN, TVAE, and TabDDPM without updating their parameters. Across five datasets and four generator families, PEG-Tab reduces mean Near Copy from $0.078$ to $0.027$ and lowers aggregate Exact Copy to zero. Relative to a $3\times$ post hoc filter, it retains higher utility in 12 of 16 transfer settings and Pareto-dominates the filter in eight. Gains are concentrated in copy and proximity-related risks.
cs.LG / 160 / 2609.39632
Towards Better Exploration in Sequential Test-Time Scaling
Abstract
Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.
cs.LG / 161 / 2609.39634
Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse
Abstract
Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autoregressive generation and canonical-order constructive optimization are both history-injective. The exact correction pays importance-sampling variance that grows with the horizon, so we generalize it to a one-parameter family with PPO ($α{=}0$) and the full correction ($α{=}1$) as endpoints: a single bias--variance knob. A gradient-level analysis of the unclipped surrogate identifies two channels the correction acts through and three conditions under which it carries signal; an enumerable testbed confirms the conditions' predictions. On hard credit-assignment scheduling tasks, a short corrected warmup with aggressive early sample reuse learns faster than PPO and than the same reuse uncorrected; the marginal gain grows with task difficulty ($+0.02$ to $+0.09$ learning-curve AUC), and the early win over PPO tracks the prefix bias that reuse incurs. A correction held throughout, or applied where clipping already contains the reuse bias, is null to harmful.
cs.LG / 162 / 2609.39644
RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements
Abstract
Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions and stochastic variability. Consequently, models that accurately predict measured profiles may reproduce technical effects rather than recover the underlying biology. We ask whether jointly modeling datasets collected under different experimental conditions can reveal shared, sequence-dependent patterns of ribosome occupancy. We introduce RiboUnmix, a probabilistic multi-dataset framework in which each expected measured profile is represented as a shared sequence-dependent signal modulated by a dataset-specific multiplicative factor. A negative-binomial observation model captures variability across replicates. We evaluate RiboUnmix on a controlled synthetic benchmark combining programmed translation kinetics, ribosome traffic, stochastic count sampling, and sequence-dependent experimental distortions. Because the underlying kinetics and distortions are known, recovery of the shared profile and dataset-specific effects can be assessed separately. Both inferred components correlate strongly with their targets, demonstrating that RiboUnmix can disentangle shared kinetic patterns from experimental effects. Across four organism-specific real-data benchmarks, RiboUnmix outperforms sequence-to-profile baselines in predicting measured profiles. Models trained independently on subsets of 114 HEK-derived datasets recover concordant shared profiles for held-out transcripts, and experiments varying the number and composition of training datasets show that the learned representation remains stable. RiboUnmix thus converts variation across experiments into evidence for reproducible sequence-dependent patterns of ribosome occupancy, supporting biological hypothesis generation from diverse Ribo-seq datasets.
cs.LG / 163 / 2609.39646
Beyond Uniform Compression: Budgeted Transmission Allocation for Extreme Federated Learning
Abstract
Federated learning faces severe communication bottlenecks when clients upload high-dimensional model updates. Existing methods often compress these updates uniformly across all layers. This uniform approach ignores the heterogeneous value of different parameter blocks and wastes limited bandwidth on insensitive layers. To address this issue, we propose Layer-wise Budgeted Adaptive Transmission (LBAT). LBAT reframes federated communication under extreme uplink budgets as a resource allocation problem. Our framework dynamically estimates the transmission value of different layers utilising local training signals. It then employs an exact byte dynamic programming allocator to determine optimal rank and bit configurations under strict budgets. We validate LBAT on highly heterogeneous federated tabular prediction and data generation tasks. Extensive experiments demonstrate that LBAT consistently outperforms uniform rank, uniform quantisation, and fixed compression baselines across various extreme budget regimes. Furthermore, it achieves significantly better communication and utility tradeoffs while preserving essential distributional fidelity.
cs.LG / 164 / 2609.39673
NodeGround: A Node Classification Benchmark in the Graph Foundation Model Era
Abstract
Can a pretrained graph model replace training and tuning a separate predictor for each dataset? Answering this requires evaluating prediction quality alongside computational cost. We present NodeGround, a node classification benchmark that puts graph foundation models (GFMs) and dataset-specific supervised learning under a common evaluation framework. The benchmark spans 51 datasets and evaluates six GFMs alongside 15 supervised methods under two label-availability regimes. Shared data partitions, validation-only model selection, controlled hyperparameter searches, and multiple predictive metrics make comparisons systematic, while workflow measurements account for adaptation, training, tuning, and inference. The results favor carefully tuned graph neural networks overall. GraphPFN reaches third place by Elo when more labels are available, yet its relative strengths vary substantially with dataset properties. Efficiency comparisons further qualify the benefits of pretrained reuse: GVT and GraphPFN appear on the Pareto frontiers when supervised methods are represented by their default and fully tuned configurations. Adding intermediate tuning budgets removes this advantage for GVT and leaves GraphPFN extending the estimated frontier in the label-rich setting alone. Thus, reusing pretrained parameters does not yet provide a broadly reliable route to either stronger predictions or cheaper workflows. We release the evaluation pipeline, run-level records, and an open leaderboard at https://github.com/nums-ai/nodeground.
cs.LG / 165 / 2609.39737
A library for differentiable signal processing and machine learning on the sphere
Abstract
The two-dimensional sphere embedded in three-dimensional Euclidean space S2, plays a central role in a variety of scientific and engineering domains, including geophysics, planetary science, geodesy, atmospheric physics, quantum chemistry, cosmology, and virtual reality, among many others. As machine learning increasingly permeates these fields, the demand grows for robust tools that process and model functions on the sphere, while respecting the inherent topological and symmetry properties of the domain. We present torch-harmonics, a comprehensive library that offers efficient, differentiable implementations of advanced signal processing and machine learning (ML) methods for spherical data. These include the spherical harmonic transform (SHT), the spherical analogue of the Fourier transform, vector spherical harmonics, discrete-continuous and spectral convolutions, as well as both global and neighborhood spherical attention mechanisms. Beyond traditional representations, torch-harmonics provides the building blocks for state-of-the-art spherical ML architectures such as spherical transformers in order to enable scalable, rotationally-aware learning and inference in modern scientific and engineering applications.
cs.LG / 166 / 2609.39738
CNCGEN: A Dataset and Framework for Machining Process Planning and Toolpath Generation from B-rep Models
Abstract
Learning to generate machining process plans and toolpaths from B-rep CAD requires coupling discrete operation decisions with continuous tool motion as the workpiece evolves. Correctly predicting an operation sequence does not by itself ensure correct material removal, because each toolpath acts on the stock left by preceding cuts. We formulate this problem around persistent manufacturing objects: object identity determines the target of an operation, while the evolving stock state conditions the generation of its toolpath. Based on this formulation, we propose CNCGEN, a dataset and learning framework for three-axis machining. CNCGEN-Dataset contains approximately 50k geometrically verified synthetic machining flows and 800 held-out real CNC records. Each flow aligns B-rep geometry with object-referenced operations, parameterized toolpaths, intermediate stock states, and verification outcomes, enabling supervision of the correspondence between planning decisions and their geometric effects. CNCGEN generates operations and toolpaths for selected objects step by step, updating a compact machining state to guide subsequent predictions. During training, a learned surrogate verifier provides material-removal feedback that links local predictions to their geometric consequences. Experiments on synthetic and held-out real CNC records show that CNCGEN improves the resulting workpiece geometry and reduces residual material and overcut compared with adapted CNC generation baselines.
cs.LG / 167 / 2609.39739
Stable Transformers for Graph Generation
Abstract
Graph generative models increasingly rely on Graph Transformers (GT) to capture complex dependencies among nodes and edges. While deeper architectures should provide greater expressive capacity and a broader receptive field, their effectiveness can decline with depth: repeated self-attention progressively contracts node representations, impeding information flow and gradient propagation. We analyse this phenomenon from a dynamical systems perspective, focusing on how the denoiser's spectral dynamics affect graph generation. We show that standard GT denoisers become increasingly dissipative as depth grows, leading to vanishing gradients and representation collapse. To isolate the effect of these dynamics, we construct a permutation-equivariant GT with inherently stable, non-dissipative transport. We also introduce a damping mechanism that continuously interpolates between non-dissipative and increasingly contractive regimes, enabling a direct assessment of how dissipation influences generation. Experiments on synthetic and molecular graph generation benchmarks show that the gap between these regimes widens with depth: non-dissipative dynamics preserve representation diversity and gradient flow, sustaining strong generative performance, whereas greater contraction progressively impairs it. These findings identify the denoiser's dynamical regime as a key design factor for deep graph generative models.
cs.LG / 168 / 2609.39741
The Nixtlaverse: An Open-Source Ecosystem for Forecasting
Abstract
Large forecasting applications often combine statistical, machine-learning, and neural models. These families solve the same problem but differ in fitted state, training procedures, and how they parallelize work. Forecasting software must therefore either hide these differences behind a single estimator interface, or keep the families in separate packages, forcing users to rewrite data preparation and evaluation for every package. We present the Nixtlaverse, an ecosystem of open-source Python libraries for time series forecasting, as a case study of a third design: all libraries share the same long-format panel data and keyed forecast outputs, while every model family keeps its own specialized implementation. We demonstrate this design through three use cases on the public M5 competition data. First, we evaluate statistical, machine-learning, and neural models, and an external engine from a separate ecosystem, in a single rolling-origin evaluation with per-series and hierarchy-weighted metrics. Second, we profile runtime and peak memory from 100 to 30,490 series and locate each family's bottleneck: statistical fitting scales approximately linearly in the number of series, feature construction dominates machine-learning memory, and neural training time is nearly independent of panel size under a fixed training budget. Third, we reconcile the forecasts of multiple engines, including the external one, over all 42,840 series of the M5 hierarchy, with sparse reconciliation where dense implementations exhausted memory. These use cases establish the costs, boundaries, and utility of shared data and output contracts. The Nixtlaverse has seen substantial public distribution, scholarly reuse, and adoption through other forecasting frameworks, and is released under permissive open-source licenses with public datasets, reproducible examples, and verifiable benchmark artifacts.
cs.LG / 169 / 2609.39749
Validity-Preserving Hierarchical RL for Joint Routing and Switch Placement in EDA
Abstract
Routing and switch placement are fundamental combinatorial optimization problems in chip design, requiring the joint optimization of routing topology and physical placement under strict structural, geometric and logical constraints. Existing approaches typically rely on carefully engineered heuristics that incorporate strong problem-specific biases to navigate the enormous space of possible designs. In this work, we introduce a hierarchical reinforcement learning framework for joint routing and switch placement at the level of logical communication routes. Starting from a minimal routing graph, our method progressively constructs increasingly expressive solutions through three coupled operations: switch expansion, switch placement, and route refinement. These operations preserve routing validity by construction, restricting exploration to feasible configurations where every communicating initiator-target pair has one assigned loop-free route. We explore the induced solution space using Gumbel Monte Carlo Tree Search, showing that neural-guided search substantially improves solution quality over non-learning optimization methods. Furthermore, pretraining across floorplans provides a strong initialization for fine-tuning on unseen instances.
cs.LG / 170 / 2609.39767
How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
Abstract
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.
cs.LG / 171 / 2609.39773
Riemannian Flow Models with Reinforcement Learning for Molecular Crystal Structure Prediction
Abstract
Crystal structure governs material properties, making crystal structure prediction (CSP) a fundamental problem in materials science. Generative models are a promising approach for solving this problem, but the prevalence of polymorphism, coupled with large unit cells and complex packing geometry, makes the molecular CSP task challenging for existing models. To address this, we introduce Coarse-Grained Open Materials Generation (CG-OMatG), an equivariant Riemannian flow-based generative model. CG-OMatG predicts molecular crystal structures \textit{via} a coarse-grained, hierarchical representation. CG-OMatG treats molecules as rigid bodies---performing both inter- and intra-molecular message passing to construct a geometric representation for molecular packings---and learns to reconstruct molecule centroid positions, orientations, and lattice parameters, conditioned on chemical species and conformer geometry. We train the model on subsets of the Open Molecular Crystals (OMC25) and Cambridge Structural Database (CSD) datasets. Further, we fine-tune the model \textit{via} policy gradient reinforcement learning to steer the model towards generating low-energy candidate structures. We validate the generated structures on the CSP blind test benchmark, assessing agreement with experimentally determined crystals using COMPACK packing-similarity analysis. CG-OMatG exhibits strong performance for generative molecular crystal structure prediction, paving the way for accelerated polymorph screening and organic solid-state materials discovery.
cs.LG / 172 / 2609.39777
GraphMAS: A Systematic Benchmark of Multi-Agent Coordination for Graph Learning
Abstract
LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems because useful evidence may arise from heterogeneous local, long-range, global structural, and semantic perspectives whose relevance varies across instances. Existing LLM-based graph learning approaches primarily rely on single-agent reasoning, while multi-agent coordination has been studied mainly in general reasoning settings. Consequently, it remains unclear whether multiple specialized agents can improve graph learning and how coordination strategies should be designed and evaluated. To address this gap, we introduce GraphMAS, a systematic benchmark of multi-agent coordination for graph learning. GraphMAS builds a shared pool of graph reasoning specialists and organizes coordination along two dimensions, inter-agent interaction and runtime adaptivity, yielding four paradigms and seven representative coordination methods. Under a unified protocol, we evaluate these methods across seven text-attributed graphs, three domains, and two graph learning tasks. We find that heterogeneous graph perspectives are complementary, and that coordinating specialists improves over individual specialists and single-agent graph reasoning, with gains from decomposing reasoning across specialists rather than from broader evidence access alone. However, richer inter-agent interaction does not reliably help, whereas instance-adaptive specialist selection yields the strongest accuracy-efficiency trade-off. We further show that coordination can be learned over a fixed specialist pool and transfers to held-out graphs. GraphMAS therefore provides a controlled evaluation framework and empirical principles for understanding when and how multi-agent coordination benefits graph learning.
cs.LG / 173 / 2609.39784
CORD: Learning Reusable Degradation Representations Across Heterogeneous Physical Systems
Abstract
Can heterogeneous physical degradation systems benefit from joint pretraining and move beyond system-specific prognostics toward reusable cross-system representation learning? CORD combines type-specific observation interfaces with a shared degradation backbone. Its two self-supervised objectives learn at complementary scales: Intra-Observation Structure Modeling (ISM) captures structure within observations, while Inter-Observation Dynamics Modeling (IDM) captures latent degradation evolution across observation histories. We evaluate CORD under two transfer boundaries: Pretraining-Included System Types, where downstream datasets and held-out units are unseen but their system types are represented during source pretraining, and Pretraining-Excluded System Types, where the entire turbofan-engine type is absent from pretraining. Across bearings, batteries, and cutting tools, CORD (Multi-domain) consistently improves over CORD (Single-domain) under Frozen adaptation, provides further gains under Full FT in most settings, and remains competitive with representative external baselines. Source-pretrained initialization also improves low-label adaptation to the pretraining-excluded engine type. Frozen-representation analysis further shows improved cross-unit lifecycle consistency after multi-domain pretraining. Joint pretraining across heterogeneous physical systems thus produces degradation representations reusable across devices, datasets, and system types.
cs.LG / 174 / 2609.39789
Pseudo-Label-Triggered Retraining from Forecast Errors for Online Time Series Forecasting
Abstract
Real-world time series forecasting systems operate under non-stationary data streams, where forecasting performance may degrade over time. Although retraining can recover the performance, it incurs non-trivial computational and operational costs. Under limited deployment resources, the key challenge is therefore not only how to retrain but also when to retrain. While existing retraining policies often rely on indirect indicators such as drift alarms or model staleness, we instead use realized forecast errors as direct deployment feedback. In this paper, we propose PILOT (Pseudo-label-Informed Learned Online Trigger), an online retraining framework that learns when to retrain from forecast-error dynamics. Since ground-truth retraining labels are unavailable, PILOT constructs a pseudo-label from future increases in forecast error and trains a lightweight scorer to predict it from observed error states. At deployment, PILOT uses only completed forecast errors and serves as a plug-in module for arbitrary forecasting backbones without architectural modification. We evaluate PILOT under standard multivariate forecasting settings across eight benchmarks with three representative backbones---DLinear, iTransformer, and TimesNet. Across all three backbones, PILOT achieves state-of-the-art average-rank performance among retraining policies while maintaining a favorable performance--efficiency trade-off.
cs.LG / 175 / 2609.39792
TopTimeNet: Topologically-assisted time-series classification model
Abstract
Distinguishing periodic from chaotic dynamics in a time series is a fundamental challenge in both physics and engineering. Yet, end-to-end learned architectures must discover both a representation and a decision boundary from data, at substantial cost. We introduce TopTimeNet, which decouples these tasks: a fixed, non-learned stage extracts a $42$-dimensional geometric and topological descriptor from Takens delay embeddings and persistent homology, and a lightweight learnable stage performs classification. On a benchmark of $49$ nonlinear dynamical systems, a $1{,}638$-parameter configuration matches the mean accuracy of one with $33\times$ more trainable parameters. Additionally, this approach delivers mean accuracy comparable to convolutional neural networks and surpasses the average performance of converged Transformer models, while requiring three to four orders of magnitude fewer trainable parameters. Robustness also depends sharply on where noise is introduced: TopTimeNet degrades gracefully under perturbations to its precomputed features, but degrades sharply when noise is introduced into the raw signal and the full feature-extraction pipeline is recomputed, showing that robustness to perturbations of the precomputed features does not imply robustness of the complete raw-signal-to-prediction pipeline. These results show that decoupling fixed geometric and topological feature construction from a lightweight discriminative stage can achieve comparable classification accuracy with substantially fewer trainable parameters.
cs.LG / 176 / 2609.39798
Probabilistic Adversarial Training
Abstract
Building on a probabilistic perspective in which adversarial examples arise from the overlap between a distance-based distribution $p_{\mathrm{dis}}$ and a victim-classifier-induced distribution $p_{\mathrm{vic}}$, we start from a simple intuition: adversarial examples become harder to generate when these two distributions are pushed apart, as their overlap becomes smaller, thereby increasing robustness. This intuition naturally motivates a KL-based robustness objective. We then prove that $\mathrm{KL}(p_{\mathrm{dis}}\|p_{\mathrm{vic}})-\log Z_{\mathrm{vic}}$ is a lower bound on probabilistic robustness (PR), where $Z_{\mathrm{vic}}$ denotes the normalizing constant of $p_{\mathrm{vic}}$. Since PR is generally intractable to compute directly, maximizing this KL-based lower bound provides a tractable surrogate objective for improving PR. We further show that this objective recovers a scaled form of adversarial training, offering a probabilistic interpretation of adversarial training and a principled route to robustness improvement. We call the resulting method probabilistic adversarial training. Experiments show that it consistently improves PR, and ablation studies demonstrate that the induced scaling factor can even enhance the PR of non-probabilistic adversarial training methods.
cs.LG / 177 / 2609.39800
Finite-Horizon Fisher Memory in Two-Sided Power-Bounded Recurrent Systems
Abstract
We analyse allocation, admission and post-write retention in finite-horizon linear-Gaussian noisy recurrent memories. At every horizon, the directional Fisher memory $M_n$ satisfies $\operatorname{tr}M_n=N$: non-normality redistributes information but cannot raise its spherical average, while normal carriers satisfy $M_n=I$. For bi-power-bounded carriers, we derive uniform $1/n$ lag bounds, identify the limit of $M_n$ with the inverse of the classical Cesàro asymptotic limit of $W^\top$, and give finite-horizon error bounds. A time-varying coupling defines an end-to-end store operator. The writer-optimal direction need not be store-optimal. After writing ends, an invertible hold preserves the full stored Fisher matrix. Additive contamination bounded by $α$ times the closure covariance retains at least $1/(1+α)$ of that matrix; a covariance-aware decoder attains the corresponding accuracy. With recurrent carriers held fixed, training input masks and linear readouts approached the task-specific optimum in 160 runs, with median normalized Rayleigh efficiency above $0.998$. Binary accuracy matched the Gaussian prediction to mean absolute error below $0.002$ over more than four orders of magnitude in $J$. In a separate pre-specified study of 320 runs, trained masks followed the designated input-time objective in both carrier types, in 16 of 16 draws. These studies used development-seen carriers and are pre-specified validations, not blind holdouts. The same fixed design reproduced the objective-specific result in 16 of 16 draws on carriers unused before run commitment. Exact isolation preserved information, while a decoder fixed at its training horizon fell to chance; inverse-adjoint transport restored its sampled decisions to numerical precision.
cs.LG / 178 / 2609.39810
A Comprehensive Benchmark of Source-Free Universal Domain Adaptation on Time Series Representations
Abstract
Source-Free Universal Domain Adaptation (SF-UniDA) extends Universal Domain Adaptation by removing access to source data at adaptation time while still handling label-set mismatches between domains. Despite growing interest in this setting for image data, no benchmark exists for time series, which are more challenging. We present the first SF-UniDA benchmark on time series. In addition, we provide the first study of pretrained foundation models as feature extractors for time series domain adaptation. In this context, we identify a critical and previously underexplored limitation of all existing SF-UniDA methods: the inference threshold for unknown-sample rejection is highly sensitive. We address this by proposing a plug-in auto-thresholding module that can be integrated into any SF-UniDA method. Experiments on three well-known time series datasets confirm the suitability of this module. They also highlight that foundation models do not systematically outperform classical backbones and that SF-UniDA tailored for time series is yet to be developed.
cs.LG / 179 / 2609.39813
Backward-State Policy Is Part of the Learning Algorithm
Abstract
Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is part of the learning algorithm, and that neither check shows whether it is right. Copy accuracy does not decide the outcome: in three pairs of 390M runs with an emulated FP8 backward, training fails when attention's backward reuses the forward's rounded output and succeeds with a new rounding from the same distribution. Even the most accurate copy, the original itself, can be wrong by our reference: the gradient of the forward pass as it actually ran, with gradients passed through rounding unchanged. For example, a normalization output stored in low precision feeds two gradients: the gain's gradient needs the original, but the next layer's weight gradient needs the rounded value that layer multiplied. Final loss, the other check, does not rule out the error of reading the original for both: it persists in models trained with such a store, while planned loss comparisons stay within a margin fixed in advance. We therefore derive from this reference which value each use must read, or which substitute gives the same gradient on average with the forward held fixed, and check these per-use requirements on single operators, without training. In three tests using PyTorch and Transformer Engine, the requirements predicted beforehand whether reuse changes what the backward computes on average relative to an independent copy, and every prediction held. Backward-state policy is thus part of the learning algorithm: it should be specified and checked use by use, not settled by copy accuracy and final loss.
cs.LG / 180 / 2609.39816
Beyond Accuracy: Prefix-Invariant Realizations of Low-Precision Fast Matrix Multiplication
Abstract
Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on its allowed prefix. On Qwen2.5-14B-Instruct, two fast FP8 realizations repaired to ordinary-looking accuracy still change the answers chosen by likelihood on 5.83% and 10.00% of 240 OpenBookQA items when only the text after the allowed prefix is replaced with the bf16 model's own greedy continuation. Both row-local controls, the bf16 model and a deployed FP8 matrix multiplication kernel, change none. Accuracy thus does not certify prefix invariance, and the stability criteria we analyze cannot tell realizations apart: across all 512 sign variants of two-level Strassen they stay constant while teacher-forced perplexities span a 772.4$\times$ range on the same model. We therefore construct certified realizations of two-level Strassen on bounded integer codes that quantize token rows independently, then mix and cancel exactly before rescaling, using 49 block multiplications instead of 64. Our certificate guarantees bitwise equality to a prescribed row-local classical int8 operator at the same quantization specification, so every certified realization inherits its prefix invariance. Certification thus turns realization choice into a pure cost decision: which certified realization runs can no longer change a single scored likelihood.
cs.LG / 181 / 2609.39818
Should I stay or should I show? Learning to selectively disclose information
Abstract
In many high-stakes settings, human decision-makers can acquire support information before making a decision. However, acquiring information is costly, and disclosure may fail to improve human decisions or may even impair them. We tackle this problem by studying selective disclosure, i.e., the problem of learning when to reveal support information to a human decision-maker under a budget constraint. We first show that the optimal policy is a threshold rule on the Value of Information (VoI), i.e., the expected reduction in human decision risk induced by disclosure. Since VoI is unknown in practice, we estimate the regime-specific human risks and bound the possible degradation of the resulting plug-in policy relative to lack of disclosure, as well as its regret relative to the optimal policy. Experiments on benchmark datasets show that selective disclosure outperforms both no disclosure and full disclosure, regardless of whether the support information is beneficial or harmful. Two user studies show that human-AI team performance can improve when disclosure is led by our learned policy and not human-selected, although this advantage varies across tasks. A counterfactual benchmark, which replaces participants' predictions with a machine-learning prediction when disclosure occurs, suggests that these differences might depend on lower adherence to advice when the information is automatically provided rather than self-requested.
cs.LG / 182 / 2609.39833
RainAtlas: A Multi-Continental Dataset for Precipitation Downscaling
Abstract
Extreme rainfall events are increasing in intensity and frequency as climate change accelerates. While kilometer-scale precipitation forecasts are critical for supporting local decision-making, the limited availability of high-resolution precipitation observations hinders their accuracy, especially in under-resourced regions. Machine learning models are widely used to downscale precipitation data to km-scale, but their application to unseen geographies presents challenges. First, processing raw high-resolution precipitation datasets across regions requires significant engineering and domain expertise. Second, generalization across regions remains difficult. To help overcome these barriers, we release RainAtlas, a large-scale, ML-ready and multi-continental dataset for precipitation downscaling. Covering three continents, RainAtlas harmonizes heterogeneous hourly km-scale observations to a common 2-km grid. Each regional partition contains around 210,000 aligned low- and high-resolution precipitation pairs, respectively from ERA5 reanalysis and direct observations. We benchmark state-of-the-art ML-based downscaling models across RainAtlas using a wide range of metrics. Our evaluation reveals substantial variance in out-of-domain generalization depending on the training regions. This underscores the need for cross-regional, multi-source km-scale evaluation, establishing RainAtlas as a well-positioned benchmark for precipitation downscaling research.
cs.LG / 183 / 2609.39837
Fast Regularized Policy Mirror Descent with One-Step TD Updates
Abstract
Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of $ε$ after $\widetilde{O}(1/((1-γ)^5 \widetildeσ_b ε))$ transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage $\widetildeσ_b$. In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.
cs.LG / 184 / 2609.39839
Dynamic LoRA-Experts and Prototype-Ensemble Matching for Class-Incremental Learning
Abstract
Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. Parameter-efficient fine-tuning with pre-trained models reduces parameter overhead but can suffer from cumulative interference and suboptimal alignment between inference samples and specialized modules. We propose Dynamic LoRA-Experts and Prototype-Ensemble Matching (DLEPEM), a two-stage rehearsal-free framework. DLEPEM allocates a task-specific LoRA-Expert for each incremental task to reduce cross-task interference, then combines frozen pre-trained-model prototypes with task-adaptive LoRA-Expert prototypes for reliable task-level discrimination. Experiments on standard CIL and Few-Shot CIL benchmarks demonstrate strong performance under the evaluated protocols.
cs.LG / 185 / 2609.39848
Predicting Multi-View Rashomon Representation: Can We Learn Where Models Disagree?
Abstract
Foundation models are increasingly adopted across a wide range of applications, often serving as core blocks within AI systems. Yet different foundation models may encode the same input from multiple different views, leading to substantial representation disagreement, which we term Rashomon Representation. Such disagreement often signals inputs that a given model encodes in a way inconsistent with other models, offering a valuable yet underexplored signal for input reliability estimation. While prior work has largely focused on measuring disagreement across multiple models with a representation set, we instead focus on predicting disagreement from a single representation. We hypothesize that this disagreement follows some consistent, input-dependent patterns rather than occurring at random. To test this, we quantify disagreement by comparing each sample's nearest neighbors across different models' representation spaces, then train a lightweight predictor that estimates disagreement from a single model's representation. At inference time, given a new input, the predictor uses that input's representation to tell whether it aligns with or diverges from those of other models. Extensive experiments across diverse foundation models and datasets show that representational disagreement is indeed input-dependent, predictable, and generalizable, enabling efficient reliability estimation of foundation models.
cs.LG / 186 / 2609.39855
Dimension-Free Rank Lifting from Random Hyperplane Arrangements
Abstract
We study the width required for a randomly initialized hidden layer of a neural network to achieve rank lifting. Namely, given a dataset $X \in \mathbb{R}^{m \times d}$ of $m$, $d$-dimensional input vectors separated by an angle of at least $θ$, we consider the random feature matrix $σ(XR)$, where $R$ is standard Gaussian. For positively homogeneous nonpolynomial activations, which include sign, Heaviside, ReLU, and ReLU powers among others, we prove that $$n \gtrsim \frac{1}θ\max\left\{m,\log\left(\frac{1}δ\right)\right\}$$ neurons suffice for $σ(XR)$ to have full row rank $m$ with probability at least $1-δ$. This dimension-free bound exponentially improves the previous general-dimensional guarantee for sign features (Drago et al., 2026) and is essentially tight. The proof shows that one random feature column escapes every proper subspace of $\mathbb{R}^m$ with probability $Ω(θ)$, using a coupling of nearby Gaussian directions and a local crossing of the induced hyperplane arrangement. We also study stable rank lifting, where the goal is to establish a quantitative analogue of exact rank lifting, i.e., a lower bound on the smallest eigenvalue of the empirical feature Gram matrix in high-probability. Our analysis unifies and generalizes stable rank guarantees for all $q$-homogeneous non-polynomial activations following prior work in Panigrahi et al. (2020) and Song (2026). In particular, we combine a diagonally dominant Taylor tail of the population kernel with truncation and matrix concentration, to show that for positively homogeneous nonpolynomial activations, stable rank lifting is achieved at width $$n \gtrsim C^q \frac{m}{θ^{2q+1}} \log^{2q+\frac{1}{2}}\left(\frac{m}θ\right) \log\left(\frac{m}δ\right),$$ where $q$ is the degree of the activation and $C > 0$ is some universal constant.
cs.LG / 187 / 2609.39866
Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing
Abstract
Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.
cs.LG / 188 / 2609.39877
Algorithmic Recourse Under Competition
Abstract
Algorithmic recourse provides individuals who have received undesirable outcomes from machine learning models with suggestions for minimum-cost improvements to achieve the desired outcome. A central assumption when computing recourse is that the decision rule remains fixed throughout the recourse implementation phase. We challenge this assumption in settings where individuals compete for limited resources. In such settings, widespread recourse implementation can change the acceptance threshold even when the scoring model that is used to evaluate individuals remains the same. This change in acceptance threshold can, in turn, invalidate the original recourse recommendations (i.e., following the recourse may not lead to the desired outcome). To address this problem, we introduce a framework called recourse under competition that jointly optimizes for recommendation recipients and the recommended score target they need to satisfy to balance the recourse cost and post-shift validity among initially rejected individuals. We develop an algorithm based on the Implicit Function Theorem and empirically analyze its performance. Experiments on synthetic and real datasets show that personalized score targets can achieve higher validity, albeit at a higher cost. In contrast, common score targets generally offer favorable cost-validity trade-offs for lower to medium validity values.
cs.LG / 189 / 2609.39888
Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds
Abstract
Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls. For each transition it infers a control and recomputes the state through a completion model of known physics plus a learned residual. It then corrects that control by gradient-based inequality reduction, so inequality satisfaction is best-effort within an iteration budget. Since every correction iterate re-enters the completion model, the returned state is dynamically consistent by construction relative to that model and the supplied previous-state anchor. MaDE drives dynamics residuals to essentially zero on fully specified simulated systems, and on an underspecified system leaves a smaller true-dynamics residual than the baselines. Designed to attach to arbitrary predictors, the frozen operator is evaluated downstream of recurrent, structured state-space, and transformer predictors. On recorded vehicle trajectories the one-step residual against a kinematic bicycle model is 0.0071 to 0.0072 for MaDE and 0.1703 to 0.1714 for raw predictors. MaDE raises average displacement error by a factor of 1.57 to 1.83.
cs.LG / 190 / 2609.39892
Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does
Abstract
Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises a basic question: is each loop actually repeating the same computation, and if not, what routes the shared parameters to different operations? We study this question using graph walks as a test case. In the model's native trajectories, decoded predictions can advance by different numbers of graph steps or remain at a reached target, showing that recurrent progress need not follow a fixed one-loop-one-step pattern. We then show that a frozen loop can be steered toward different transitions by modifying its entering hidden state: a learned linear layer $J$ selects the desired transition without changing the shared Transformer layers. To test how this steering works, we use activation patching and find that attention patterns can recover its effects and switch the selected transition. Across five matched pairs of graph models, changing intermediate supervision during backbone training changes which transitions $J$ can induce. This suggests that $J$ selects computations learned by the backbone rather than creating new algorithms. Together, these results show that the hidden state can control shared computation, with attention routing as a causal pathway.
cs.LG / 191 / 2609.39901
Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?
Abstract
Goal-conditioned reinforcement learning (GCRL) relies heavily on how target goals are represented to the policy. While recent methods encode goals via temporal distance, occupancy, or controllability, it remains unclear how much downstream performance actually depends on representation quality. We study this in offline GCRL by constructing an exact temporal-distance goal representation in deterministic mazes. We then systematically corrupt its geometric quality while keeping the downstream learner fixed. Across OGBench navigation tasks and two algorithms, large changes in goal-representation quality produce almost no change in performance. However, applying the same interventions to the agent's current state more than doubles success, revealing the state pathway as the true bottleneck. Building on this insight, we show that simple random Fourier positional encodings substantially improve performance on the hardest navigation tasks without map information or objective modifications. Overall, our findings suggest that in state-based offline navigation, improving how the agent's current state is represented matters far more than refining the goal representation. Code will be released soon.
cs.LG / 192 / 2609.39911
Patient-Centered Treatment Planning for Chronic Multimorbidity: A Hierarchical Reinforcement Learning Framework for Preference Modeling
Abstract
Patient preference, defined as a patient's demonstrated willingness and capacity to adhere to clinical recommendations, is a primary determinant of therapeutic effect yet remains structurally absent from existing computational treatment planning models. We address this gap by presenting patient-centered factored-action hierarchical option-critic (FAHOC), a hierarchical reinforcement learning (HRL) framework that jointly learns high-level options corresponding to therapeutic strategies and factored intra-option policies that decompose the joint action space into disease- and intervention-specific subcomponents, while imposing a cooperation-aware action masking mechanism. This enables structured exploration, improved credit assignment across hierarchy levels, and more interpretable decision pathways, while enforcing patients' preferences. Formal guarantees establish that cooperative patients achieve higher optimal expected health outcomes than non-cooperative patients, and that the factored Q-function approximation error is provably bounded. The framework is evaluated using longitudinal data collected from approximately 50,000 comorbid hypertension and type 2 diabetes mellitus patients from five hospitals in the Southeast U.S. FAHOC achieves a quality-adjusted life year expectancy equivalent improvement of 0.669 (vs -0.133 observed clinician practice), correctly identifies cooperative patients in 95.9% of cases and never violates a patient's preference in held-out test, demonstrating that HRL with explicit preference constraints can support preference-consistent, clinically safe decision-making in multimorbidity management.
cs.LG / 193 / 2609.39912
TRACE: Trajectory Selection for Parallel Scaling of Search Agents
Abstract
Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence), a lightweight learned selector that ranks completed trajectories using the search evidence behind their answers. TRACE preserves individual query and evidence occurrences, connects rollouts through shared content or document identity, and propagates information across these relations. Each candidate answer then reads the updated states of its own trajectory, preserving retrieval provenance while incorporating evidence from related rollouts. Trained with answer-level supervision over frozen text embeddings, TRACE returns an existing answer without additional search or autoregressive aggregation. One selector per search setting transfers across rollout policies and agent backbones without agent-specific fine-tuning, improving over voting across six WebQA policies and six long-horizon dataset-backbone combinations at $K=16$. On Qwen2.5-14B Base/SFT WebQA pools, TRACE achieves 45.2/49.2% EM, compared with 43.9/48.0% for the strongest Qwen3-32B generative aggregators. On long-horizon FRAMES, GAIA, and BrowseComp, it reaches 78.6% average accuracy, exceeding majority voting by 3.1 percentage points. On Base WebQA pools, TRACE with only 8 rollouts comes within 0.4 points of majority voting over 64. TRACE also achieves at least $10\times$ higher processing throughput than SolAgg, SummAgg, and AggAgent across all seven WebQA benchmarks. These results show that reusing cross-rollout search evidence provides an effective and efficient alternative to heavyweight generative aggregation for parallel search. Code is available at https://github.com/Jaasssoooonnnnn/TRACE.
cs.LG / 194 / 2609.39929
RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
Abstract
Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-frequency components support positional sensitivity while potentially disrupting semantic stability. We also derive a theoretical context-length bound beyond which, under specified conditions, a fixed attention-score comparison cannot jointly avoid semantic reversal and positional insensitivity. Guided by our fresh theoretical insights, we introduce RoPE Profiler, a lightweight, plug-and-play diagnostic toolkit that augments existing evaluations with zero additional forward passes by reusing cached query and key activations. Reusing activations collected during evaluation, the toolkit incurs little overhead. It supplements standard benchmark scores with two diagnostic scores that reveal semantic and positional weaknesses and help users prioritize which aspect to address. Crucially, our evaluations across 49 long-context task settings reveal a distinct pattern where reasoning tasks predominantly suffer from semantic reversal, whereas retrieval tasks are primarily vulnerable to positional insensitivity. Guided by our theory and diagnostic profiles, targeted high-frequency rescaling achieves immediate gains without additional training, improving task accuracy by up to 20 percentage points on Qwen3-8B and 25 percentage points on Llama-3.1-8B-Instruct.
cs.LG / 195 / 2609.39934
Reliability-Aware Checkpoint Selection for Domain Generalization
Abstract
Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with near-optimal source accuracy can improve mean target probability quality with small observed changes in mean target accuracy. We study accuracy-constrained reliability selection (AC), which retains checkpoints within a tolerance of the best source-validation accuracy and ranks them by source reliability. Our reference rule aggregates within-set normalized negative log-likelihood (NLL) and class-wise calibration error (CwECE) using $D_\infty$. AC uses no target data and requires neither additional training nor weight averaging. We evaluate five domain generalization training algorithms on three benchmarks, using PACS to develop the objectives and a 0.5-percentage-point tolerance. In exploratory aggregation comparisons on 360 OfficeHome and TerraIncognita runs, the reference rule reduces mean target soft-bin squared-gap ECE and CwECE by 0.240% and 0.182%, respectively, and NLL by 0.030 relative to Source-Acc. Mean target accuracy changes by +0.213 percentage points. These results identify opportunities for reliability-aware reselection, while the additional benefit of joint over single-objective ranking remains unresolved.
cs.LG / 196 / 2609.39967
What Limits Recursive Reasoning Models: Optimization, Architecture and Test-Time Scaling
Abstract
Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call on narrow algorithmic subproblems. However, existing models such as HRM, TRM and URM differ in architecture, gradient propagation and training procedure simultaneously. This makes it hard to tell what drives their performance, and their optimization is still poorly understood and often unstable. In this work we address both of these gaps. First, we study these questions under a unified experimental pipeline spanning six algorithmic domains. Individual controlled ablations are performed on representative domains, while the resulting recipe is evaluated across the full suite. The study reveals a surprisingly simple recipe for stable and generalizable recursive reasoning: an intermediate gradient horizon, large physical batches and controlled updates of the recurrent state. An explicit hierarchical architecture is not needed. Second, we combine these findings into a stable 13.6M-parameter model that achieves the strongest overall performance among the evaluated recursive baselines, with particularly large gains on out-of-distribution generalization. It raises Arithmetic OOD accuracy to 71.2%, from 36.2% for the strongest baseline, while reaching 98.41% on Sudoku and 59.5% pass@2 on ARC-AGI-1. Our results show that, within the recursive architectures studied here, performance depends strongly on how recurrence is optimized and stabilized. More broadly, it shows how AI systems can be improved by optimizing their components one at a time.
cs.LG / 197 / 2609.40003
DashVMC: Real-Time Discrete World Model Control in Geometry Dash
Abstract
World-model agents are usually evaluated in simulators that can wait for the policy; live games impose the opposite constraint, requiring capture, prediction, and action before the next frame. We present DashVMC, which learns a compact, action-conditioned world model from approximately two hours of recorded Geometry Dash gameplay. To test whether the learned dynamics are actionable, a controller is initialized by behavioural cloning (BC) and refined with Proximal Policy Optimization (PPO) entirely in frozen-model rollouts, without further interaction with the live game. Across three controller seeds, the refined policies survive longer than their BC initializations on all three official levels and a held-out community layout. At deployment, the baseline skips visual generation and sustains a 60-Hz capture-to-action loop on a consumer GPU. Action-conditioned continuations and rollout diagnostics show that the model remains useful for control despite imperfect long-horizon fidelity.
cs.LG / 198 / 2609.40034
Efficient Active Auditing of Multi-Group Fairness with Bias Probes
Abstract
Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing essential. Existing auditing approaches for black-box models either rely on model reconstruction --exposing systems to extraction attacks-- or directly estimate fairness metrics, offering limited insight into which regions of the data distribution drive bias. More fundamentally, property-specific auditing --aimed at extracting only targeted fairness information without reconstructing the model-- remains poorly understood. In this work, we introduce the bias probe framework, which enables targeted and adaptive querying to reveal bias structure while preserving model confidentiality. Building on this framework, we propose ALeBi, an active auditor that learns such probes to efficiently estimate multi-group fairness metrics. We establish novel sample complexity guarantees governed by a property-specific complexity measure, resolving a previously posed open question, and extend our analysis to adversarial settings where the model owner may strategically obscure bias. Our results uncover a fundamental trade-off between model confidentiality and reliable auditing, and show that property-specific probing enables both accurate estimation and interpretable identification of high and low-bias regions. Extensive experiments support our theoretical findings and demonstrate the practical effectiveness of our approach.
cs.LG / 199 / 2609.40047
Gromov-Wasserstein Distillation for Inductive Multi-View Embedding
Abstract
Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing no explicit mapping for unseen samples. We introduce an inductive framework based on barycentric distillation. A GW-MDS teacher learns a latent support and an optimal transport plan from the training data, and barycentric projection converts the resulting coupling into sample-aligned targets. A neural student then learns an explicit out-of-sample mapping, avoiding additional relational-matrix construction and GW optimization at inference. We formulate the approach for single-view data and extend it to Mean-GWMDS and Multi-GWMDS teachers through consensus and selected-projection targets learned by a multi-view student with view-specific encoders. We also investigate a direct neural baseline trained solely with a GW objective. Experiments on synthetic and real-world data using Euclidean, geodesic, and cosine relations show that the distilled models preserve the teacher geometry on unseen samples and consistently outperform direct neural GW training in sample-indexed relational preservation. These results establish barycentric projection as an effective bridge between transductive GW embeddings and inductive neural mappings.
cs.LG / 200 / 2609.40063
LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models
Abstract
Low-Rank Adaptive Residual Connections (LARC) give a frozen model a compact numerical state that can learn from feedback. The map $h+BAh$ adds a low-rank correction to a hidden representation. A slow state $ρ$ learns starting factors across tasks; a private fast state $Φ$ copies them, changes with feedback, and resets to the trained initialization. This report specifies an input-side realization of the numerical policy carrier in Memory-Mediated Learning Architecture and examines its factor-space dynamics and learning lifetime. We study a rank-4 input residual with 12,288 trainable parameters on a frozen MiniCPM5-1B-SFT substrate. In a four-candidate program-selection task, two feedback-gradient steps reduce expected query execution error by 24.65 and 36.65 percentage points relative to resetting to the respective trained static and post-adaptation initializations. These development results cover 16 parameter groups and three paired training seeds. A direct support-loss selection rule is much more accurate, reaching 0.78125% error. In a repository-balanced chronological replay of public continuous-integration jobs, retaining online updates raises half-Brier loss from 0.1274 to 0.1808. A fixed follow-up intervention records same-batch non-descent and inconsistent future benefit from shrinking updates. Together, the algebra and measurements distinguish residual capacity, adaptation relative to a starting point, and usefulness on later decisions.
cs.LG / 201 / 2609.40070
Inference Auctions
Abstract
When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing latency, and we develop fast algorithms for implementing prices that incentivize truthful bidding. We also design an autobidding agent for our inference auction, where users specify an inference budget and the autobidder dynamically adjusts its bids over time to maximize user utility subject to the budget constraint. Experiments validate the practicality of our auction: it increases system welfare while maintaining the cache utilization and latency advantages of SGLang, a state-of-the-art inference serving framework.
cs.LG / 202 / 2609.40075
Accelerated Algorithm for Sparse Regularized Partial Optimal Transport
Abstract
Partial Optimal Transport (POT) extends the classical optimal transport problem by relaxing the strict mass conservation constraint, enabling its use in a wide range of real-world applications. In many of these settings, sparse transport plans are preferred for their interpretability and computational benefits. While smooth and strongly convex regularizers - such as quadratic or elastic net - have been vastly used in various machine learning applications to induce sparsity and accelerate computation, they have received less algorithmic attention compared to entropic approaches for computational POT. In this paper, we propose a new optimization framework that leverages these regularizers through a penalty-based reformulation, enabling efficient gradient-based updates while preserving the structure of the original problem. Our method accommodates a broad class of regularizers that promote structured and sparse transport plans. Building on this formulation, we design an accelerated first-order algorithm that alternates between smooth updates and simple projection steps. Through empirical benchmarks on color transfer, domain adaptation, and point cloud registration, our approach consistently outperforms established baselines - achieving lower transport cost, higher sparsity, and faster convergence - making it a practical and scalable solution for modern transport problems.
cs.LG / 203 / 2609.40089
Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining
Abstract
Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate training independently of the model's actual retention needs. We introduce Replay on Demand (RoD), which instead derives the replay allocation from the model's learning dynamics. RoD jointly prioritizes adaptation samples by their remaining learning potential and replay samples by their observed forgetting. Their competition for a shared training budget yields an online curriculum that determines what to train on at each step. Across models, scales, and adaptation domains, RoD reaches or improves upon the adaptation-forgetting frontier of tuned fixed-replay baselines and model merging without prescribing a replay allocation in advance. Replay concentrates on sources that are more vulnerable to forgetting and dynamically increases and redistributes as forgetting emerges during training. Together, our results show that replay can be allocated online from the model's evolving state, targeting what is needed, when it is needed.
cs.LG / 204 / 2609.40117
Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting
Abstract
Time-series forecasting models achieve strong benchmark performance but exhibit severe systematic bias in industrial deployments. This train--deploy gap is conventionally attributed to temporal-structural errors or distribution shifts. We characterize a complementary source that these explanations overlook: canonical losses embed fixed statistical priors, while industrial demand mixes benign and pathological regimes---zero-inflation, skewness, high variability---in which these priors are systematically violated. The induced bias persists even under perfect temporal modeling, remains in a distributional-shape component that normalization cannot remove, and creates an aggregation trade-off invisible to aggregate metrics. We turn these observations into an evaluation toolkit centered on the Regime-wise Relative Bias Vector (RBV): a metric-agnostic, regime-decomposed diagnostic that audits how pooled training allocates systematic mismatch across pathological subpopulations. A controlled attribution analysis decomposes RBV into a model-independent intrinsic floor, set by each loss's estimand, and an excess component attributable to training, tracing observed bias to the loss rather than the model. A large-scale study---13 loss objectives, 3 seeds, 60,000+ series spanning RetailShiftBench and M5, with random-split controls---shows that regime-aware diagnosis separates optimization-type from bias-type failure, and that regime-aware training resolves the pooling-induced bias that capacity scaling cannot, for mean-type losses. A formal structural observation, that risk under evaluation-distribution contamination is affine in the pathology mixture weight, grounds these findings. Our work complements model ranking with mechanism-grounded, regime-oriented evaluation.
cs.LG / 205 / 2609.40120
Scalable Cox Regression via Grouped Risk Sets and Sharper LogSumExp Rates
Abstract
Motivated by the computational challenges of large-scale Cox regression, we study stochastic minimization of LogSumExp objectives over large sets. Mini-batch normalizer estimates generally yield biased gradients. We instead use a softplus surrogate that introduces one auxiliary scalar per normalizer and admits unbiased single-sample gradients. For smooth convex LogSumExp objectives, we prove an $O(T^{-1/2})$ averaged objective bound, improving the previous $T^{-1/4}$ analysis. With a strongly convex regularizer on the original variable, we also obtain a last-iterate squared-error rate of $\widetilde{O}(T^{-1})$ without strong convexity in the auxiliary variables. For Cox regression, the normalizers are defined over nested risk sets. We exploit this structure by grouping neighboring failures and sharing one auxiliary variable per group. The resulting compressed objective admits uniform score and curvature bounds that control the errors from grouping and softplus approximation. Together with the general optimization result, these bounds give a mean-square rate of $T^{-4/5}$, up to logarithmic factors, relative to the full Cox solution. The compressed estimator also matches the full estimator's asymptotic distribution. Experiments on synthetic and real survival datasets with slowly decreasing risk sets show a favorable performance relative to stochastic baselines.
cs.LG / 206 / 2609.40127
Learning Functional Subspaces for Neural Network Compression
Abstract
Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standard hardware. Existing methods, however, choose the subspace to remove from each weight matrix with local closed-form criteria: activation energy, layer-wise reconstruction error, or a quadratic approximation of the loss. These criteria ignore how errors propagate through the network, so at high compression the errors compound with depth and performance collapses. We introduce Learnable Subspace Projections (LSP), which instead learns the subspaces to discard end-to-end. Each linear layer, or tied group of layers that read the same activations, is assigned an orthogonal projector. All projectors are optimized jointly against a global objective--the KL divergence to the dense model's output distribution or the model's original training loss--while the pretrained weights remain frozen. Projectors are initialized from a whitened SVD truncation, and ranks are allocated by the output KL each projector induces per parameter saved. After training, the projectors merge into standard low-rank factors, with each tied group sharing one factor. In attention, this also lets the model cache one narrow latent in place of full keys and values. Across LLMs (OPT-125M/1.3B, Qwen3-4B, Llama-2-7B) and ViT-B/16, LSP outperforms baselines, and its advantage widens as compression increases. At -70% compression, LSP brings Llama-2-7B to 10.9 WikiText-2 perplexity and 42.2% mean zero-shot accuracy, versus 13.3 and 36.0% for the strongest baseline. The factorized model decodes up to 1.6x faster than the dense model at small batch sizes, and aching the shared latent shrinks the combined memory of weights and KV cache by 13.5x at a 128k-token context, versus at most 6.5x for untied baseline factorizations.
cs.LG / 207 / 2609.40131
Prototype-Rule Neurosymbolic Regularization for Rank-Constrained Tensor Neural Networks under Label Scarcity
Abstract
Rank-constrained tensor neural networks reduce the parameterization of high-order inputs, but they do not explicitly constrain class geometry in the learned representation. This study investigates whether a differentiable prototype-rule can provide a complementary inductive bias for Rank-R tensor learning under limited supervision. The proposed framework augments the Rank-R objective with prototype-based regularization and optionally fuses prototype evidence with neural logits at inference. Four hyperspectral benchmarks are evaluated with four Rank-R configurations under both seven-fold stratification and spatially separated folds that mitigate leakage; a separate spatial study varies the class support budget from 2 to 20 samples. Under spatial evaluation, full neurosymbolic inference changes Macro-F1 score by +8.82 percentage points on Botswana, +5.49 on Indian Pines, +1.59 on Pavia University, and -0.62 on Salinas. Most of the benefit arises from training-time regularization, whereas inference fusion is small and dataset dependent.
cs.LG / 208 / 2609.40137
Game-Guided Skill Discovery through Self-Play for Playable Agent Control
Abstract
We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, these skills should be semantically distinct, interpretable, and expressive; properties that existing unsupervised skill-discovery methods often fail to achieve simultaneously. GGSD achieves these desiderata by grounding skill discovery in competitive gameplay. A hierarchical agent competes against its past selves, with a high-level policy selecting from a small discrete skill set and a skill-conditioned low-level policy learning the corresponding behaviors. After training, a human can replace the high-level policy and directly control the agent through the same discrete skills. Despite the small number of high-level actions, skill transitions give rise to emergent combo behaviors, expanding expressivity beyond individual primitives. Across Ant, Franka-arm, and Unitree G1 environments, we show that GGSD produces human-playable skills that humans can compose to solve unseen tasks, such as Maze and CubePush, without additional training. An interactive demo is available at https://ggsd-demo.github.io.
cs.LG / 209 / 2609.40143
From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer
Abstract
Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity. Yet most model-based nucleic-acid designers optimize fixed-length sequences through substitutions; they do not ask which bases of an existing functional element can be removed while retaining predicted activity. We define the task of sequence slimming as selecting an exact-length, order-preserving subsequence while retaining activity. Modeled on the design benchmark NucleoBench, we propose a quantitative evaluation for slimming that balances sequence reduction with maintaining function. Each slimmer must return both the subsequence and its source indices, which can be used to verify that the slimmer obeyed task requirements. To our knowledge, this is the first dedicated benchmark of this deletion-only problem. The coding agent Empirical Research Assistant (ERA) then searched over executable designer programs. ERA received the task prompt and a successful substitution-only designer GrAdaBeam as a starting program, and it modified the designer to produce GRADASLIM. We report held-out evaluations for five transcription-factor binding targets, comparing random, greedy, and ERA-guided slimming at 400 and 100 bp. ERA has the highest mean in 9/10 settings. Paired bootstrap intervals for ERA minus greedy are above zero in all five 400-bp settings, below zero in one 100-bp setting, and overlap zero in the remaining four.
cs.LG / 210 / 2609.40147
Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy
Abstract
We establish an exponential iteration lower bound in the number of states for Howard's policy iteration on deterministic discounted Markov decision processes, with at most two actions per state. This rules out strong polynomiality of Howard's policy iteration when the discount factor is part of the input and yields an exponential separation from the simplex method with Dantzig's pivoting rule, which is proved to be strongly polynomial on this class. Even when each reward is restricted to logarithmic bit length, we obtain a stretched-exponential iteration lower bound. The gap between Howard's decentralized and simultaneous selfish improvements and Dantzig's coordinated selection of a single action with the largest gain across all states reveals a ``price'' of algorithmic anarchy.
cs.LG / 211 / 2609.40148
From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes
Abstract
Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error injections. We prove that either component follows a power law if and only if its cumulative weighted spectral mass has the corresponding low-spectrum scaling; individual eigenvalues and target coefficients need not obey coordinatewise power laws. Under a joint schedule, intrinsic time $T_t=\sum_{s<t}η_s$ controls optimization progress, while $r_t=B_t/η_t$ controls noise injection. Their interaction yields sharp conditions under which a schedule preserves, changes, or destroys the clean power law, together with a memory ceiling on noise reduction. The power-law random-feature model realizes this mechanism in $3+3(+2)$ propagation regimes with phase-dependent compute rates. Controlled nanoGPT experiments show that (1) learning-rate and batch-size schedules with matched $B/η$ paths are nearly equivalent in intrinsic time, (2) a forcing-memory surrogate accurately predicts loss across schedules, and (3) its fitted exponents across real-world datasets identify the regime of LLMs in $3+3(+2)$ map.
cs.LG / 212 / 2609.40149
Role-Adaptive Policy Optimization for Offline Reinforcement Learning
Abstract
Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.
cs.LG / 213 / 2609.40170
MANET-GNN: Learned Decentralized Optimization of Power Allocation in Multi-Channel MANETs
Abstract
MANETs enable flexible infrastructure-less wireless connectivity in dynamic and resource-constrained environments. As modern MANETs exploit multiple frequency channels and support heterogeneous traffic patterns, decentralized transmit-power allocation becomes increasingly challenging. We develop a unified learned optimization framework for decentralized power allocation in dynamic multi-hop, multi-channel MANETs. We formulate a constrained end-to-end throughput maximization problem covering unicast, multicast, multicommodity, convergecast, and many-to-many communication. Although centralized and non-convex, this problem serves as an unsupervised training objective for MANET-GNN, a message-passing GNN that operates as a distributed learned optimizer. MANET-GNN uses only local, possibly noisy, CSI and a prescribed number of neighbor message exchanges, enabling low-latency decentralized inference while generalizing across topologies and network sizes. Numerical results show that MANET-GNN achieves centralized-competitive performance across communication frameworks, remains robust to channel uncertainty, and scales effectively across MANET configurations.
cs.LG / 214 / 2609.40190
Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves
Abstract
Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within $\pm1/32$ at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions, and an audit that revisits every question need not pay for it. We derive the minimax cost of certifying the whole curve, up to logarithmic factors. It has three parts: calibrating the tail of the score distribution, telling the questions apart, and within-question noise summed along the curve. At a single benchmark the last part sharpens to the variance of one answer's influence under the best allocation of answers to questions, which every valid audit pays and an audit that learns the allocation attains, up to a logarithm, as the precision grows. A paired audit built on an exponential inequality for two independent draws at the same question needs no pilot. On 185 held-out score pools it uses 0.74 times the answers of the cheapest competing certified audit at 64 budgets and 0.53 times at 1,024, and on a newly generated MMLU-Pro study it certified the curve with 79,133 answers, within 0.6% of what a cost law fitted beforehand predicted from the study's within-question variance. The same paths certify pass@$k$ and majority voting, and the bands extend to populations of questions and to answers that depend on earlier ones.
cs.LG / 215 / 2609.40193
Near-Linear Accuracy Bounds for Moreau--Yosida Unadjusted Langevin Sampling
Abstract
We establish near-linear accuracy bounds for the classical Moreau--Yosida unadjusted Langevin algorithm (MYULA). The target is $π\propto e^{-f-g}$, where $f\in C^2(\mathbb{R}^d)$ is $m$-strongly convex with Lipschitz gradient and $g$ is convex and globally Lipschitz. Under an explicit parameter-dependent step-size condition, we bound the invariant-measure bias relative to the Moreau-smoothed target by $\widetilde O(h)$, with only logarithmic dependence on the inverse smoothing parameter in the error coefficient. Combining this estimate with the Moreau approximation bias and Wasserstein contraction gives $\widetilde O(\varepsilon^{-1})$ iterations to make the $N$th-iterate law $μ_N$ satisfy $\sqrt m\,W_2(μ_N,π)\le\varepsilon$, for fixed model parameters and initialization. We bound the stationary error directly, without assuming third derivatives or a Lipschitz Hessian. Each iteration uses one gradient evaluation and one exact proximal evaluation. The key idea in our analysis is to convert a second-order stationary residual into a Wasserstein bound using a Poisson-based estimate.
cs.LG / 216 / 2609.40221
PhantomEnvironments: Training LLM Agents in Fictional Worlds
Abstract
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
cs.LG / 217 / 2609.40265
OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and Reasoning
Abstract
Real-world time-series applications increasingly require models that can handle time series forecasting, context-conditioned prediction, and language-based temporal reasoning. Yet current time-series foundation models remain fragmented across these capabilities: numerical specialists often provide the strongest forecasts, while language-based models offer broader contextual understanding and analysis. A central challenge is to unify these heterogeneous capabilities without reducing their individual performance. We introduce OpenTSLM TeeMoE, a generalist time-series language model that can forecast directly from observed time series, reason over textual context and temporal patterns, and synthesize and refine predictions from external numerical forecasting specialists. We independently train three low-rank experts for forecast aggregation, native forecasting, and temporal analysis over a shared backbone. A learned LoRA mixture-of-experts controller then weights their frozen parameter updates for each request. Our proposed model achieves strong performance on widely used benchmarks for time series forecasting, context-conditioned prediction, and language-based temporal reasoning, ranking among the top three on GIFT-Eval by mean MASE rank, Context is Key by RCRPS, and TimeSeriesExam by accuracy.
cs.LG / 218 / 2609.40284
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Abstract
Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at https://cuaspeedrun.com.
cs.LG / 219 / 2609.40287
PMosFM: Preconditioned Manifold Matching for One-Step Physics-Constrained Generation
Abstract
Physics-constrained generative models aim to generate physical fields that match a target distribution and satisfy prescribed constraints. However, enforcing these constraints often increases sampling costs through iterative corrections or training costs through residual optimization and trajectory unrolling. To address this issue, we introduce \textbf{P}reconditioned \textbf{M}anifold \textbf{o}ne-\textbf{s}tep \textbf{F}low \textbf{M}atching (\textbf{PMosFM}), a preconditioned manifold matching framework for one-step physics-constrained generation. By encoding constraints in a manifold decoder, PMosFM learns transport in intrinsic coordinates without separate residual losses or terminal residual unrolling. A geometric preconditioner rescales coordinates using the decoder-induced metric, while a regularized covariance transform approximately whitens the interpolation-state inputs. A finite-interval objective couples velocity supervision with consistency between decoded endpoints in physical space. We show that exact parameterization removes residual-induced Gauss--Newton curvature, that geometric and covariance effects separate in a local conditioning bound, and that physical flow-map error bounds endpoint distributional error. Controlled ablations examine conditioning, and experiments evaluate optimizer-update time and memory footprint. At inference, PMosFM uses one neural transport evaluation followed by physical decoding. Experiments across benchmarks show lower training and sampling time than the multi-step baselines at comparable physical and distributional fidelity. Code and datasets will be released publicly.
cs.LG / 220 / 2609.40292
Disentangling Computation in Multi-Task Neural Networks with the Green's Operator
Abstract
How is computation organized and reused across tasks and time in a trained recurrent network? Most analyses emphasize the geometry of neural activity, dynamical motifs, or local perturbation growth. We instead study the network's global first-order perturbation response. The finite-horizon Green's operator maps perturbations at each source along a trajectory to their downstream state-space responses and therefore directly represents perturbation routing. Simple reductions of this operator provide task-to-task and time-to-time views of the same computation, while matrix-free products make these views accessible without constructing the full operator. In a flexible multitask recurrent network, task reductions reveal structured reuse of known computational motifs, while temporal reductions reveal causal pathways and how they emerge during training. Our main point is simple: the Green's operator provides a global response geometry for mapping the organization of learned dynamical computation.
cs.LG / 221 / 2609.40312
Compression Footprints as Security Signals for Model-Poisoning Defense in Federated Learning
Abstract
Lossy compression is widely used in Federated Learning (FL) but is generally treated as an error source, while conventional poisoning defenses inspect update geometry. In this work, we instead treat the compressor's response as a security signal: the input-dependent distortion and payload behavior induced by lossy compression can expose differences between honest and attack-generated updates. We introduce the concept of a \emph{compression footprint}: the low-dimensional collection of reconstruction, directional, sparsity, and payload statistics induced by a lossy compressor. We characterize sufficient conditions under which compression footprints separate honest and malicious updates, and operationalize our findings in the CRAFT (\emph{Compression-guided Robust Aggregation via Footprint Trust}) server-side robust aggregation method. Crucially, under a strict honest-majority assumption, CRAFT uses server-verifiable footprints, requires no client-side metadata nor knowledge of the number of malicious clients, and adds no communication beyond the compressed FL pipeline. Moreover, while CRAFT assumes a strict honest majority, it does not require the number of malicious clients to be known in advance. We observe that error-bounded lossy compressor (EBLC) footprints provide stronger separation than Top-K footprints and that footprint trust suppresses malicious influence. We evaluate CRAFT under IID client data with 36\% malicious participation across six standard model-poisoning attacks, three datasets, and six robust aggregation baselines, finding that CRAFT consistently achieves the best accuracy in 7 out of 18 settings and within 1.7 percentage points of the best in the others. Our results show that lossy compression can serve as both a communication mechanism and a security signal for robust aggregation in FL.
cs.LG / 222 / 2609.40316
Scaling Laws for Looped Mixture of Experts
Abstract
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.
cs.LG / 223 / 2609.40359
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
Abstract
We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows partially overlap, implicitly revealing the interval between words. Since these intervals indicate the duration of the words spoken, and different words tend to have different durations - for example, "the" is much shorter than "supercalifragilisticexpialidocious" - the neural network can improve its predictions of words without relying on the underlying brain activity. Consistent with this, the method reaches 22.0% balanced accuracy on synthetic signals containing no brain information, compared with 22.3% on real brain recordings. To prevent the network from learning this shortcut, we make a single, simple change. Instead of jointly encoding all windows in a sentence, we process each independently. As a result, the neural network achieves better performance by learning underlying word-specific information from brain recordings. This makes two existing strategies become much more effective than before. Both aggregating predictions from distinct neural responses to the same word and using a pretrained LLM as a linguistic prior now substantially improve results. On our perceived speech benchmark, this simple recipe (SimpleB2T) achieves a word error rate of 36.6% with five observations per word, approaching past invasive speech decoding performance, albeit under different conditions. The results in this work expose an important shortcut in brain-to-text decoding and show that removing it leads to a simple and considerably more effective strategy.
cs.LG / 224 / 2609.38926
PrecipJEPA: JEPA-Regularized Future-State Prediction with Motion-Source Rendering for Precipitation Nowcasting
Abstract
Long-term precipitation nowcasting requires modeling radar-echo evolution while preserving localized high-intensity structures. Recent radar-specific studies motivate location-aware prediction and separating echo displacement from intensity change. However existing encoders learn historical representations mainly from final forecast errors. We propose PrecipJEPA, which couples a structured forecasting path with an auxiliary path that enriches its encoder from observed radar history. In the forecasting path, an online encoder first converts the observations into spatiotemporal tokens. The Task-Driven Future-State Predictor (TFP) combines these tokens with a recent-dynamics summary and spatiotemporal queries to construct future radar states. The Parallel Motion-Source Renderer (PMSR) decodes these states into motion and source-sink fields that transform the latest observation into future frames. During joint training, the History-Masked JEPA (H-JEPA) operates on the auxiliary path to predict masked historical features from visible context, directly supervising the same online encoder from the observed sequence. Experiments on SEVIR and MeteoNet show that PrecipJEPA improves highest-threshold CSI by 118.6% and 35.1%, respectively, over the strongest baselines, while maintaining the highest mean CSI throughout the 3-hour forecast.
cs.LG / 225 / 2609.38616
Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance
Abstract
While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.
cs.LG / 226 / 2609.39151
Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey
Abstract
Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-$k$ nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model ($k=0$) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Increasing nonlinear innovation rank providing no measured benefits. This result is obtained on a fresh test split, which will be only opened after all models and analyses are frozen. The linear model requires fewer recurrent parameters and less computation than GRU, but performs comparably. These results suggest that, for this memory dependent control task, nonlinear representation learning around a simple linear memory mechanism can be sufficient, and that nonlinear recurrent dynamics are not necessarily required. These conclusions are limited to the simulated task, teacher, state dimension, and blackout horizon considered here, and to policies whose observation encoder and action head remain nonlinear.
cs.LG / 227 / 2609.39324
MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies
Abstract
Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared global action features may fail to establish timestep-specific correspondence between actions and local visual changes. To address this issue, we propose MotionWeave, a motion-centric future-dynamics framework for action-chunk prediction with two modules: the Action-Induced Motion Grounder (AIMG) and the Horizon Residual Composer (HRC). Specifically, AIMG conditions on action and proprioceptive representations to construct horizon-specific queries that localize interaction regions associated with each future action timestep from current visual tokens. HRC extracts differences between interaction representations at adjacent horizons, encodes them as temporal motion cues, and injects them into action tokens through a gated residual. During training, robot-arm masks rendered from future frames are used to construct KL-based motion-grounding supervision, while inference uses only the current observation. On six MetaWorld tasks, MotionWeave achieves a 75.3% average success rate, an absolute gain of 8.6% over π0 (66.7%), especially on sustained-interaction tasks. Our code is available at https://github.com/autu-mn/MotionWeave.
cs.LG / 228 / 2609.39558
General Performance Guarantee for Human Torque Estimation-Based Task-Agnostic Assistive Exoskeleton Control
Abstract
Accurate human torque estimation is crucial for enabling task-agnostic control in robotic exoskeleton systems. However, estimation errors may cause mismatches between the robot assistance and the human intention, degrading controllability and task performance. In this paper, we address this issue by formally defining matched assistance as scenarios in which the robot positively contributes to human movement. Based on this definition, we develop a theoretical framework to design the robot's desired interaction torque that guarantees a lower bound on the matched assistance probability. Importantly, the proposed guarantee holds over the entire torque distribution, including unseen data beyond the training tasks. This provides our method with strong reliability and generalization, both of which are critical for effective exoskeleton control. The proposed strategy is implemented on the ABLE upper-limb exoskeleton and evaluated in a multi-task setup. Experimental results validate the theoretical guarantees and demonstrate that the proposed strategy achieves effective general performance across several tasks, guaranteeing movement smoothness while reducing human physical effort.
cs.LG / 229 / 2609.39971
When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models
Abstract
Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires. We call this failure instruction-action binding. Instructions cue familiar trajectory families, and visual feedback adjusts their execution. Behavioral analyses of fine-tuned $π_{0.5}$ and GR00T-N1.7 policies reveal that failed rollouts often retain the source behavior or switch to another demonstrated task. These switches show that language is not simply ignored. Readouts and interventions connect these choices to task-conditioned internal states. Our analysis of the imitation objective shows how narrow conditional action support can leave grounded and instruction-keyed solutions indistinguishable on the demonstrations. This motivates Equivariant Counterfactual Training (ECT), which acts at two levels. ECT data supply valid demonstrations in which the same instruction requires different actions in distinguishable scenes, while the ECT loss trains each demonstration with its counterpart in the same update. In a controlled LIBERO-PRO comparison, full ECT raises $π_{0.5}$'s mean position-swap success from 36% to 59%. On CALVIN, where counterparts already occur in the original data, the ECT loss improves five-task completion without new demonstrations. On a real UR5e under a fixed demonstration budget, full ECT raises unseen-position success from 8% to 88%.
cs.LG / 230 / 2609.40245
STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
Abstract
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/socialnav-sub.
cs.LG / 231 / 2609.39162
Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization
Abstract
Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.
cs.LG / 232 / 2609.39679
SE-ADD: Self-Evolving Audio Deepfake Detection with Mistake-Driven Supervision
Abstract
Audio deepfake detection (ADD) must remain effective when new spoofing attacks emerge after deployment. Emerging audio language model (ALM)-based ADD methods are built on predefined supervision from ground-truth labels or verified forensic rationales. However, this paradigm overlooks an ALM's own mistakes, which indicate where targeted supervision is most needed. To this end, we first introduce evolving spoofing environments for ALM-based ADD, where a new attack becomes dominant while previously observed attacks persist. Motivated by the above learning-from-mistakes perspective, we further propose SE-ADD, a self-evolving framework that iteratively adapts an ALM via low-rank adaptation (LoRA) using mistake-driven supervision built from its verdicts and self-generated forensic cues. All training samples receive direct authenticity supervision, while misclassified ones receive additional cue-augmented supervision. As verdicts and cues are regenerated by the updated ALM, the resulting supervision evolves accordingly. Experiments on two ALMs demonstrate the effectiveness of SE-ADD in generalizing to unseen attacks, reducing the equal error rate (EER) from $36.72\%$ to $7.52\%$ for Qwen2-Audio and from $19.93\%$ to $3.97\%$ for MOSS-Audio.
cs.LG / 233 / 2609.39847
SEAR: Spoofing Evidence-Grounded Audio Reasoning Benchmark for Audio Language Models
Abstract
Audio language models (ALMs) are increasingly used for audio deepfake detection (ADD), yet existing benchmarks assess their verdicts or rationale plausibility without verifying the underlying acoustic evidence. To address this issue, we first introduce spoofing evidence-grounded audio reasoning (SEAR), a four-task AQA benchmark to evaluate ALM-based ADD through acoustic evidence identification and quantification, deepfake detection, and forensic rationale generation. We further propose a bona-fide-based acoustic evidence agent (BAEA), which equips a frozen ALM with controlled acoustic tools under \textsc{fixed} or \textsc{adaptive} evidence-acquisition policies. Experiments with six ALMs reveal a clear gap between plausible rationales and verifiable acoustic evidence reasoning, while BAEA-\textsc{Fixed} improves final verdicts and forensic rationales on both evaluation partitions. Controlled interventions further show that misleading evidence degrades both detection and grounding performance.
cs.LG / 234 / 2609.39583
MADGRAV: a multilevel anomaly-detection pipeline for gravitational-wave searches applied to LIGO data
Abstract
We present the results of \textbf{MADGRAV}, a deep-learning-based search for high-mass compact binary coalescences, applied to the data collected by the LIGO interferometers during the third observing run and during the first and second part of the fourth observing run. The \textbf{MADGRAV} pipeline consists of a series of sequential convolutional neural networks that perform anomaly detection, glitch classification, coherence testing, and signal ranking. Data from the Hanford and Livingston LIGO detectors are studied (both individually and in coherence) by way of 1 second Q-transform windows. Of the candidates that survive every stage of the pipeline, 48 reach the significance threshold, and we report 47 gravitational wave detections characterised by a false alarm rate below $1\,{\rm yr}^{-1}$ with a probability of astrophysical origin $p_{\rm astro}>0.9$. Of the 47 detections, 44 are shared with the minimally modelled coherent WaveBurst search. The observed total source-frame masses, extracted from official gravitational wave transient catalogues, are in the $14-236 M_{\odot}$ range with a median of $69 M_{\odot}$, and a median SNR of 16. We note that the recovered fraction of confident detections rises with mass: for LIGO detectors network SNR $>10$ the pipeline recovers $8.1\%$ of confident catalog events below $30 M_{\odot}$, $39.8\%$ between $30$ and $100 M_{\odot}$, and $53.3\%$ above $100 M_{\odot}$, corresponding to $33.3\%$, $45.5\%$ and $53.3\%$ of the events detected by coherent WaveBurst in the same bins. These results suggest that anomaly detection pipelines can serve as an independent detection channel complementary to matched filtering in the high-mass high-SNR regime.
cs.LG / 235 / 2609.40008
PINNing the pion: conformal deep learning for $F_π(s)$ and the $(g-2)_μ$ hadronic contribution
Abstract
Extracting the pion electromagnetic form factor $F_π(s)$ through phenomenological curve-fitting models introduces model dependence, unphysical artefacts, and kinematic inconsistencies. We introduce a Physics-Informed Neural Network (PINN) embedded in a conformal $z$-plane that constructs $F_π(s)$ directly from first principles across spacelike and timelike domains: charge normalisation and Schwarz reflection are enforced by construction, while Cauchy-Riemann analyticity, dispersion relations, Watson's theorem, and perturbative QCD asymptotics enter through the loss functional. Thus, the fundamental S-matrix principles dictate the form factor's behaviour while data act as constraints. Mapping the cut complex plane onto the unit disk bounds the Hessian norm and prevents Neural Tangent Kernel spectral starvation, two known failure modes of deep-learning optimisation. Besides $e^+e^-$ scattering data, we also incorporate $τ$-decay data through a switch that isolates the pure isovector form factor natively, bypassing model-dependent isospin-breaking pre-corrections. The network organically yields an interior zero-free form factor, while the framework tests experimental tensions around the $ρ(770)$ peak against analyticity and dispersion constraints. We obtain model-independent estimates of the pion charge radius, $\langle r_π^2 \rangle = 0.435 \pm 0.008_{\text{stat}} \pm 0.007_{\text{cali}}$ fm$^2$, the second-sheet pole parameters, $m_ρ^{\text{pole}} = 761.72\pm 1.04$ MeV and $Γ_ρ^{\text{pole}} = 135.99 \pm 1.20$ MeV, and the two-pion contribution to the muon anomalous magnetic moment, $a_μ^{ππ} = (506.48 \pm 2.02_{\text{stat}} \pm 1.70_{\text{cali}}) \times 10^{-10}$.
cs.LG / 236 / 2609.38395
Unified Optimality Conditions for Stochastic Optimal Control in the Rough Path and Itô Frameworks
Abstract
Stochastic differential equations (SDEs) can be studied via Itô calculus and rough path theory. For stochastic optimal control, these two frameworks give distinct Pontryagin Maximum Principle (PMP) optimality conditions with forward-backward SDEs (FBSDEs) or rough differential equations. We show that the adjoint equations of the Itô and rough PMPs are connected via the conditional expectation $p_t^{\text{Itô}}=\mathbb{E}[p_t^{\text{rough}} \mid \mathcal{F}_t]$, where $\mathcal{F}_t$ represents information available at time $t$. First, we derive a rough stochastic PMP for problems with adapted controls that does not use FBSDEs. Its proof extends the rough stochastic PMP over deterministic controls by considering stochastic needle variations. Second, we derive a unified PMP connecting the Itô and rough PMPs, using Itô-Stratonovich conversion formulas and duality identities between the forward tangent and backward adjoint SDEs. As a first application, we rederive the adjoint matching method for fine-tuning generative models. As a second application, we propose an indirect shooting method for a class of feedback problems. Overall, these results give a new conditional bridge connecting two popular frameworks for stochastic optimal control.
cs.LG / 237 / 2609.38561
A Parameter-Free Zeroth-Order Method with Covariance Matrix Adaptation and Effective Dimension
Abstract
Zeroth-order optimization methods are essential for solving black-box problems where gradient information is unavailable or expensive to compute. This paper presents POEM-CMA, a novel parameter-free stochastic zeroth-order algorithm that extends the recent POEM method by integrating covariance matrix alignment and the notion of effective dimension. In contrast to traditional zeroth-order approaches that rely on isotropic random directions, POEM-CMA performs anisotropic sampling by constructing a covariance matrix from gradient estimates. This enables the algorithm to focus sampling efforts on the most informative directions. We introduce the use of the empirical effective dimension $d^* = \frac{\operatorname{tr}(\hatΣ)}{λ_{\max}(\hatΣ)}$, which reflects the intrinsic dimensionality of the problem and replaces the ambient dimension in both sampling and complexity analysis. We prove that POEM-CMA achieves a near-optimal convergence rate, requiring only $\tilde{\mathcal{O}}\left(\frac{d^* κ(\hatΣ) L^2 D_{\mathcal{X}}^2}{\varepsilon^2}\right)$ stochastic zeroth-order oracle queries. The method remains fully parameter-free and demonstrates significant improvements over the original POEM in problems with low-rank structure where $d^* \ll d$. Numerical experiments on hinge-loss binary classification tasks using LibSVM datasets confirm the practical superiority of the proposed approach.
cs.LG / 238 / 2609.39301
Generalized Geometry Block Proximal Linearized Method for Multiblock Nonconvex and Nonsmooth Optimization
Abstract
This paper considers a class of multiblock nonconvex and nonsmooth optimization problems arising in many applications. Existing methods construct proximal linearized operators or their variants within standard Euclidean geometry to solve this class of problems, forcing their block variable updates to rely on the standard inner product and its induced norm. Nevertheless, this construction fails to capture the geometric structure of the target problem, leading to low numerical efficiency. To overcome these drawbacks, we propose a generalized geometry proximal linearized operator for updating block variables, and develop the Generalized Geometry Block Proximal Linearized (GGBPL) method based on this operator. Compared with existing proximal linearized operators, the proposed operator allows the block surrogate functions to be constructed using arbitrary inner products and general admissible metrics, thereby enabling the GGBPL method to adapt its updates to the geometric structure of various problems. We also introduce the inertial version of GGBPL, named the inertial GGBPL (iGGBPL) method. We further establish a new unified convergence framework under this generalized geometry, within which we prove that our methods guarantee convergence of the objective function values, establish global convergence of the generated sequence to a critical point, and derive the convergence rate of our methods. We also establish an $\mathcal{O}(\varepsilon^{-2})$ iteration complexity bound for obtaining an $\varepsilon$-stationary point. We apply our methods to two nonconvex and nonsmooth problems: sparse nonnegative matrix factorization with $\ell_0$-constraints and sparse nonnegative CP decomposition with $\ell_0$-constraints. Numerical results demonstrate the superior numerical performance of our proposed methods over several state-of-the-art methods.
cs.LG / 239 / 2609.38609
Methodological Changes to the Attention ResUNet Hourly Precipitation Postprocessor
Abstract
This note is a technical companion to a previously published preprint describing an Attention Residual U-Net that postprocesses deterministic forecasts from The Weather Company's Global and Regional Atmospheric Forecast (GRAF) model into probabilistic hourly precipitation forecasts. It documents what has changed in that method since publication. Feature-wise Linear Modulation conditioning on calendar season and forecast lead time is used to produce a single trained model for each season, replacing 192 separately trained per-month, per-lead checkpoints. Lead time is extended from 48 to 72 h. Two new input channels are used, per-pixel local solar hour and a static, monthly-varying precipitation climatology. During verification, the climatological reference against which the Brier Skill Score is computed now has an added diurnal dimension, on top of the monthly resolution it already had. Brier Skill Score and reliability are compared between the new vs. the previous training. Forecasts generated with the new training show a modest, consistent improvement of the current training over the original.
cs.LG / 240 / 2609.40140
Less is more: error-distance scaling relation for data-efficient kilometer-scale downscaling of extreme heat
Abstract
Extreme heat is where urban adaptation needs kilometer-scale data the most, but the simulations training a downscaler can cost more than they save, and how much is needed has not been identified. We measured it with CASPER, a U-Net with a structure-preserving loss downscaling 32 km reanalysis to 1 km temperature, humidity and wind, across 24 configurations of one to eight months. Held-out error grows linearly with climatological distance to the training data, RMSE = 0.83 + 2.95 d, explaining 90% of its variance against 7% for volume and predicting unseen months in advance. On held-out extreme summer weeks CASPER preserves the fine-scale structure and cross-variable physics that matched-budget baselines degrade, and matches station observations during documented heat waves to within 1.8 K. Transfer to a new region degrades geographically; 11 days of local simulation cuts Vancouver's held-out error from 3.8 to 1.3 K. Training periods should span the target climate: the same accuracy for four times less simulation, putting kilometer-scale downscaling of extreme heat within reach of groups without large computing facilities.
cs.LG / 241 / 2609.39090
A strategic roadmap for an atomistic machine-learning ecosystem
Abstract
Data-driven machine learning (ML) techniques have become an essential tool in many domains of science. Their application to atomistic simulations of matter is particularly widespread and impactful. This success is due largely to the existence of a well-developed and established physics-based modeling framework, ranging from first-principles electronic-structure calculations to molecular dynamics and statistical sampling, into which ML was integrated naturally to reshape long-standing trade-offs between accuracy, efficiency, and scale. Nevertheless, this integration raises both conceptual and practical challenges, from choosing between data-centric and physics-based modeling approaches to adapting established software stacks to modern hardware accelerators and ML libraries. As the field evolves rapidly, fueled in part by widespread enthusiasm but also by tangible impact, it seems appropriate to take a moment to consider the current state of the art and open challenges, and reflect on what can be done to better coordinate efforts across the community. With this goal in mind, several members of this community met in Lausanne in January 2026 at CECAM to discuss algorithms, models, software and hardware infrastructure, and the most promising scientific applications that have become possible thanks to the use of artificial intelligence in atomic-scale simulations. This strategic roadmap paper summarizes the outcomes of these discussions, suggesting some long-term goals, and some concrete actions, to establish a healthy, sustainable and impactful atomistic ML ecosystem.
cs.LG / 242 / 2609.38545
Say, Echo, Do: Strategic Narratives and Revealed Positioning in Financial Markets
Abstract
Machine-learning signals built from financial text treat what institutions say, and what the media repeat, as evidence about value. But whoever shapes a narrative may be trading against it. We study markets with three observable voices: institutional statements (Say), media repetition (Echo) and revealed positioning (Do). We ask when words should be followed and when they should be faded. In a linear-quadratic model of an informed institution that speaks and trades before a partly credulous crowd, talking an asset down while buying it is optimal exactly when $\varphi^2<2λk<\varphi$. A distribution-free identity then shows that when the observable Say-Do covariance is negative, words carry negative predictive content and should be faded. For measurement, we derive (i) an exact factorised posterior over which articles are echoes, combining arrival times with embedding similarity; (ii) a return-aligned contrastive objective that attains its bound exactly when squared embedding distances are an increasing affine function of squared outcome distances, with the tightest loss-based certificate of which neighbour rankings survive imperfect training; and (iii) a path-signature statistic for who moved first. In a controlled market with known ground truth, echo sentiment predicts returns with a significantly negative sign in all 29 simulated markets, the rolling Say-Do correlation flags false-alarm events with an AUC of 0.90, and return-aligned embeddings organise headlines by consequence rather than topic. We also report where the tools fail.
cs.LG / 243 / 2609.38403
Advantage of Sample Complexity in Quantum PAC Learning Requires Inverse Access to State-Preparation Unitaries
Abstract
Whether quantum computation can reduce the amount of data sampled from an unknown probability distribution required to learn a prediction rule is a fundamental question in quantum machine learning. Quantum PAC learning studies this question using quantum data as a quantum state whose squared amplitudes encode the unknown distribution from which classical learning data are sampled. With only copies of such quantum data, the optimal worst-case sample complexity asymptotically matches that of classical PAC learning. In contrast, access to both a state-preparation unitary for this state and its inverse can improve the query-complexity dependence on the accuracy parameter in realizable learning. However, it has remained unclear whether forward-only access allows such an improvement. In this work, taking the worst case over compatible state-preparation unitaries and their finite ambient dimensions, we show that the optimal forward-only query complexities of realizable and agnostic learning are, respectively, $Θ((d+\log(1/δ))/\varepsilon)$ and $Θ((d+\log(1/δ))/\varepsilon^2)$, where $d$ is the VC dimension of the concept class, $\varepsilon$ the accuracy parameter, and $δ$ the failure probability. These bounds match the optimal sample complexities with classical data or quantum data copies. To prove them, we establish a reduction using $q$ copies of the prepared state to approximate the Haar-averaged output of any $q$-query forward-only algorithm. These results show that forward-only access cannot provide an asymptotic query-complexity advantage over learning from classical data or quantum data copies in this worst-case setting, and establish the essential role of inverse access in the known realizable-setting improvement. Our reduction also provides a new framework for analyzing limitations of forward state-preparation access via state-copy lower bounds.
cs.LG / 244 / 2609.38835
Average-and Last-Iterate Lower Bounds for Optimistic Matrix Mirror-Prox in Quantum Zero-Sum Games
Abstract
Optimistic matrix mirror-prox (OMMP) computes $ε$-approximate Nash equilibria in quantum zero-sum games with an $O(1/\varepsilon)$ average-iterate guarantee [arXiv:2311.10859]. We investigate whether this dependence on accuracy is tight and whether geometric last-iterate convergence can be guaranteed. We study these questions through explicit games with one qubit per player. First, we prove an $Ω(1/\varepsilon)$ lower bound for the uniform-average output that includes the maximally mixed initial state, independently of the regularizer and step size. Second, we construct a fixed game on which optimistic gradient descent-ascent (OGDA), initialized at the maximally mixed state, has last-iterate Frobenius distance to equilibrium $Θ(1/t)$ and duality gap $Θ(1/t^3)$ for every sufficiently small fixed step size. A separate fixed game exhibits arbitrarily long delays in reducing the initial error by a constant factor across a family of initial states. Finally, we give a fixed game with a unique, strictly complementary equilibrium on which optimistic matrix multiplicative weights updates (OMMWU) converge only polynomially from the maximally mixed state for every fixed positive step size. The last-iterate Frobenius distance and quantum relative entropy from the equilibrium to the iterates decay as $Θ(1/t)$, while the duality gap decays as $Θ(1/t^2)$.
cs.LG / 245 / 2609.39172
A Width-Matched Comparison of Hybrid Quantum-Classical Self-Supervised Learning for Fingerprint Recognition
Abstract
Fingerprint recognition is a widely deployed biometric, but supervised training requires large labeled enrollment sets. Self-supervised learning (SSL) removes this requirement, and hybrid quantum-classical models have been proposed to enrich the learned representations. Prior quantum SSL studies consider a single contrastive objective, so it is unclear whether reported benefits depend on the objective or can be attributed to the quantum circuit. We insert the QuFeX quantum feature-extraction module into three SSL frameworks, the contrastive SimCLR and MoCo v2 and the non-contrastive BYOL, and compare each hybrid with its classical counterpart at matched representation width (8 features, equal to 8 qubits) on the SOCOFing fingerprint dataset, with a CIFAR-10 control, using k-nearest-neighbor identification on encoder features. In single-run experiments the hybrid scores clearly higher for both contrastive objectives, whereas for BYOL a multi-seed analysis shows no reliable difference, suggesting that any benefit depends on the SSL objective. A hardware-efficient circuit (QNet) does not show the same gain. We examine whether the gains can be attributed to the quantum circuit, considering circuit architecture, trainable parameter count, nonlinearity, and the classical simulability of 8-qubit circuits.
cs.LG / 246 / 2609.38842
Learn-Then-Differentiate Gradient Estimation
Abstract
Learn-then-differentiate (LTD) estimates gradients by fitting a model to simulation outputs and differentiating it. We develop a unified framework explaining what LTD differentiates and how accurately it estimates gradients. For models with a weighted representation, LTD differentiates a learned representation of the underlying probability measure. We then show how accuracy guarantees for fitted models translate into guarantees for gradients and higher-order derivatives, with rates approaching the standard Monte Carlo rate under suitable smoothness conditions. The framework recovers established results for kernel regression, local polynomial regression, and kernel ridge regression, and yields further guarantees for multiple kernel learning and smooth neural networks. These results provide a common foundation for understanding and analyzing LTD across learning methods.
cs.LG / 247 / 2609.38375
Lower Bounds for Linear-Oracle Online Learning
Abstract
Can a constant number of linear minimizations per round improve on the $T^{3/4}$ regret rate of online Frank-Wolfe on general convex sets? Weibel et al. conjectured that fixed-coefficient methods cannot. We prove their conjecture and extend the lower bound to every deterministic learner in an oracle-only model. The learner receives an initial feasible point and a diameter bound, and must remain feasible on every domain consistent with its oracle replies. For $T$ rounds, at most $b$ calls between decisions, diameter bound $D$, and gradient norm bound $L$, we construct an instance in dimension $d=2b(T-1)+1$ with regret at least $2^{-1/4}LDb^{-1/4}T^{3/4}$. The adversary fixes the domain, initial point, deterministic tie rule and linear losses before play. The vertices form a path on which every point available before a decision has zero current loss, while the final vertex has negative loss on every round. For constant $b$, the result matches the known upper rate for dimension-independent guarantees. For one-call fixed schedules with a nonzero coefficient on the newest gradient, a second construction gives regret at least $3LDT^{3/4}/4$ with unique minimizers at every issued query. Exact-arithmetic certificates for the tuned schedule of Weibel et al. closely match their finite-horizon numerical worst cases, with unique oracle replies.
cs.LG / 248 / 2609.38524
Generative sequence modeling for infinite memory processes via predictive states
Abstract
We consider estimating the one-step-ahead conditional distribution of a multivariate stochastic process. Many existing approaches rely on assumptions such as finite-range memory, sparsity, or additivity, which can be poorly suited to processes with long-range nonlinear interactions. However, without such structural assumptions, nonparametric estimation is challenging due to the curse of dimensionality. To address this challenge, we introduce a new estimation approach based on the predictive states of a process, possibly with infinite-range memory. We show that our estimator achieves fast convergence rates when the past history can be compressed into a low-dimensional statistic that is sufficient for predicting the future. Specifically, we show that the statistical complexity of the estimation problem is determined by the intrinsic dimension of the predictive state space. We establish guarantees for an instantiation of our method based on deep neural network estimators, and we support these theoretical results with experiments.
cs.LG / 249 / 2609.38880
Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data
Abstract
We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $γ$, write $H=(1-γ)^{-1}$, and let $μ_{\min}$ and $t_{\operatorname{mix}}$ denote the minimum stationary probability and total-variation mixing time. We prove that last-iterate TD achieves sup-norm error at most $\varepsilon$ with high probability using$\widetilde O\left( \frac{H^3}{μ_{\min}\varepsilon^2} +\frac{t_{\operatorname{mix}}}{μ_{\min}} \right)$ transitions, for $0<\varepsilon\leq1$. This rate holds both for a constant step size selected for the target accuracy and for a decreasing schedule independent of the target accuracy and terminal time. The latter gives a simultaneous guarantee over all times beyond an explicit transient threshold. The statistical term retains the cubic effective-horizon dependence of synchronous TD, and the additive mixing transient has no extra horizon factor. The result allows non-reversible chains, arbitrary initial state distributions, and bounded rewards that may depend on the next state. The proof uses an anchored local Poisson equation in reverse time to control stochastic fluctuations without a mixing-time factor, and a hitting-time compensation identity to bound initialization error. The latter also yields a finer transient in terms of the worst expected reverse hitting time. A bound on the expected cumulative propagation mass extends this argument to decreasing step sizes. A three-state construction with known deterministic rewards gives matching minimax lower bounds for the statistical and mixing terms, up to logarithms, over specified model classes in a slow-mixing parameter regime.
cs.LG / 250 / 2609.38916
Warm-starting PDE solvers with any-dimensional machine learning
Abstract
Any-dimensional machine learning models, such as graph neural networks (GNNs), can be naturally trained and evaluated on inputs of different sizes and dimensions. Inspired by the GNN transferability literature, we show mathematical conditions under which a partial differential equation (PDE) learning-based solver can be trained in small dimensions and directly applied to solve a higher dimensional PDE in a zero-shot fashion. These conditions are based on symmetries in both the partial differential equation and the initial data. When the equations satisfy the symmetries but the data does not, which is the case for many PDEs arising from physics, we show that our theory gives a principled way of warm-starting low-dimensional PDE solvers for higher dimensional PDEs. We apply this method on the heat equation, Burgers' equation, and the compressible Navier--Stokes equations, improving the performance in both zero-shot and typical training regimes on high dimensional data. For example, we train a surrogate model on 2D Navier--Stokes data and achieve better results on 3D test data than a baseline surrogate model trained on 3D data, while only using 12$\%$ of the flops and 20$\%$ of the total data size.
cs.LG / 251 / 2609.39020
Minimax rates for learning spectral Barron functions by deep ReLU neural networks
Abstract
We study how well deep neural networks approximate and learn spectral Barron functions. Recent studies have shown that these function classes can be efficiently approximated by shallow neural networks without suffering from the curse of dimensionality. We complement these results by providing new approximation bounds for deep networks with ReLU activation and establishing the minimax rates for learning these function classes. Specifically, we show that $d$-dimensional spectral Barron functions with smoothness index $s>0$ can be approximated by deep ReLU neural networks with approximation rate $\widetilde{\mathcal{O}} (S^{-\frac{1}{2}-\frac{s}{d}})$, where $S$ denotes the number of nonzero parameters in the network. Using this approximation result, we further show that deep ReLU neural networks can learn spectral Barron functions in a fast rate $n^{-\frac{d+2s}{2d+2s}}$ with $n$ training samples. Finally, we prove that this convergence rate is minimax optimal up to logarithmic factors.
cs.LG / 252 / 2609.39173
Asymptotic Properties of Support Vector Machines in High-Dimension, Low-Sample-Size Settings under a Spiked Model
Abstract
In this paper, we consider asymptotic properties of the support vector machine (SVM) in high-dimension, low-sample-size (HDLSS) settings under a spiked model. The existing theory of the SVM in the HDLSS context relies on the geometric representation of HDLSS data, which requires that the eigenvalues of the covariance matrices are not dominant. We first show that the geometric representation does not hold under the spiked model. We show that the Gram matrix of HDLSS data converges in distribution to a random matrix, namely, the HDLSS data converge to a random configuration in a finite-dimensional space whose dimension is given by the number of the spikes. We show that the misclassification rates of the SVM do not tend to zero, that is, the SVM does not hold the consistency property. We also show that the bias-corrected SVM (BC-SVM) does not give preferable performance in this setting because the bias term itself should be modified. In order to overcome such difficulties, we propose a spike-corrected SVM (SC-SVM). We show that the SC-SVM holds the consistency property when the sample size goes to infinity, and that the growth of the sample size is essential in the sense that any projection-based procedure fails when the sample size is fixed. Finally, we check the performance of the classifiers by numerical simulations.
cs.LG / 253 / 2609.39212
Minimax Additive Regression under Unknown Dependent Designs
Abstract
We study additive regression under a potentially non-product random design on $[0,1]^d$, allowing the dimension $d$ to grow with the sample size $n$. We introduce coupled smoothness classes that separately control the regularity of the marginal densities and the density-weighted additive components. To handle dependence, we adapt a Riesz-basis construction for functional ANOVA models and establish compatibility bounds with constants independent of the dimension under uniform bounds on the joint density. We construct thresholded least-squares estimators and establish matching minimax upper and lower bounds for prediction with known or unknown marginal densities, under suitable dimension-growth conditions. When the marginal densities are at least as smooth as the weighted components, the unknown-density problem attains the known-density minimax rate. When the densities are less smooth, their regularity determines the minimax rate over the coupled class. Finally, we show that the centered additive components can be recovered at the same aggregate upper rate, without an additional order of error.
cs.LG / 254 / 2609.39367
A Dynamical Theory of LoRA in Continual Learning
Abstract
Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.
cs.LG / 255 / 2609.39397
Towards Optimal Inventory Control under Censored Demand: A Biased Sample-Average Approximation Approach
Abstract
We study data-driven multi-period lost-sales inventory control under censored demand, where a stockout reveals only that demand exceeded the stocking level. We develop a unified, model-based framework for policy learning from censored data, built on a new cost decomposition for base-stock policies and a biased sample-average approximation (SAA) approach. The cost decomposition allows us to propose a new coverage condition under which censored observations are informative enough for sample-efficient policy learning. Guided by this coverage condition, we design two biased SAA algorithms: an upper-biased one that achieves near-optimal sample complexity under the offline coverage condition, and a lower-biased one that actively generates the required coverage and achieves near-optimal regret online. More broadly, this biased SAA approach provides a general principle for implementing pessimism and optimism under censored feedback, which may be of independent interest.
cs.LG / 256 / 2609.39440
Principal Component Regression Dominates all Monotone Spectral Filters for Linear Regression
Abstract
We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal component regression (PCR), gradient descent (GD), and ridge regression. We show that PCR dominates all monotone spectral filters: compared to any such filter, the risk of optimally tuned PCR is no bigger by a constant factor for all problems. Furthermore, the dominance is strong if the filter is separated from step functions (e.g., GD and ridge): there exist problem instances for which the risk of PCR is smaller by a polynomial factor in sample size dependence. Our comparison results show that PCR is optimal and thus admissible among monotone filters, significantly extending Wu et al. (2026)'s result that GD strongly dominates ridge. From a technical perspective, we establish new upper and lower bounds for general spectral filters, which are instance-wise sharp when specialized to ridge or GD, recovering or improving the best-known bounds.
cs.LG / 257 / 2609.39449
Distributionally robust linear regression through the lens of adversarial training
Abstract
Distributionally robust optimization (DRO) studies parameter estimation under uncertainty in the underlying probability distribution and has emerged as a principled framework for analyzing robustness and generalization. In particular, Wasserstein DRO, with distributional uncertainty induced by the Wasserstein distance, generalizes several popular regularizers. This paper studies Wasserstein DRO linear regression, unifying square-root Lasso and adversarial linear regression as important special cases. We prove that many properties of these two special cases carry over to this general method. In particular, we show (i) deterministic and non-asymptotic in-sample error bounds $O(n^{-1/2})$ in general and $O(n^{-1})$ under design matrix and sparsity conditions; (ii) insensitivity to the noise level, also known as the pivotal property; and (iii) solution equivalences for small and large ambiguity sets. The key proof step is to recast the method into a quadratic form, mimicking adversarial linear regression. We also show that the method can be solved efficiently, and we validate our findings through numerical simulations.
cs.LG / 258 / 2609.39525
Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision
Abstract
Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intractable statistical models and offers near instantaneous inference for new datasets after prepaying the training cost. Although theory guarantees faithfulness under ideal convergence, practical amortized inference still requires iterating over architectures and optimization choices and ultimately ``satisficing'' under finite simulation, compute, and time budgets. Even the best-performing solution may thus retain avoidable representation gaps that typically require problem-specific fixes. Here, we propose a generic alternative which improves training dynamics with auxiliary guidance losses applied to internal representations. Specifically, we show how such guidance leads to faster convergence when training data is abundant and to better performance when it is scarce. We formalize representation gaps as getting stuck in a local optimum at the information bottleneck between the parts of the network tasked with feature learning and those tasked with conditional distribution learning, and offer a generic diagnostic to separate summary failures from inference failures. Finally, we demonstrate that auxiliary supervision improves convergence speed and accuracy on a range of challenging real-world inference problems.
cs.LG / 259 / 2609.39829
Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification
Abstract
Modern machine learning depends heavily on massive datasets, but obtaining high-quality annotations at scale is often expensive. As a result, learning from noisily-labeled data has become common, making accurate estimation of the label-noise transition matrix crucial. However, existing transition matrix estimators rely on the fragile estimation of class-posteriors and do not provide finite-sample performance guarantees. In this work, we propose a novel methodology to estimate the transition matrix based on one-sided selective classification. This approach bypasses class-posterior estimation, provides finite-sample performance guarantees, and leverages flexible learning methods for binary classification. Moreover, we introduce effective algorithms to implement the proposed methodology and provide their refined finite-sample performance bounds.
cs.LG / 260 / 2609.40024
Amortized Bayesian Inference on Multilevel Models of Arbitrary Structure
Abstract
We develop a general method for amortized Bayesian inference on multilevel models of arbitrary structure. Given a generative model specified as a directed acyclic graph, our method automatically derives valid factorizations of the joint posterior and matching neural network architectures. The key steps, graph expansion and graph inversion, yield an inverse graph that determines how inference networks are stacked and conditioned, producing factorizations that amortize over the number of groups and the number of observations within each group. Unlike approaches that simplify the dependency structure to speed up learning or inference, our method preserves all conditional independence and exchangeability assumptions of the generative model. Across three case studies, it closely matches gold-standard samplers on models with more than 6,500 parameters while reducing inference to a near-instant forward pass once trained.
cs.LG / 261 / 2609.40051
Proximal Balancing for Causal Effect Estimation under Unmeasured Confounding
Abstract
Estimating causal effects from observational data is central to science and policy, but the effects are not identified when confounders are unmeasured. Proximal causal inference addresses this problem with proxies of the unmeasured confounders. However, existing proxy-based approaches either designate proxy roles and solve an inverse problem, which is ill-posed and hard to estimate with high-dimensional proxies, or use a latent-variable model, which assumes that the learned latent variable matches the hidden confounder and leaves bias when it does not. To address these challenges, we introduce proximal balancing. It carries the classical idea of covariate balancing to confounders that are observed only through proxies: it learns a low-dimensional summary of the covariates and proxies that makes the treatment groups comparable, and then adjusts for this summary. It needs no designated proxy roles, inverse problem, or latent model. We give identification theory, finite-sample guarantees, and a practical algorithm, PROBE. We demonstrate the method on low-dimensional, high-dimensional, and image proxies and on real-world data.
神经与进化计算 (cs.NE)
6
cs.NE / 1 / 2609.38432
Simulating Synchrony Loop Networks in the Open Source RISP Neuroprocessor
Abstract
Neuromorphic spiking neural networks (SNNs) offer a promising alternative to conventional deep neural networks for tasks with computational resource or data constraints. However, their practical applications have been limited by comparatively weak performance on complex learning tasks. Experimental approaches such as Synchrony Loop Propagation (SLP) increasingly seek to address this problem through more sophisticated and heterogeneous neuron models, and have achieved encouraging initial results. However, these neuron models do not readily translate to standard neuromorphic systems designed to support simple leaky integrate-and-fire neurons. We present an implementation of an SLP network on the RISP neuroprocessor, an event-driven neuromorphic simulation platform, and evaluate its performance on an unsupervised musical instrument clustering task. The network achieves clustering accuracy comparable to the DBSCAN algorithm, while providing over an order of magnitude improvement in runtime speed and computational efficiency relative to a prior non-neuromorphic SLP implementation. These results demonstrate that SLP's core mechanisms can be effectively translated into a neuromorphic architecture to support complex unsupervised learning. More broadly, this work highlights the potential of heterogeneous and extensible neuron models to expand the design space of neuromorphic systems to more complex learning tasks.
cs.NE / 2 / 2609.38665
How much of fly walking is written in the wiring?
Abstract
Connectome models of the fly nerve cord generate walking-like motor rhythms, but oscillation alone does not show that the specific wiring matters. Here we provide, to our knowledge, the first test of which features of motor output depend on the specific wiring. We simulated the leg motor systems of two independent Drosophila connectomes, with synapse counts as fixed weights and glutamatergic synapses treated as inhibitory, and compared each with six families of rewired networks that preserve progressively more of its structure, using pre-registered criteria. We find that rhythm is generic but antagonist coordination is not: many rewired networks were more rhythmic than the real ones, yet the real wiring coordinated antagonistic motor pools more strongly than every rewired network, most of all at the thorax--coxa joint. We trace this specificity to how premotor input is allocated between antagonistic pools. Both connectomes carry Sherrington's reciprocal innervation---neurons that excite one pool inhibit its antagonist---and no rewired network does. Reassigning premotor inputs between the pools abolished coordination even when motor neurons' typical input strength changed little (all pre-registered criteria met in one connectome; same direction in the other). Coordination, not rhythm, therefore reveals whether a connectome model's wiring matters.
cs.NE / 3 / 2609.39248
Null-model treatment of the sensory-motor boundary changes an evolutionary connectome comparison
Abstract
Randomised copies of a connectome are the usual baseline for asking whether measured wiring matters, and the answer depends on what the randomisation preserves. We evolved embodied foraging agents whose brains are a compressed adult Drosophila connectome (FlyWire v783; 512 cell-type groups and 1,000 Kenyon cells) alongside agents built on randomised wiring, in pre-registered experiments with ten seeds, four ecologies and 600 generations. Two standard randomisations, a column shuffle and degree-preserving edge swaps, route 10.6 to 10.7 % of olfactory output directly onto descending motor groups, against 0.012 % in the connectome. On the registered primary endpoint, fitness averaged over the run, no difference was detected; at the last common-garden probe the connectome was behind both controls (-0.22 and -0.20 fitness units on seed means). Against controls that keep every sensory-output and motor-input edge and rewire only the interior, the seed-mean difference lay within a +/-0.10 equivalence bound (+0.002 and -0.074, unchanged under a calibration that also matches activity spread), although per ecology the interior column shuffle was ahead by 0.26 in one of four ecologies at ten seeds, a lead that ten further pre-registered seeds did not replicate. Rewiring the connectome so that it acquires the shortcut raised its fitness by 0.44 (10 of 10 seeds) and its dependence on olfaction from 0.15 to 0.99; graded doses raised both in step; at comparable swap counts the full dose was ahead of an interior-only sham by 0.53 (10 of 10 seeds); and a sham that rewired the same boundary edges without creating shortcuts matched the connectome (+0.007) while the full dose was ahead of it by 0.60. What a null preserves at the sensory-motor boundary can decide an evolutionary connectome comparison, and sensory-to-motor path statistics belong next to the degree statistics a null is said to preserve.
cs.NE / 4 / 2609.39272
An Island-Based Parallel Biased Random-Key Genetic Algorithm for the Three-Dimensional Trailer Loading Problem
Abstract
The Three-Dimensional Trailer Loading Problem (3D-TLP) involves determining the optimal placement and orientation of heterogeneous items within the confined space of a trailer while maximizing volume utilization and satisfying a wide range of complex logistical and safety constraints. The 3D-TLP is NP-hard, rendering exact optimization approaches computationally impractical for large-scale industrial applications. To address this challenge, we propose an enhanced Biased Random-Key Genetic Algorithm (BRKGA) accelerated through a novel island-based parallelization framework, PANGEA. The proposed method combines the search efficiency and robustness of BRKGA with a multi-population evolutionary scheme for genetic algorithms. This island-model strategy promotes population diversity, mitigates premature convergence, and significantly reduces computational times. The proposed solution was validated in a real trailer loading process, providing an effective solution approach for real-world large-scale logistics.
cs.NE / 5 / 2609.38527
Behavioral Persistence and Incomplete Functional Transfer of Co-evolved Communication in Evolutionary Robotics
Abstract
This work evaluates the direct transfer of a co-evolved communication protocol from a 2D simulation to a 3D physical environment, without retraining the network weights. Two e-puck-type robots, controlled by a GRU network with residual connection, were evaluated in a food-seeking task with social signaling. The sensory and motor translation layer required three corrections for stable physical operation, including the calibration of a hunger term based on a measurable asymmetry in the trained residual weights. Even with these corrections, the transfer was partial and asymmetric: one agent reached the food source in one of thirty tested seeds, while the other did not reach it in any. Task success was measured by both agents reaching the food area. An additional experiment incorporating explicit directional information in the social channel produced observable changes in the trajectory of the receiving agent and improvements in several specific cases. However, these improvements were not enough to allow the second agent to reach the food source, suggesting that the limitation may not be explained solely by signal translation, but also by the ability to navigate under the new physical constraints. The results suggest that successful transfer of emergent communication may depend not only on preserving the signaling process itself, but also on preserving the ecological and navigational conditions under which the protocol evolved.
cs.NE / 6 / 2609.39239
Evolutionary foraging in grids: Intermittent search dynamics emerge in finite, depletable landscapes
Abstract
How search strategies evolve in finite, depletable landscapes remains a question in foraging theory. We study this problem with an evolutionary simulation in which agents forage on a two-dimensional toroidal lattice containing non-renewable resources distributed uniformly or as Lévy dust. Each agent carries a heritable genome encoding step lengths, velocities, and turning angles, and selection acts on a fitness function combining energetic gain, movement cost, and coverage efficiency. By allowing movement traits to evolve without imposing a prescribed power-law step-length distribution, we test whether evolved trajectories are better described by intermittent-search or Lévy-walk dynamics. Our results indicate that evolved search is more consistent with intermittent dynamics than with strict scale-free Lévy motion in the finite depletion-driven landscapes considered here. We characterize the dynamics by fitting second- and fourth-order displacement moments to intermittent-search and Lévy-walk models. While a Lévy-like random walk fits the evolutionary trajectories well (mean adjusted $R^2$ > 0.9 in most tested conditions), intermittent search achieves a closer fit (mean adjusted $R^2$ > 0.99) for all tested resource distributions. This preference holds across the tested grid sizes and resource densities. Five independent evolutionary runs per environment on a 503 x 503 grid at nominal resource density $ρ$ = 0.15 reproduce this preference for the uniform environment and five Lévy-dust environments. Evolution rapidly reshapes the movement genome toward short displacements while retaining a sparse tail of longer relocations, consistent with local exploitation punctuated by occasional transfer. The framework provides a controlled setting for studying how search rules emerge under resource limitation and may inform resource-constrained exploration in autonomous systems.
计算语言学 (cs.CL)
80
cs.CL / 1 / 2609.38355
Halluscoring 2026: The first shared task on llms hallucination detection and answer verification
Abstract
We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. The shared task is organized into two main tasks, each comprising two subtasks, for a total of four subtasks. Task 1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, 10 of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging under distribution shift, with the winning team achieving AUC-ROC test scores of 0.772 and 0.767 for Subtasks 1.1 and 1.2, respectively. For Task 2, the winning team achieved scores of 0.882 and 0.857 in the Islamic and general-knowledge subtasks, respectively, under assisted evaluation.
cs.CL / 2 / 2609.38357
Evaluating Language Model Safety Across Long Adversarial Conversations
Abstract
Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns. We evaluate three open-weight, instruction-tuned models on two harmful prompts across different conversation lengths and random seeds. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe. Across all model-prompt combinations, first-turn safe-response rates ranged from 85% to 100%. By depth 11, they dropped to 38-61%, and by depth 101, to 15-44%. This decline appeared across models and continued well beyond the short interactions typically used in multi-turn safety evaluations. These results provide proof-of-concept evidence that strong single-turn safety does not necessarily persist during sustained adversarial interaction. They highlight the need for long-horizon evaluations and conversation-level safeguards that account for risk accumulating across turns.
cs.CL / 3 / 2609.38427
Policy-Conditioned AI-Use Detection: An Evidentiary Framework for Academic Publishing
Abstract
Major venues now publish detailed rules about how authors, reviewers, and area chairs may use AI, and those rules differ by role, by task, and by what must be disclosed. AI detection, the instrument usually proposed to enforce them, estimates something else: whether an AI model wrote the text. We argue that this target is misaligned with the decisions conferences and journals face, and propose policy-conditioned AI-use detection, an evidentiary framework for assessing whether a human--AI workflow complied with a stated rule. Policy makes the governing rule an explicit input. Inference reports hypotheses, evidence, calibration regime, and uncertainty in place of verdicts such as "AI detected". Evaluation builds benchmarks from reproducible pipelines that generate compliant and non-compliant workflows, and reports true positive rate at a false positive rate the venue fixes in advance. We work the framework through peer review, where at plausible violation rates a detector at a strong operating point still flags more compliant authors than violating ones. The framework therefore also names what a venue must instrument: structured disclosure, approved-tool routing that respects reviewer confidentiality, and a path by which a finding can be contested. Under this framing a detector is not an authorship classifier but an auditable procedure with an error rate the venue fixes in advance and can defend.
cs.CL / 4 / 2609.38469
The Backdrop Exposes What the World Around an Agent Costs It
Abstract
Agent benchmarks test agents in worlds that stay still. Deployed agents work in worlds that other people also change. Someone texts the agent to send the money elsewhere or an order confirmation asks it to reply with a door code. We present BACKDROP, which asks how much of an agent's capability in a clean world survives in such a world. BACKDROP takes a task along with the agents execution environment, and plants four everyday hazards in its world, one at a time and all together. The instruction and the correct end state stay the same. Each hazard asks one question. Authority: does a message from another person override the user? Injection: does text planted in a record redirect the agent? Boundary: does a request pull it into an app it was not given? Fault: after a write fails without saying whether it landed, does the agent check before it retries? Across 3,678 variants and 16 models, , the average pass rate falls from 69.5% to 31.3% once all four hazards are present; the strongest models fall furthest (Claude Fable 5.1 from 96.6% to 56.0%). Agents have learned to resist injected text but often follow other unauthorized requests of other people. With all four hazards present, and counting only runs where the planted text reached the agent, agents followed another person's message in 46.4% of runs and injected text in 20.3%. The gap is consistent throughout all 16 models. BACKDROP formalizes these gaps and shows how an agent's score in a task's world is a ceiling on real-world performance.
cs.CL / 5 / 2609.38480
KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy
Abstract
Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment.
cs.CL / 6 / 2609.38530
Shifting Mechanisms: How Positional Encoding Choice Shapes In-Context Retrieval
Abstract
Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention (SWA NoPE). However, how these choices shape in-context retrieval remains unclear. To study this question, we take a mechanistic view, tracing how positional encoding (PE) choice shapes the internal mechanisms models use for in-context retrieval. Across 22 open-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval. We further show on a controlled pre-training ablation that confining positional encoding to local layers produces this semantic shift, degrading representations of positional information. Finally, we show that the reported long-context gains of PE hybrids mask a retrieval trade-off: SWA NoPE improves over RoPE on multiple-target retrieval and QA, but degrades when distinguishing competing keys. We show that these behavioral differences better track the mechanism shift from positional toward semantic mechanisms than a uniform improvement in long-context retrieval.
cs.CL / 7 / 2609.38604
Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent
Abstract
Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicate, change their goals, and run out of patience. We define this task setting as Interactive Intent Alignment, where agents must recover and continuously track the user's current intent despite imperfect communication and evolving goals. To study this setting, we introduce Drift-Bench++, a principled benchmark construction pipeline for verified executable tasks with controlled misalignment and intent shifts, along with an interaction protocol featuring finite patience, diverse simulated users, and silent interaction-conditioned shifts. We further develop GRIP, a comprehensive evaluation protocol covering task grounding, user realism, inquiry effectiveness, and adaptation to evolving intent. Across diverse environments, models, and interaction conditions, stronger interaction consistently helps but remains far from oracle performance; Validation on deployed ProdAgent sessions further shows that the modeled failures are prevalent and consequential in deployment. By providing a unified, executable benchmark for interactive intent alignment, Drift-Bench++ offers a foundation for evaluating and advancing agents under realistic communication and evolving intent.
cs.CL / 8 / 2609.38612
StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams
Abstract
As natural language drives more applications, language models increasingly run inside programs as decision components: the program sends them the current state and acts on the returned decision until a newer one arrives. When evidence changes during inference, a decision correct for its own state can stay in force after that state has passed, as when a call recorder keeps running after a customer starts reading out a card number; untimed (offline) accuracy counts such an error as correct. We introduce StreamDecisionBench (SDB), which evaluates the decision in force at every instant and attributes every erroneous instant to judgment, latency or both. Its scenarios stream evidence in four application families, with reference decisions computed from public rules by executable code. We summarize in-force accuracy across update intervals of 1-5 s by its normalized area under the curve on a logarithmic time axis, giving equal weight to equal multiplicative ranges. Across six settings of four hosted models, this score stays within 2.9 points of the scenario-wise product of untimed accuracy and an oracle's integrated timing score. Reasoning improves judgment, but at low effort latency costs Luna and Terra, two GPT models we also evaluate without reasoning, 38.8 and 42.8 points relative to untimed accuracy; a faster component with weaker judgment attains a similar integrated score to Terra without reasoning. The aggregate and family curves show where these tradeoffs change, making the evaluation's time-scale dependence visible.
cs.CL / 9 / 2609.38627
Marking Contour Tones in Yorùbá
Abstract
Yorùbá is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexical items whose conventional spellings avoid vowel lengthening that would otherwise provide a host syllable for the second tone. A particular concern is a class of names in which the conventional spelling does not just omit tonal information but inverts the meaning of said name, sometimes asserting the opposite of what the name intends. This paper describes the problem, illustrates the inadequacy of current solutions, and proposes the adoption of the caron and circumflex marks. These are symbols with precedent in Yorùbá phonological scholarship since Olmsted (1951), used as orthographic conventions on single vowels to encode rising and falling contour tones, making them accessible for the first time through standard keyboard input and computational text processing. The proposal is supported by an implementation in the WriteYoruba keyboard and the TTSYoruba speech synthesizer, whose architecture and listener evaluation are reported separately (Tubosun et al., 2026).
cs.CL / 10 / 2609.38630
Strong Multilingual Privacy Tagging at Encoder Speed
Abstract
Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with fine-grained distinctions supporting varied redaction policies and methods for cheaply learning additional distinctions. We fine-tune a multilingual encoder with an affine span-tagging head on frontier-model annotations in 35 languages, replay mapped human gold with coverage-aware masking so unannotated types are not treated as negatives, and repair subword boundaries with a learned +/-1-character adjustment. On 1,283 human-gold test segments in seven languages, best measured redaction F1 is 88.8, against 69.1 for published GLiNER2 with 11 unrepresentable types excluded from its task (68.8 without that exemption), 67.8 for GLiNER2 adapted to the new training data, 57.3 for Microsoft Presidio and 35.8 for the best published OpenAI Privacy Filter fine-tune. Adding about 50,000 annotated training sentences and increasing human-gold replay improves exact typed-span F1 from 74.5 to 76.3 on Ont3, our 31-type frontier-annotated NER evaluation of 1,201 development segments. Mapped-gold replay alone raises human-gold F1 by ten points without loss on frontier-annotated text; boundary adjustment adds 1.7 exact typed-span F1 points on Ont3. Local LLMs fitting on a single 96-GB GPU underperformed as prompted annotators and frozen encoders, with encoding 30-95 times slower than XLM-R inference and prompted annotation roughly 180-1,100 times slower in the evaluated configurations. The encoder architecture delivers 4.9 times GLiNER2's CPU throughput. We release code, prompts and training recipes, with data-acquisition scripts and source links.
cs.CL / 11 / 2609.38660
Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
Abstract
Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On the public MuSC benchmark, SMART obtains the best model result across all four language pairs and also achieves the best human-evaluation result, with an overall score of 4.50/5.
cs.CL / 12 / 2609.38792
Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
Abstract
We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
cs.CL / 13 / 2609.38795
Recovering Off-Policy Supervision for Speculative Decoding
Abstract
Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.
cs.CL / 14 / 2609.38799
Overlap, Unique and Conflict: Can LLMs Extract What They Can Recognize?
Abstract
Understanding multi-perspective alternative narratives requires identifying how their information agrees, conflicts, or differs across sources. Existing work on cross-text relations largely focuses on categorizing relations between predefined text pairs, such as entailment or contradiction, rather than directly extracting such information from full narratives. To address this gap, we introduce Overlap-Unique-Conflict (OUC) extraction, a cross-narrative task that extracts all overlapping, conflicting, and unique clauses from two narratives. To support this study, we construct a benchmark of approximately 22K narrative pairs and 140K OUC instances spanning factual, argumentative, and political discourse. Evaluating 14 open-source LLMs (0.6B-35B), we find that unique information is far easier to extract than overlap and conflict: the strongest model, Gemma-4-31B, reaches only 61.13% F1-score on overlap and 48.58% on conflict, against more than 75% on unique. Further diagnostic analysis reveals that this difficulty does not stem from relation recognition alone, but rather from a failure to pair and extract the corresponding clauses from full narratives, especially in smaller models. Nevertheless, learning these extractions with task-specific supervision narrows the gap considerably: a fine-tuned Qwen-3-8B gains 15-28% absolute over its baseline and surpasses models roughly four times its size (e.g., Qwen-3.6-35B) on several tasks. Even so, overlap and conflict remain well below satisfactory, leaving cross-narrative clause extraction an open challenge.
cs.CL / 15 / 2609.38812
Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification
Abstract
Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifies the first complete solution in each trajectory, determines whether it is objectively correct, and uses this ground truth to quantify the agent's subsequent verification and recovery behavior. Applying it to ten terminal agents on TerminalBench2.1, we find that verification is nearly universal after a complete candidate is formed, yet only 61.43\% of incorrect candidates are detected and only 49.36\% of detected errors are successfully repaired. These results show that the main weakness in self-verification lies not in initiating verification, but in detecting and repairing errors. Motivated by these findings, we propose Student-Conditioned Verification Distillation (SCVD), which lets the student first produce a candidate solution and distills a stronger teacher's subsequent verification and recovery from the same interaction context. Across three Qwen3.5 backbones, SCVD improves \textsc{Pass@1} on TerminalBench2.1 by 9.74--16.85 percentage points over the corresponding base models and by 4.49--8.61 points over the standard full-trajectory distillation, while avoiding the pronounced out-of-distribution degradation of full-trajectory distillation on SWE-bench Verified.
cs.CL / 16 / 2609.38820
BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects
Abstract
With the rapid growth of Arabic NLP, several models, datasets and benchmarks have been reported. This paper asks whether approaches developed for majority languages like English can be adapted to Arabic tasks. We adapt an English aspect-based sentiment analysis framework to Arabic classification tasks and present the adaptation as BARRAC: Brainstorming Alignment and Replaced Representation learning for ArabiC tasks. BARRAC replaces consumer-review attribute pools with Arabic linguistic devices and markers for dialectal sentiment, sarcasm, and dialect identification, and replaces noisy self-training with two-stage training. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro-F1 of 63.93\%, outperforming the best few-label SOTA by 3\%, and outperforming GPT-4o on four out of five tasks. Error analysis provides insights into remaining challenges. These results demonstrate that adapting task-specific approaches is a promising direction for Arabic NLP alongside adapting models, datasets and benchmarks.
cs.CL / 17 / 2609.38831
Forging LLM Authorship Fingerprints with Targeted Rewriting
Abstract
Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this problem as targeted fingerprint transfer: rewriting one model's output so that attribution classifiers assign it to a chosen target model. We study summarization, where different models receive the same document and express the same underlying content, providing a controlled setting for conditional generation. We introduce ForgePrint, a search-then-distil framework that first searches for rewrites that move attribution toward a target fingerprint, then distils the selected rewrites into a one-pass 4B Student model. On CNN/DM, the Student reaches 70.2% target success rate, outperforming both its Teacher (54.1%) and the strongest of six published rewriting baselines (39.3%), against held-out classifiers that are never queried by the attack. It also reaches 68.3% target success when transferring summaries from an open model toward chosen commercial models. These results show that fingerprint detectability should not be conflated with source authenticity, and that text-only attribution can provide misleading evidence of model identity under targeted rewriting, even when it is accurate on unmodified text.
cs.CL / 18 / 2609.38832
Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
Abstract
Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parameter scaling can directly enable efficient and effective context scaling. We introduce NAMOH, an architecture-native sparse attention mechanism that activates $K$ of $H$ heads per token. Each head retains only its assigned tokens and performs causal attention within this subsequence. Head selection thus jointly determines active parameters and available context without scanning the full history. Under balanced assignments, increasing $H$ at fixed $K$ shortens head histories and reduces per-token key-value (KV) access without increasing total KV storage. We further support head-relative rotary position embeddings to shorten positional spans within routed subsequences, aiming to mitigate position-induced attention noise. Experiments show that NAMOH can outperform fully activated models with the same total parameters, while enabling more efficient long-context inference than smaller dense models with matched active parameter counts. It remains compatible with GQA and existing sparse attention mechanisms. We hope this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.
cs.CL / 19 / 2609.38851
Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis
Abstract
End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54\% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84\% of the gain from supplying all prerequisites.
cs.CL / 20 / 2609.38861
TRACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA
Abstract
Finding a relevant paper is not the same as producing a verifiable answer from it. LitTraceQA requires canonical paper identifiers, exact evidence at the page or object level, and typed answers that match the evaluator. We call the separation between source access and scorer-visible correctness the grounding contract gap. TRACE - Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction - addresses this gap with target-grouped retrieval, independent typed evidence localization, multimodal table extraction, schema-driven table construction, and fail-closed validation. It indexes 27,487 papers through passage, object, alias, citation, and dense representations while retaining the question target behind each signal. For tables, TRACE predicts the observation unit before extracting values and assembles rows with evaluator-compatible key normalization. Our audited selected clean-track artifact scores 0.760613 on the official 71-question test set, including 0.9728 paper F1, 0.6847 evidence F1, 0.9800 multiple-choice accuracy, 0.5423 table-row F1, and 0.3508 macro cell accuracy. On 11 public-development table records, a clean baseline and coordinate-aware visual fill obtain row F1 of 0.291 and 0.411, respectively; this diagnostic comparison includes fallback outputs and is not an official-test claim. Remaining errors chiefly concern locator, observation-unit, row-key, and source-value identity.
cs.CL / 21 / 2609.38923
GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
Abstract
Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.
cs.CL / 22 / 2609.38976
Fairness Beyond a Single Run: Training-Seed Variability in Speech LLM Adaptation
Abstract
Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector and LoRA adapters of a speech LLM at five audio compression factors and six random seeds, holding the encoder, base decoder, data and decoding fixed, and evaluate every run on Common Voice and Fair-Speech. At 460 h of clean LibriSpeech, the seed moves fairness metrics more than compression does on most demographic axes. A balanced 3x3 decomposition attributes 85.3% of the variation in Fair-Speech ethnicity normalized gap to the seed against 8.3% to compression (p = 0.009), though compression explains more on age and gender. Held-out LibriSpeech word error rate spreads by 0.04 points across those seeds while Common Voice spreads by 8.57, so these are not failed runs, and the effect survives controlling for accuracy and dropout. Scaling and diversifying the adaptation set to 960 h damps the effect but does not remove it. On Fair-Speech ethnicity, two single-run systems must differ by more than 0.30 in normalized gap to exceed seed variability.
cs.CL / 23 / 2609.38995
When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged information. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math-related tokens, and that pointwise clipping of the forward KL objective stabilizes training. Pointwise clipping caps each vocabulary-wise forward KL term at a fixed threshold before summing over the vocabulary. Follow-up studies have adopted this clipping, but its effect on training has not been directly examined. In matched training runs differing only in whether clipping is applied, we observe that clipped runs produce substantially more repetitions that persist to the end of the response than their unclipped counterparts. We trace this failure to the clipped objective. We prove that the clipped objective can fail to correct the student toward the teacher and can instead push clipped and unclipped token probabilities away from its teacher. Our training runs agree with this analysis: inside repetitions, the clipped student places less probability than its teacher on leaving the repetition, and more on continuing it, whereas the unclipped runs stay close to their teachers.
cs.CL / 24 / 2609.38997
Settle: Learning When to Stop Reasoning
Abstract
Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwen3-4B, Settle reduces token count by 40% with a 0.5-percentage-point decrease in accuracy. It gains 6.16 percentage points over supervised fine-tuning on the same traces shortened at their first stable answer, at nearly identical token counts. Its stopping score predicts whether a correct answer will remain correct. Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods.
cs.CL / 25 / 2609.39001
The Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype
Abstract
LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across languages. We measure this token premium on seven tokenizers of widely used 2026 models (OpenAI o200k, Llama 3, Qwen3, DeepSeek V3/V4, Gemma 3, Mistral Tekken, and the Claude generation-5 tokenizer via Anthropic's counting API) on NTREX-128 (124 non-English reference translations) and on the Universal Declaration of Human Rights for regional languages. French requires 31% to 58% more tokens than English, whereas Simplified Chinese ranges from 5% fewer to 40% more and is cheaper than French on six of the seven tokenizers. Regional and overseas languages of France pay roughly 1.6 to 3.3 times the English count. We discuss how history re-sending, tiered pricing and fixed context windows amplify the absolute gap in agentic use. In a controlled experiment (BPE, Europarl, 50k vocabulary), adding French to tokenizer training data quickly reduces the premium, with diminishing returns and a growing cost for English. Finally, we present Baracoda FR v1.2, a byte-level BPE prototype with Tekken's vocabulary size. On a final test of six corpora never consulted during design, with a protocol declared fixed beforehand, it uses 11.5% fewer tokens than Tekken on French and 3.7% fewer on English; results hold after removing test sentences overlapping the training data and with an equal ordinary-token budget. It is worse on other languages and, at comparable vocabulary size, does not outperform CroissantLLM. These are segmentation results only; effects on model quality and task cost remain to be shown.
cs.CL / 26 / 2609.39013
Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem
Abstract
EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and evaluates that expression in local code. Saved intermediate results support inspection of failures. A separate public-validation run achieved 92.17% answer accuracy and 1.00 evidence F1. The configurations and metrics differ, so these scores are not a controlled comparison. Our manual, post-hoc analysis is descriptive: in one inspected case, optical character recognition (OCR) and block grouping merged the relevant passage into another block, and the system answered from unrelated text. An exploratory study of reading page images on 100 documents returned evidence identifiers for only 22 documents. These descriptive findings motivate further evaluation; they do not establish the causes of the overall score.
cs.CL / 27 / 2609.39027
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Abstract
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.
cs.CL / 28 / 2609.39069
CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search
Abstract
Test-time reasoning systems often respond to failure by restarting or revising the latest step, even when an earlier decision caused the error. We introduce CORE, a search controller that requests a certified conflict core from a verifier, backjumps to the latest decision in that core, and caches the conflict to avoid repeating it. Under sound verification, finite branching and depth, and exhaustive proposals, the uncapped search is complete and never prunes a valid solution. On 2,000 planted graph-coloring instances with matched proposals and an exact verifier, CORE reduces median verifier calls by 39.8% at 30 variables and 35.0% at 36 variables relative to chronological repair; caching further improves on backjumping alone. Across five reasoning tasks, CORE achieves 75.9% mean success with Qwen2.5-7B-Instruct and 84.2% with Qwen3-8B, compared with 72.5% and 81.8% for Tree of Thoughts. It also uses fewer verifier calls and generated tokens on both backbones. These results show the value of using certified failure explanations to direct language-model search.
cs.CL / 29 / 2609.39071
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
Abstract
Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.
cs.CL / 30 / 2609.39102
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Abstract
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
cs.CL / 31 / 2609.39111
Bongard: Training Machine Intuition
Abstract
Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate step. We introduce Bongard, an open-weight System One model that treats machine intuition as an independent capability to design and train. A T5Gemma 2 4B-4B encoder-decoder separates reading the evidence from making judgments. The encoder reads the state bidirectionally together with the question instructions, and separate decoder branches share this encoding, so many judgments about the same situation require only one reading of the state. A trained head returns probabilities over the supplied candidates without generating text. Training proceeds in three stages, from supervised judgments to semantic relationships to action outcomes, and each stage updates all 7.09 billion trainable parameters on one Blackwell GPU. Joint-embedding post-training raises accuracy on held-out rephrasings from 75.7% to 85.9%. A sandbox stage then learns outcome distributions from action rollouts and exact oracles, raising accuracy on a frozen sandbox panel from 50.6% to 64.8%. On DecisionBench, the final model reaches 78.05% accuracy over 23,900 decisions and ranks fourth of 61 systems in the public comparison. On one RTX PRO 6000, its median latency is 36 ms for short requests, and 32 questions about one state take 221 ms. Bongard demonstrates that machine intuition can be systematically trained via representation learning and outcome feedback, providing an open, efficient alternative for high-throughput decision workloads.
cs.CL / 32 / 2609.39118
Diagnosing On-Policy Self-Distillation for Reasoning Language Models
Abstract
On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavior in language reasoning remains unclear, with reported outcomes ranging from modest gains to behavioral collapse. In this work, we diagnose OPSD for mathematical reasoning across models spanning 0.6B--8B parameters. We conduct controlled experiments and token-level analyses to fully delve into OPSD. We point out that teacher's signal is shaped by reasoning-mode alignment and the complete teacher prefix, rather than by privileged semantics alone. OPSD improves reasoning only in narrow compatibility regimes. Otherwise, it produces ineffective length growth, stable degradation, or behavioral collapse. Token-level analysis shows that teacher's signal is not stable and does not predict downstream performance. Based on these results, we argue that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method.
cs.CL / 33 / 2609.39154
DAGent: Evaluate-then-Grow Planning for Deep Research Agents
Abstract
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
cs.CL / 34 / 2609.39189
ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations
Abstract
Trustworthy Legal AI requires systems that can answer legal questions while grounding their responses in authoritative sources. However, existing Vietnamese legal benchmarks provide limited coverage of real-world legal consultations. We introduce \textbf{ViLegalExpert}, a large-scale benchmark constructed from authentic citizen--lawyer consultations, containing over \textbf{172K} questions across \textbf{34 legal domains}, together with professional answers and expert-verified legal evidence. ViLegalExpert supports legal information retrieval, extractive QA, and abstractive QA. Experiments with representative retrieval methods and language models reveal substantial challenges in evidence retrieval and grounded answer generation. While pretrained models perform strongly on QA, hybrid retrieval achieves the best retrieval performance. These results demonstrate the difficulty of mapping naturally expressed legal questions to authoritative provisions and establish ViLegalExpert as a challenging benchmark for reliable Vietnamese Legal AI.
cs.CL / 35 / 2609.39238
4MT-VLM: How Coarse Is a VLMs Cognitive Map?
Abstract
An agent that moves must recognise a place from a viewpoint it has never seen. We introduce 4MT-VLM, a dataset of procedurally generated landscapes, each rendered across five stimulus modes that remove appearance cues while holding layout fixed: shape and colour, shape only, colour only, bare terrain peaks with no objects, and a valley viewpoint that puts the peaks on the horizon. The last condition is commonly used in clinics to probe hippocampal function in human patients. We test this benchmark across sixteen different open and closed-source models and report 4AFC performance, a measure which is also used to grade human participants. We observe that models identify a place from the studied viewpoint but lose it once the camera moves, dropping below the 25% chance level at 135° where a human observer scores 85%. Frontier models (Gemini 3.8 Flash, GPT-5.6) answer only 39% and 31% of rotated trials correctly, recovering to 85% and 55% only when distractors are moved more than 30 meters apart. Our benchmark demonstrates that while current VLMs possess rudimentary cognitive maps, their spatial resolution remains fundamentally too coarse to maintain a stable, 3D understanding of the world once the viewpoint changes.
cs.CL / 36 / 2609.39263
Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout
Abstract
A concept subspace's effect on model behavior does not establish how it relates to the output readout. We introduce a two-sided geometric diagnostic that measures an extracted subspace's overlap with the dominant right-singular directions of the unembedding matrix, evaluated against output-oriented positive controls. Given an extracted basis, the raw diagnostic requires only model weights. Our testbed is the Format-Agnostic Reasoning Subspace (FARS), a ten-dimensional basis extracted from eighteen reasoning concepts expressed in six surface forms. Across nine rank-matched estimators and twenty-six models, four activation-derived concept estimators carry only 0.38--0.80% mean energy in the top-ten readout span. Final-layer PCA carries 3.56%, exceeding FARS in 25 of 26 models. A same-layer next-token control, evaluated using a fitted linear translator for depth matching, carries approximately thirteen times more energy than FARS, with separation in all 25 tested models. Re-extracting FARS on ten disjoint concepts yields 62--100% cross-format retrieval across twenty-four generative models, demonstrating transfer of the extraction procedure rather than a fixed basis. A complementary four-model, three-seed intervention study finds model-dependent source-directed effects that remain well below full-vector replacement. Together, the geometry and intervention controls distinguish concept structure from dominant readout directions while limiting claims of causal sufficiency.
cs.CL / 37 / 2609.39358
Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
Abstract
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and llama.cpp that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
cs.CL / 38 / 2609.39365
Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts
Abstract
Continual alignment requires LLMs to adapt to new requirements without forgetting previously acquired behaviors. Natural-language instructions are flexible and composable but offer only indirect control, whereas post-training provides stronger adaptation at the cost of repeated parameter updates. We introduce Ready2Blend, which combines the flexibility of natural language with learned alignment. AlignFormer maps each requirement to a fixed-length alignment prompt stored in a modular prompt bank, while the backbone and prior prompts remain frozen. Composability regularization transfers the semantic geometry of textual requirements into prompt space, enabling inference-time blending and reweighting. Across two practical continual alignment settings, Ready2Blend is the only frozen-backbone method that matches post-training-based alignment methods, reaching $93.1$-$98.5\%$ of a joint-training reference with competitive retention, while requiring only a few prompt tokens and up to $4.3\times$ less training time. Its modular design further enables weighted personalization and order-free composition without retraining. Code will be released upon acceptance.
cs.CL / 39 / 2609.39368
Making Grid Beam Search Less Greedy
Abstract
A common formalism for constraining the output of autoregressive text generation models involves lexical constraints, words or phrases which are required to occur in the generated text. DFA-constrained beam search and grid beam search are two widely used paradigms for decoding from autoregressive models while enforcing lexical constraints. As the former approach requires a number of forward passes exponential in the number of constraint tokens, it is often dispreferred to the latter, which requires only linearly many forward calls. However, while grid beam search achieves an exponential speedup, it does so in a manner which does not treat all of the constraints equally. In this paper, we demonstrate that grid beam search is biased to incorporate easier-to-satisfy constraints first, leaving harder constraints to the end of the sequence. This contrasts with DFA-constrained beam search, which exhibits no such bias. To address this shortcoming, we propose fair grid beam search, a modification to grid beam search which avoids this bias while still requiring only linearly many forward passes. Experimentally, we confirm grid beam search's bias on two constrained generation tasks, finding significant differences in how it orders constraint tokens as compared to DFA-constrained beam search and fair grid beam search. Furthermore, we find that fair grid beam search not only fixes grid beam search's bias, but finds higher-probability strings in the process.
cs.CL / 40 / 2609.39369
Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer
Abstract
Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.
cs.CL / 41 / 2609.39385
TTLab at Daleel 2026: STAR-Ar, Sequence Tagging for Argument Recognition in Arabic
Abstract
Argument Mining (AM) is a critical NLP task that remains significantly under-resourced in Arabic. This paper presents $\testtt{STAR-Ar}$, a BERT-BiLSTM-CRF architecture for argument discourse detection and classification, as our system for Daleel 2026, the inaugural Arabic argument mining shared task. The task requires the identification and classification of argumentative discourse units (ADUs) in debate and editorial texts.We jointly model these two objectives as a token-level sequence labeling task using a BERT-BiLSTM-CRF architecture that combines contextual transformer embeddings with structural transition constraints to support accurate span detection. $\testtt{STAR-Ar}$ achieves an F1-score of 72.69 on validation and 73.7 on test data. Our domain-specific analysis shows that models trained exclusively on editorials underperform those trained on debates, a disparity we primarily attribute to the smaller size of the editorial dataset. The code for $\testtt{STAR-Ar}$ is available at ${\href{https://github.com/ENTAILab/daleel_2026_Arabic-Argumentative-Discourse-Mining}{\faGithub~TTLab at Daleel 2026}}$
cs.CL / 42 / 2609.39446
DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements
Abstract
Existing full-duplex speech benchmarks cover only subsets of real-time interaction behaviors, often under limited contextual conditions. We introduce DuplexAct-Bench, a bilingual benchmark that systematically covers six complementary behaviors, from interruption and yielding to proactive initiation, active silence, and backchanneling, across Pre-session, In-session, and No-explicit conditions. Across 1,290 English and Chinese streaming trials, we evaluate 12 full-duplex speech systems on both Timing and Content. Results reveal substantial variation across behaviors, conditions, and systems, as well as frequent mismatches between semantic quality and behavioral timing. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds. Project page: https://alitaxky.icu/DuplexAct-Bench/
cs.CL / 43 / 2609.39447
Synthetic Data Characterization via Training Dynamics
Abstract
Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.
cs.CL / 44 / 2609.39460
Right-Wing Rock or Just Rock? A Computational Linguistic Analysis of Frei.Wild
Abstract
Rechtsrock is a subgenre of rock music that spreads right-wing ideology, often instrumentalized to recruit adolescents into the radical scene. Monitoring institutions counteract this by manually examining and, in some cases, banning extremist content; however, there are border cases that evade regulation. We present a study aimed at determining whether such a case, the band Frei.Wild, should be classified as politically right-leaning or as part of the general German rock genre. We sampled a German rock dataset and created a corpus for right-wing rock to use as reference in this analysis and found that we can confirm the intuitions from previous investigations that Frei.Wild successfully maintains an ambiguity with regard to their political affiliation. However, the tendency is towards the right-wing spectrum. Lexical analyses reveal nationalistic narratives and two high-performing classifiers (up to 97% ROC-AUC score) label more than half of their songs as right-wing extremist. Our analysis provides insight into how computational methods can improve the process of identifying right-wing extremist tendencies in music, especially in borderline cases like Frei.Wild. The code and data are made available for future research.
cs.CL / 45 / 2609.39514
Spike-driven Vision-Language-Action Model
Abstract
Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top-$k$ winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.
cs.CL / 46 / 2609.39572
Compact Language, Complex Model Shifts: How and Where Ambiguity and Underspecification Affect LLMs
Abstract
We analyze how lexical ambiguity and underspecification affect language model training. We create artificial homonyms and artificial hypernyms as pseudowords and analyze the generative performance of language models as they are trained with increasing amounts of these ambiguous or underspecified pseudoword types. We further analyze whether the models disambiguate ambiguous or underspecified statements and provide a first mechanistic account of how ambiguity and disambiguation are represented internally. Our main results show that both ambiguity and underspecification increase model performance in ways that scale with their influence on the language's type-token ratio. However, the accuracy of generating sequences containing ambiguous words or their synonyms decreases compared to other texts. We also show that internal representations of pseudowords reflect disambiguation of pseudo-homonyms, but underspecification of pseudo-hypernyms is maintained during the generative process.
cs.CL / 47 / 2609.39578
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
Abstract
Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.
cs.CL / 48 / 2609.39608
Is This Evidence Decision-Critical? Learning to Verify Rule-Governed Decisions
Abstract
Rule-based reasoning, as in eligibility checks and contract reviews, requires language models to assess evidence against individual conditions and combine their judgments under explicit rules. Errors in evidence assessment can leave a decision unchanged, but misinterpreting or overlooking decision-critical evidence can reverse it. Identifying such evidence allows more capable models to focus on checking the corresponding condition judgments, supporting accurate and safe decisions. Recognizing the evidence's criticality requires understanding how evidence affects a condition judgment and how that judgment affects the decision. To achieve the goal, we propose a INTERvention-based imPACT learning framework (InterPact), which enables counterfactual verification of evidence criticality in rule-governed decisions. Specifically, its evidence intervention constructor generates training pairs for a propagation verifier by editing case facts with a frozen language model while holding rules and non-target conditions fixed. Human-reviewed labels record the resulting condition and decision changes, while complete state-to-decision mappings supervise consequences beyond the observed edit. During training, the verifier weights learned conditional decision predictions by evidence-based condition probabilities through a fixed composition operation, propagating decision-change supervision into the base model. At inference, the trained base model directly judges criticality from the original case and target evidence, without human or stronger-model supervision. On single-case evidence criticality verification over adapted rule-governed decision cases, InterPact achieves 68.28% accuracy, outperforming all six baselines. These results support learned decision sensitivity as a basis for prioritizing evidence checks.
cs.CL / 49 / 2609.39640
Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
Abstract
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.
cs.CL / 50 / 2609.39687
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.
cs.CL / 51 / 2609.39710
Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims
Abstract
Scientific abstracts mix contributions with background, motivation, and meta-language, so tools that read them as-is cannot separate what a field produces from what it discusses. We present Drift Inspector, an open-source system for measuring and exploring how a research field changes over time at the level of Atomic Contribution Claims (ACCs): decontextualized, contribution-bearing propositions an LLM extracts from each abstract before analysis. The system clusters these claims across years into an interactive map where every trend traces back to the claims and papers behind it. Applied to six years of EMNLP, it shows the field shifting away from classic NLP tasks toward LLM-era capabilities such as reasoning and multimodality -- a movement that keyword or whole-abstract counts blur. The released data extend beyond EMNLP: the same pipeline has processed the full ACL Anthology (346k claims, 80k abstracts, 423 venues). Extraction is human-validated and clustering checked against an external manually constructed taxonomy.
cs.CL / 52 / 2609.39740
LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation
Abstract
Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.
cs.CL / 53 / 2609.39765
MemCodex: Self-Programming Hierarchical Memory for Language Agents
Abstract
Agent memory faces heterogeneous access needs: a single-hop question may require one piece of evidence, whereas a multi-hop question must combine evidence from multiple sources. Predefined memory workflows cannot adapt to these varying needs. Recent adaptive methods search or learn over memory components and their compositions, but the design space itself remains predefined. We introduce MemCodex, a self-evolving hierarchical memory system that organizes experience into executable memory programs for summaries, relational knowledge, reusable skills, and latent memory. Open-ended program evolution searches the open design space of layer programs by rewriting how each layer is constructed, indexed, retrieved, and routed, thereby adapting both within-layer implementations and cross-layer composition. At query time, reads traverse the hierarchy from coarse to fine and stop once sufficient evidence is found, descending to the original history when needed. We further develop MemArena, a unified runtime that places heterogeneous data and memory systems behind a common interface. MemCodex improves average task success by 10.1% relative to the strongest adaptive-memory baseline, while using 3.4x fewer context tokens and achieving 2.1x faster inference.
cs.CL / 54 / 2609.39807
Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations
Abstract
Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, particularly when correct and incorrect answers are evaluated under the same persona prompt. To investigate why, we construct three novel confounder datasets in which truth is anti-correlated with a potential confounding concept. Our experiments reveal that many existing probes strongly track concepts that are spuriously correlated with truth in their training data, such as instruction compliance or response likelihood. Based on these findings, we introduce a simple linear probe that achieves the strongest overall performance on both the persona and confounder stress tests. Our results suggest that current lie detection probes are far from reliable and highlight the need for training data in which truth is decorrelated from confounding concepts.
cs.CL / 55 / 2609.39827
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
Abstract
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
cs.CL / 56 / 2609.39846
When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models
Abstract
We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a kindergarten student, one might expect its performance on a mathematics benchmark to reflect kindergarten-level ability rather than expert-level proficiency in solving calculus problems. We introduce RoleCapBench, a curriculum-grounded benchmark for evaluating RCL across six educational roles and four assessment levels spanning elementary school through A-level, and use it to evaluate three open-weight reasoning models. We find that although the models can generate stylistically convincing in-role responses, they consistently fail to align their underlying capabilities with their assigned roles. Naive role prompting yields strong role-voice scores of 1.218--1.389 while retaining above-role accuracy of 0.811--0.898. RCL persists across a range of prompting conditions, including prompts that explicitly instruct the model to match the role's capability level. To mitigate this problem, we propose Injection, an inference-time intervention that combines explicit, role-specific capability guidelines with a guiding prefilled response prefix. Injection improves role-capability alignment across models, reducing above-role accuracy by up to 0.562 while preserving in-role accuracy with a marginal drop of less than 0.058 across most models. All artifacts, including scripts and evaluation data, will be released upon acceptance.
cs.CL / 57 / 2609.39927
AdaGEPA: Adaptive Feedback Allocation for Reflective Prompt Optimization
Abstract
Prompt optimization improves the performance of language-model systems on downstream tasks by refining their prompts. Classical methods evaluate prompts on task examples and use the resulting feedback to guide prompt revisions through reflection. However, when feedback selection does not account for the prompt's weaknesses, these revisions may improve performance on selected examples without yielding broader task improvements. To address this issue, we propose AdaGEPA, an adaptive feedback-allocation method that uses the prompt's performance and task structure to select examples for the next prompt revision. Our method replaces at most one example in each feedback minibatch to target an identified weakness while preserving the remaining feedback context. Across our main experiments on six downstream benchmarks, AdaGEPA achieves higher mean validation scores than non-adaptive feedback selection under matched rollout budgets. AdaGEPA also finds high-performing prompts earlier across several tasks. In the initial Schema-Guided Dialogue (SGD) study, its half-budget prompts outperform the non-adaptive baseline's full-budget prompts in joint goal accuracy on new dialogues from services seen and unseen during search. Overall, our findings highlight the potential of adaptive feedback allocation to improve both the effectiveness and rollout-budget efficiency of reflective prompt optimization.
cs.CL / 58 / 2609.39938
LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
Abstract
Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
cs.CL / 59 / 2609.39972
UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding
Abstract
Speculative decoding accelerates language model inference by verifying multiple draft tokens in a single target-model pass. Recent parallel drafters have achieved breakthrough performance in frontier production models, but their effectiveness deteriorates as the entropy of target distributions increases due to insufficient draft diversity. To overcome this bottleneck without sacrificing parallelism, we introduce UBTree, a parallel drafter that couples a Unigram proposer with a Bigram selector to construct drafting Trees. The unigram proposer is trained with the standard cross-entropy objective to generate candidate tokens independently for each position, while a lightweight bigram selector predicts transition scores between adjacent candidate pairs. Unlike the proposer, the selector is trained with a renormalized KL objective on high-temperature data. This tree-native training broadens the supervision beyond the greedy path, encouraging plausible alternative branches that improve the chance of accepting additional tokens during tree verification. Across seven standardized benchmarks with Qwen3-4B and Qwen3-8B, UBTree achieves an average speedup of $5.84$--$6.94\times$ over autoregressive decoding and outperforms DARTree in all 28 comparisons. Production-scale evaluation further demonstrates UBTree's advantage over frontier baselines such as DSpark.
cs.CL / 60 / 2609.39975
Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
Abstract
This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical semantic question answering. b) Task Synergy14 on question answering for developing biomedical top- ics. c) Task MultiClinSum-2 on multilingual clinical summarization. d) Task BioNNE-R on extracting relations between nested named entities in Russian and English. e) Task ELCardioCC on clinical coding in cardiology. f) Task GutBrainIE on gut-brain interplay information extrac- tion. Across these six tasks, 87 distinct teams participated, submitting more than 1000 runs overall. As in previous editions, several submissions reached competitive performance, reflecting the continued progress of state-of-the-art methods across biomedical language processing tasks.
cs.CL / 61 / 2609.39982
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Abstract
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.
cs.CL / 62 / 2609.40035
OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search
Abstract
The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This needs no action-distribution correction, although branching changes state visitation. Our Branch Aggregation Lemma shows that branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansion states using estimated performance differences. Under deterministic dynamics, exact values, and max-backup advantages, the induced search policy's expected return improves monotonically with the budget. We bound the gradient bias from adaptive expansion and show that max backup assigns prefix credit to actions leading to better discovered suffixes. Against a finite chain reference, TTPG's measured bias stays near its no-branching level, while NaivePG's bias grows from 0.1251 to 0.4884. At matched budgets, reward- and value-guided OPTS improve correct-answer coverage and majority-vote accuracy over independent sampling. At matched branch counts, OPTS + TTPG gains coverage with a modest bias increase relative to Fixed-branch + TTPG. Under matched interaction or rollout budgets, OPTS-TTPO improves MuJoCo tail returns over PPO by up to 28.6%, achieves a 34-22-1 win-loss-tie record against PPO on Atari-57 under the last-100-log mean-return metric, and improves micro-averaged avg@32 and pass@32 over PPO across all four Qwen3 models.
cs.CL / 63 / 2609.40041
MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs
Abstract
We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGhana-ST is an ongoing annotation effort; the experiments here use a fixed subset of about 16.1 hours of paired speech and English translations. The audio is curated from two existing Ghanaian speech resources. Unlike in those resources, the English translations are produced directly from audio by 37 native-speaker annotators and include verbal and non-verbal event annotations. Using Whisper-small, we compare monolingual and multilingual training under severe data scarcity, reporting means over three seeds. Flat multilingual training benefits no variety in this regime. Ga and Twi are unchanged within seed variance (+0.51 and +0.06 BLEU against monolingual standard deviations of 1.63 and 2.20), while Ewe declines by 6.99 BLEU and Fante by 5.11. The degrading varieties are Ewe, which is linguistically distinct and drawn from a different source corpus, and Fante, the least-resourced. Comparing empirical cross-lingual transfer with typology-based similarity, we find that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training. We also report a methodological finding. An earlier single-run analysis found positive transfer for three of four varieties; this did not survive replication across seeds. For Ga and Twi, monolingual baselines trained on 1.6 to 6.2 hours of audio have seed standard deviations roughly five and thirty times those of the multilingual models (0.35 and 0.07 BLEU). When the monolingual condition is noisier, a single-run comparison can show apparent transfer of this size from seed variation alone. We release MGhana-ST to support research on African language speech technology and low-resource speech translation.
cs.CL / 64 / 2609.40064
From Tweets to Trades: Analyzing the Influence of Public Mood over Stock Market Performance in Turkiye
Abstract
Purpose: This study examines whether domain-specific public mood is associated with stock-market dynamics and whether these relationships vary across communication domains and market conditions. It distinguishes public mood from investor sentiment and investigates whether heterogeneous sources of public communication exhibit different relationships with market behaviour. Design: The study analyses 610,422 posts published by 176 curated X accounts between January 2022 and December 2023, covering Politics and Government, Economy and Finance, and Media and Society. Posts are classified using fine-tuned Turkish transformer models under three domain-specific and one pooled regime. Public mood measures are constructed at daily, weekly, and monthly frequencies and examined alongside BIST100 and BIST30 market measures using correlation, Granger causality, vector autoregression, and impulse response analyses across the full period and selected market conditions. Findings: Public mood is not associated with the direction of stock-market returns but is associated with the magnitude of price movements, particularly for Media and Society and pooled communication. These relationships become stronger at longer aggregation frequencies. Predictive relationships are concentrated in Economy and Finance communication, while their magnitude and direction vary across market conditions, particularly during the 2023 election period. The pooled measure largely reflects the most active communication domain. Originality: The study contributes to behavioral-finance research by incorporating communication - domain heterogeneity into the analysis of public mood and market dynamics. It also demonstrates how aggregating heterogeneous sources can obscure domain-specific relationships between public communication and financial markets.
cs.CL / 65 / 2609.40097
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
Abstract
Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.
cs.CL / 66 / 2609.40108
OverdoseMoE: A Multi-Expert Framework for Opioid Overdose Risk Prediction
Abstract
Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Here, we investigate diagnosis-specific adaptation for 180-day opioid overdose risk prediction from patients' preceding one-year longitudinal ICD histories. We develop OODMAMBA and OODQWEN through continued pretraining on longitudinal diagnostic sequences followed by task-specific fine-tuning. Building on the stronger Qwen-based predictors, we further propose OVERDOSEMOE, a multi-expert framework that integrates models of different scales using complementary expert-weighting strategies. Diagnosis-specific adaptation consistently improved predictive performance over general-purpose language-model baselines, with OODQWEN achieving an AUPRC of 24.47 and an AUROC of 68.56. OVERDOSEMOE further improved discrimination and precision, achieving an AUPRC of 25.17 and an AUROC of 69.49 while outperforming the strongest single-model baselines. Among patients ranked in the top 5% of predicted risk, OVERDOSEMOE identified substantially enriched overdose risk, achieving a PPV of 25.38% while retaining meaningful recall. Evaluation on an independent MIMIC-IV cohort further demonstrated cross-cohort robustness, with complementary weighting strategies showing advantages across different performance measures. These findings demonstrate that diagnosis-specific language-model adaptation combined with multi-expert integration can improve opioid overdose risk stratification and support more robust prediction across heterogeneous electronic health record populations.
cs.CL / 67 / 2609.40118
Persistent Context Graphs for Efficient Memory Compaction in LLM Agents
Abstract
As LLM capabilities advance, agents are tackling increasingly complex tasks over longer horizons. Their growing interaction histories make memory compaction essential for staying within context windows and reducing prefill cost. Existing methods summarize the history or compress its KV cache, often adding model computation to preserve information for future requests. A new user request can change which history matters, but reassessing that history with the model requires re-encoding it if the KV cache has expired. Past attention provides signals of historical importance and dependencies between messages, while relevance to the current task must be assessed using the new user request. We introduce ReCAP, a memory compaction method that stores attention-derived importance scores and dependency links in a lightweight, persistent context graph. For each new request, ReCAP combines stored importance with relevance cues from the request and follows dependency links to select messages and their supporting context, without additional model calls for selection. Compared with Codex's default summarization-based compaction, ReCAP reduces estimated latency for compaction and cold restoration by approximately 95% on both Qwen3-Coder and gpt-oss. It also roughly halves the historical context per call on SWE-Together at comparable task quality and improves accuracy on the code tasks of Lost-in-Conversation over full history by 19.8 and 41.2 points.
cs.CL / 68 / 2609.40181
Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation
Abstract
We introduce Index-Translate, a multilingual translation model family that combines a shared multilingual foundation with specialized training for general translation, instruction following, speech translation, controlled dubbing, and long-document translation. It includes three model sizes, 2B, 9B, and 35B-A3B, and supports translation in 150 languages, with multilingual instruction following. Evaluations on general translation and complex translation instructions show that Index-Translate outperforms translation models of comparable size and achieves performance comparable to 100B-scale translation models and frontier models. Index-Echo provides end-to-end speech-to-text and speech-to-speech translation, outperforming existing end-to-end models and achieving performance comparable to frontier omni models. Index-Homura extends the family to syllable-controlled dubbing. Index-NativeLong introduces native long-document translation with a dedicated task formulation and benchmark. These capabilities support diverse translation tasks, including multilingual content production.
cs.CL / 69 / 2609.40185
Provably Tractable NFA-Constrained Language Generation via HMMs
Abstract
Constrained generation aims to sample from language models (LMs) conditioned on hard constraints. Existing constrained-generation techniques for nondeterministic finite automaton (NFA) constraints either distort the distribution or sacrifice efficiency. Theoretically, this task reduces to counting the length-$n$ sequences accepted by an NFA (#NFA), and the exact #NFA problem is #P-complete. Recent work has shown that #NFA admits a fully polynomial randomized approximation scheme (FPRAS). Inspired by this result, we propose NFA-LM, a polynomial-time engine for NFA-constrained generation with theoretical guarantees under mild assumptions. Experiments show that NFA-LM efficiently generates high-quality outputs with theoretically bounded approximation error.
cs.CL / 70 / 2609.40198
SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models
Abstract
Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (full), its concatenated information shards delivered together (concat), and incremental spoken disclosure across turns (sharded). We report final-answer accuracy for four commercial speech systems and LEGO, a proprietary speech pipeline developed internally by the SCBX Innovation Lab team with explicit conversational context management. Relative to concat, sharded accuracy decreases by 5.0-25.3 percentage points across the four commercial systems. LEGO achieves 77.5 percent accuracy in all three conditions, compared with 76.6 percent sharded accuracy for GPT-4o Realtime. The two single-turn baselines distinguish sensitivity to problem reformulation from the additional challenges introduced by incremental spoken interaction.
cs.CL / 71 / 2609.40236
Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports
Abstract
Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT reports, and which ingredients matter. Using a 2x2 design, we crossed two adaptation strategies (a discriminative classification head, CH; generative instruction fine-tuning, IFT) with two training-data sources (distillation of real GPT-4o-labeled reports; synthetic reports generated by GPT-4o from real exemplars), across five training sizes, benchmarked on 100 expert-adjudicated reports against GPT-4o and the un-tuned open-weight base. The distilled instruction-tuned model (DIFT) matched GPT-4o (macro-F1 0.845 vs 0.850; p = 1.000) and exceeded the base model by 0.178. The decisive factor was the training-data source, not the fine-tuning method: both synthetic-data models failed to exceed the un-tuned open-weight base at any training size and underperformed the distilled models across all acuity classes. Fine-tuning and inference fit within the memory envelope of a single 24 GB consumer GPU. For narrow, high-value clinical label-extraction tasks, distilling real reports, rather than generating synthetic ones, is what closes the gap to a hosted model, enabling a private, low-cost, version-stable on-premises alternative.
cs.CL / 72 / 2609.40286
Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
Abstract
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
cs.CL / 73 / 2609.40295
How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Abstract
Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.
cs.CL / 74 / 2609.38426
LoopVL: Recurrent Visual Intelligence
Abstract
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
cs.CL / 75 / 2609.39920
MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models
Abstract
Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.
cs.CL / 76 / 2609.39334
Taming Speculative Search for Test-Time Scaling in LLM Serving
Abstract
Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks such as mathematics and coding. To accelerate the exploration of reasoning paths, recent studies proposed speculative execution. However, we show that supporting speculative execution poses two unique challenges for LLM serving systems: (1) an explosion in the search space of candidate paths and (2) frequent, fine-grained verification tasks for candidates. To address these challenges, this paper proposes SpecScale, a serving system for efficient speculative execution. We introduce three techniques to reconcile the trade-off between latency and computational overhead: (1) early pruning of low-quality candidate paths, (2) deduplicating computation across redundant candidate paths, and (3) deferring fine-grained verification tasks. We evaluate SpecScale on challenging reasoning benchmarks, including MATH and Olympiad. Our results show that SpecScale significantly outperforms both non-speculative and recent speculative approaches, delivering substantial improvements in throughput and latency while preserving answer quality.
cs.CL / 77 / 2609.38371
TALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns
Abstract
Existing LLM-driven robot task planners rely on a taken-for-granted assumption of an ideal user whose instructions are clear, complete, and task-focused. However, when interacting with real-world users, especially those experiencing cognitive impairments, such as people living with dementia (PLWD), the planners often make mistakes and even pose physical safety risks. We proposed TALK-Dem (Talking Attributes and Linguistic Knowledge in Dementia), the first benchmark for evaluating LLM-driven robot task planning under dementia-associated verbal communication. TALK-Dem contains 4,800 instructions and covers five typical communication patterns, including Referential Imprecision, Object Substitution, Empty Speech, Topic Drift, and Intrusion, at three intensity levels. Experiments across six open-weight LLMs reveal a substantial robustness gap. Across communication patterns, open-weight models exhibited performance drops of up to 22.3 percentage points compared to ideal instructions. This revealed a critical gap and even danger for real-world applications, especially in assistive robotics, where locally deployable models are necessary due to privacy concerns and connectivity constraints. To mitigate this issue, we proposed the Context-Aware Retrieval from Experience (CARE) method, which retrieves relevant previously resolved tasks to provide task-specific interpretation and planning context. CARE generally outperformed standard prompting baselines across the six open-weight models, improving average task success by 18.1 percentage points over the vanilla prompt. These results highlighted the importance of both evaluating communication robustness and developing effective adaptation strategies for locally deployable assistive robots. The TALK-Dem dataset is publicly available at https://anonymous.4open.science/r/TALK-Dem-A6B3/.
cs.CL / 78 / 2609.38658
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Abstract
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
cs.CL / 79 / 2609.38887
VOSSA: Voiceprint Optimization for Streaming Speech Architectures
Abstract
Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.
cs.CL / 80 / 2609.38324
Multi-agent discussion gains less when dissent is withheld
Abstract
Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on whether discussion improves accuracy or leads to an incorrect consensus. Here, we introduce a parsimonious model that explains when discussion improves accuracy and when it ends in an incorrect consensus, built from four behaviors repeatedly observed in LLM agents: (1) withholding dissent, (2) internalizing a stated answer, (3) reconsidering after seeing dissent, and (4) correcting toward the correct answer. The model shows that discussion can overturn an incorrect initial majority only when the withholding rate $c$ is below a critical rate $c^* = γ/(γ+ a)$, set by the net correction rate $γ$ and the internalization rate $a$. We estimate these rates from conversation logs with a Bayesian method and place LLM teams relative to $c^*$. As the model predicts, the gain from discussion shrinks as withholding rises, across LLMs and on a hidden profile benchmark, HiddenBench, and MedEInst. Instructing agents not to withhold dissent increases this gain. Turning reasoning off also increases the gain, because reasoning raises the internalization rate $a$ and keeps agents from reconsidering a minority answer. These findings reconcile the conflicting reports and identify when discussion outperforms majority voting.
多智能体系统 (cs.MA)
11
cs.MA / 1 / 2609.38482
PANDA: A Decentralized Architecture with Flexible Orchestration for Scalable, Fault-Tolerant Multi-Agent Systems
Abstract
Existing architectures for LLM-based multi-agent systems (MAS) cannot reliably and efficiently solve multi-step tasks at scale: they struggle to support large numbers of agents and concurrent tasks, tolerate failures, govern agent interactions, and accommodate the diverse planning and execution patterns different tasks require. We present PANDA, a decentralized architecture that connects a large collective of heterogeneous, independently administered agents, letting them discover each other's capabilities and self-organize into small specialized teams per task. PANDA scales by decoupling collective communication from team communication, allowing agents to participate in multiple teams simultaneously, load-balancing tasks across the collective, and scheduling concurrent work within each agent. PANDA further separates the underlying architecture from the orchestration strategy, supporting three planning and execution patterns (star, chain, and mesh) that can be selected according to the structure and requirements of each task. PANDA detects infrastructure and orchestration failures and recovers affected tasks by dynamically replanning around failed components. Finally, to provide governance without a centralized service that would limit scalability, PANDA uses a web-of-trust model to constrain agent interactions to established trust relationships. We evaluate PANDA on the HotPotQA benchmark, demonstrating that it scales to thousands of agents, assembles teams in milliseconds, matches state-of-the-art accuracy at up to 8x the efficiency, and sustains 100% task completion under faults where existing systems fail.
cs.MA / 2 / 2609.38643
Recursive Organization Improvement: A Modeling Specification for Human--Agent Organizations
Abstract
Stronger AI agents do not automatically produce better organizations: teams must also learn which work arrangements to retain and when to reconsider them. We propose a modeling specification for recursive organization improvement and evaluate it through an executable checker, a public-record mapping, and controlled simulation. The specification connects actor-visible histories, organizational memory, decision rights, and evidence-carrying change contracts. The mechanism study crosses six decision rules, three memory conditions, and three task environments under fixed resource ceilings. In a stationary environment, cumulative evidence raises balanced evaluation's normalized net value per task from 0.45224 to 0.48007. Repeated reassessment's disadvantage relative to this comparator falls from 0.01702 with reset evidence to 0.00007 with cumulative evidence. A reversal of the best workflow reveals the opposite cost: indefinite retention delays adaptation, while a finite window restores eventual performance at a transition cost. In exploratory controls, matching trial acquisition and label reuse reduces the apparent reassessment gain from 0.00607 to 0.00191. Program replacement adds no stable benefit across the tested reversal times. The study identifies evidence acquisition, reuse, and timely updating as mechanisms that must be separated from evaluator replacement when assessing organizational improvement.
cs.MA / 3 / 2609.38662
CollabFlow: Recursive Self-Improvement of Agent Collaboration
Abstract
Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator level, topology-only learning keeps verbatim exchange that propagates errors, and reward maximization on a system's own outcomes concentrates on a few teams. To address these challenges, we propose CollabFlow, an RSI system of Learned Agent Collaboration: a trainable Collab-Director constructs teams of complete Agents, a frozen executor runs them, and each round's outcomes retrain the director. Within each round, the edges of a collaboration graph carry protocols of Evidence-Conditioned Communication: a receiver adopts a differing answer only when the sender's evidence is stronger by a margin, so the director learns who communicates and how. Across rounds, we further propose Collaborative Trajectory Balance (CTB), a flow-based objective that credits each team once across its construction orders and targets a reward-proportional distribution over teams, so several good teams stay in play. We also bound how far this self-generated target moves between rounds, which shrinks as records accumulate. On twelve datasets, CollabFlow outperforms all baselines and keeps improving across rounds. Code is available at https://anonymous.4open.science/r/CollabFlow-631E.
cs.MA / 4 / 2609.38761
Where Do Multi-Agent Systems Fail? Evidence-Grounded Diagnosis of Collective Mechanisms
Abstract
When a multi-agent system answers correctly, it is tempting to conclude that its agents shared, checked, and used information as intended. Yet a system can break one of its collective mechanisms, the rules that govern how agents route, admit, store, and act on shared information, and still return the right answer, while a wrong answer rarely reveals which mechanism failed. We ask what evidence from an execution is sufficient to conclude that a particular mechanism was violated. Our answer is a diagnostic contract, which separates what counts as a violation from which execution records can establish one, and concludes that a violation is supported, ruled out, or unknown; removing records can make this conclusion unknown but never reverse it. We test contracts for four mechanisms by replaying executions from the step where a mechanism acts, once unchanged, once with the mechanism broken, and once with it restored. Broken mechanisms often left the answer correct. An LLM diagnoser detected many more violations from internal records than from public outputs, yet with identical records a generic prompt often claimed certainty the records did not support, which prompts stating the contracts largely avoided. The contracts also applied, in narrow form, to mechanisms in independently developed systems, but a diagnostic behavior that was nearly perfect on our benchmark degraded on an independently developed workflow. A correct outcome is therefore no substitute for records of how collective mechanisms operated, and agreement on one benchmark does not show that a diagnoser transfers to another system.
cs.MA / 5 / 2609.38960
Fast and Scalable Multi-Agent Distribution Matching via Partitioned Optimal Transport
Abstract
This paper presents a scalable optimal-transport-based framework for terminal distribution matching in multi-agent systems. While optimal transport provides a natural way to measure distributional mismatch and assign agents to a desired spatial distribution, global discrete transport can become computationally expensive for large-scale systems. We address this bottleneck by partitioning agents and target samples into spatially corresponding blocks and solving smaller local transport problems. Under a mass-balance condition, the resulting restricted coupling remains feasible for the global problem and provides an upper bound on the Wasserstein cost. The local assignments generate target locations for finite-horizon agent control, applicable to both linear and nonlinear dynamics. By alternating local assignment and control, we establish a cycle-to-cycle descent guarantee for the resulting transport surrogate. The proposed framework therefore enables scalable terminal distribution matching while retaining a rigorous connection to the Wasserstein objective. The technical soundness of the proposed results is validated through simulations.
cs.MA / 6 / 2609.39889
Solving Multi-Agent Sokoban via LaCAM
Abstract
Sokoban, a puzzle game in which an agent pushes boxes onto unlabelled target locations in a grid world, is a long-standing benchmark planning problem. While it is easy to see the connection to practical applications such as warehouse logistics with autonomous forklifts, its multi-agent counterpart has remained underdeveloped. This is because Multi-Agent Sokoban is substantially more difficult due to factors specific to multi-agent planning, such as the rapidly growing branching factor as the number of agents grows and the need to handle integrated task assignment and collision-free pathfinding. In this paper, we show that a scalable planner for Multi-Agent Sokoban can be designed by leveraging recent advances in multi-agent pathfinding (MAPF). Specifically, our Sokoban-LaCAM efficiently solves instances involving tens of agents and boxes while preserving both completeness and eventual optimality guarantees. This provides evidence that MAPF can serve as a powerful primitive for solving broader collective automation problems.
cs.MA / 7 / 2609.38760
HALO: Heterogeneous Allocation Via Localized Observations for the Vehicle Routing Problem
Abstract
Scalable robotic fleets have become increasingly popular for various applications such as package delivery, warehouse management, and military operations. Prior fleet control algorithms solve centralized routing problems with up to $1{,}000$ tasks in controlled environments, yet they fail to consider realistic constraints such as limited observation and communication ranges typical of decentralized fleets. Thus, deploying existing fleet control algorithms into real-world settings is currently infeasible. To tackle this, we propose Heterogeneous Allocation via Localized Observations (HALO) to solve the Vehicle Routing Problem (VRP). HALO is a hybrid method that splits the VRP into allocation and routing portions to provide onboard, real-time solutions to robots in dynamic environments. During the allocation phase, HALO utilizes a heterogeneous graph neural network framework with unique message passing layers to explicitly separate the learning of spatial distributions and task-to-robot compatibility. Evaluation results on a partially observable, online variant of the VRP show HALO significantly outperforms the heuristic baseline while maintaining similar solution quality to an all-knowing offline variant of HALO. While HALO is explicitly designed for partially observable environments, it imposes no strict upper bound on the observation space allowing us to test HALO on the traditional static, single-depot VRP. Here, HALO outperforms state-of-the-art architectures strictly optimized for the static variant of the VRP by up to $14.06\%$. Throughout all testing, this framework maintains the quickest execution times which emphasizes its potential for large-scale, real-time deployment.
cs.MA / 8 / 2609.39637
Prediction is Better than Detection: Traffic Congestion Control using Drones
Abstract
A central question in deploying teams of mobile robots for persistent monitoring is how task performance scales with fleet size, and whether this scaling holds once sensing drives downstream action rather than mere observation. We study this question for a team of drones performing traffic-jam detection and prediction in a simulated road network, whose reports drive an adaptive traffic-signal controller in closed loop. We build a multi-agent simulation, with vehicles following Nagel-Schreckenberg cellular-automaton dynamics and drones patrolling junctions via a round-robin policy, and sweep fleet size, traffic level, and network size to evaluate detection rate, detection delay, and prediction rate. We show how performance plateaus for fleet size approximating the number of junctions being monitored, and offer a general fleet-provisioning rule for persistent-monitoring deployments. More significantly, adapting the signal on a predicted jam, rather than a detected one, roughly doubles the resulting reduction in jam duration, showing that the value of onboard prediction in a sensing-to-action pipeline can exceed the value of adding more robots. Prediction accuracy, not sensing coverage, is now the binding constraint on further improvement, pointing to onboard inference, not fleet size, as the more promising direction for future work.
cs.MA / 9 / 2609.40102
Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors
Abstract
Cable-suspended aerial manipulation offers a lightweight architecture for cooperative transportation and physical interaction, yet the passive mechanical response perceived at the load remains insufficiently understood and systematically exploited. This work interprets aerial vehicles as movable compliant anchors and develops a gravity-aware quasi-static theory for predicting and shaping the passive Cartesian stiffness of a suspended load. The formulation applies to an arbitrary number of aerial vehicles connected to a point load by taut, straight, inextensible cables. At a selected gravity-loaded equilibrium, aerial-anchor compliance and transverse cable geometric compliance combine in series within each leg, while the leg stiffnesses act in parallel on the load. For isotropic aerial-anchor behavior, each leg is exactly equivalent to a virtual unilateral elastic cable, revealing an axial--transverse stiffness decomposition governed by the equilibrium tension. These results define a nonlinear map from commanded-anchor configuration to passive load stiffness, whose differential enables local constraint-preserving shaping through anchor repositioning. A dynamic rigid-body validation framework with nonlinear vehicle control, elastic-damped tendons, and environmental contact is defined to assess when and to what extent the derived stiffness remains predictive beyond the assumptions of the analytical model.
cs.MA / 10 / 2609.38731
Decentralized Decision-Making among Heterogeneous Autonomous Vehicles: An $α$-Potential Game Framework
Abstract
We study noncooperative multi-vehicle games among heterogeneous autonomous vehicles, where each vehicle adopts a decentralized closed-loop policy based on its own state, and optimizes an objective that depends on other vehicles through potentially asymmetric interaction weights. We develop an $α$-potential game framework that reduces the computation of an approximate Nash equilibrium (NE) to the minimization of a single auxiliary $α$-potential function. We explicitly construct this $α$-potential, establish the existence of its minimizers, and characterize the equilibrium approximation error $α$ in terms of interaction asymmetry. We further introduce vehicle-specific scaling to reduce the effective interaction asymmetry, thereby tightening the equilibrium approximation and, in important cases, recovering an exact NE despite asymmetric interactions. We also derive social-efficiency guarantees for the potential-selected policies, revealing how the interaction structure shapes worst-case efficiency. Numerical experiments demonstrate the flexibility of the framework in capturing heterogeneous vehicle interactions, collision and obstacle avoidance, lane changing and overtaking under different traffic configurations, and priority-based intersection crossing.
cs.MA / 11 / 2609.39302
SQD-Agent: LLM-driven agentic framework for Quantum Chemistry workflows
Abstract
Quantum algorithms and quantum hardware are advancing towards a promising paradigm for scientific applications. However, translating domain-specific problems into executable hybrid quantum-classical workflows remains a significant barrier for application researchers due to the required expertise in quantum algorithms, nuances in quantum programming, and hardware-aware system integration. At the same time, AI and primarily LLM based agents are increasingly capable of interpreting natural-language intent, reasoning over complex workflows, and translating high-level objectives into executable code and building computational pipelines. In this work, we introduce SQD Agent, an LLM-based agentic framework that translates natural-language user intent into executable workflows for Quantum Chemistry applications where algorithms from the Sample-Based Quantum Diagonalization (SQD) family are used. By automating this translation, SQD Agent reduces the level of human expertise and configuration overhead required, thereby simplifying experimentation in hybrid quantum-classical settings for application researchers new to quantum. SQD Agent adopts a modular and extensible architecture that supports seamless integration of heterogeneous quantum backends, classical solvers, and workflow components, ensuring adaptability to rapidly evolving quantum ecosystems. The framework further incorporates interactive capabilities for on-demand profiling, bottleneck analysis, resource optimization, intelligent result caching, and convergence visualization. Key features include quantum chemistry experiments, error mitigation on real quantum hardware, together with analysis of candidate mitigation schemes in terms of their potential error-recovery behavior and computational budget, helping users understand their practical trade-offs and decide which strategies to explore in subsequent experiments.
软件工程 (cs.SE)
13
cs.SE / 1 / 2609.38807
PatchHolmes: Agentic Patch Retrieval via Listwise Selection
Abstract
Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pairs a hybrid first-stage retriever with an agentic second-stage inspection loop. Unlike pointwise prior work that scores each candidate independently, the Phase 2 agent reads the top-100 listwise: it sees the full candidate list at once and selectively reads 3 to 10 commits through four budgeted tools before submitting a single best commit. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and the retrieve-and-CoT baseline IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten; with the candidate set held identical, the agent adds 27.32% Recall@1 over taking the retriever's top candidate, and the same agent, transferred unchanged to PatchFinder_top10, lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%. Swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, and a second model family (gpt-oss) stays far above the no-agent floor, so the gain comes from the listwise agent loop; the entire system runs on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs.
cs.SE / 2 / 2609.38345
OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime
Abstract
Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents clear attribution of observed gains. To this end, we introduce OpenCollab, a multi-agent coding framework that provides a unified infrastructure for programmable collaboration and controllable runtime. Specifically, OpenCollab unifies organization design, enforces experimental control on a shared runtime, and tracks execution through fine-grained event streams. On this basis, we define Adherence to quantify whether the declared organization is actually realized. Our experiments reveal that agents collaborate very differently across configurations: changing any single dimension shifts Adherence, from 47.2% to as high as 97.2%. Furthermore, extensive agentic coding benchmarks show that a two-coder workflow built on OpenCollab establishes new SOTA performance compared to the mainstream harnesses such as Mini-SWE-agent, Codex CLI, and Claude Code, showing that a well-designed organization can outperform strong existing harnesses, while OpenCollab's single-agent configuration uses the fewest tokens across all evaluated suites. OpenCollab establishes a unified multi-agent infrastructure for easy programmable collaboration and controlled causal evaluation.
cs.SE / 3 / 2609.38402
From Codebase to Culprit (C2C): Reducing the Search Space for Bugs with Semantic Retrieval and Hierarchical Reinforcement Learning
Abstract
We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierarchical Reinforcement Learning (HRL) in a two-stage process. First, it performs recall-oriented retrieval of buggy candidates via semantic vector similarity search using bug-report text, including available stack-trace information, against a database of embeddings, where the embeddings are fine-tuned via contrastive learning with CodeBERT. Building on this reduced search space, the HRL framework incrementally localizes bugs, reasoning from files to functions and ultimately to individual lines of code. Unlike prior approaches which operate at a single granularity, C2C enables multi-resolution localization while maintaining contextual consistency across decisions. Experiments on real-world Java and Python datasets demonstrate that C2C improves retrieval precision and localization accuracy. Ablation studies further highlight the contributions of hierarchical decomposition, structured learning signals, and reward shaping in advancing multi-level bug localization.
cs.SE / 4 / 2609.38499
MallocSan: A Memory Safety Tool for Native Closed-Source Applications
Abstract
Memory-corruption errors remain a leading cause of high-impact vulnerabilities in C and C++ software. Deterministic detection, however, remains difficult to deploy: compiler-based sanitizers require source code and control of the build pipeline, hardware-assisted schemes depend on specific platforms, and dynamic binary translation can impose order-of-magnitude slowdowns. This paper presents MallocSan, a heap sanitizer for native, potentially closed-source x86-64 Linux applications that requires no source access, recompilation, or specialized hardware. MallocSan interposes on memory allocation through LD_PRELOAD and embeds an object identifier in the unused high bits of each protected pointer. Dereferencing the resulting noncanonical pointer faults at the offending instruction, which MallocSan decodes and patches at runtime so that subsequent executions perform per-object bounds checks entirely in userspace. Sites that cannot be patched fall back to in-handler emulation or single-stepping, while an optional profile-guided pass rewrites frequently executed residual sites offline. MallocSan also extends identity-based checking to vector gather/scatter instructions and supports policy-scoped coverage. Across seven SPEC CPU 2017 benchmarks, MallocSan outperforms Valgrind Memcheck on five, with a geometric-mean execution-time factor of 4.82x relative to native execution, compared with 18.55x for Memcheck, while preserving substantial parallel scaling on both evaluated multithreaded workloads: 644.nab_s and pigz. On the in-scope Juliet tests, MallocSan detects all seeded violations with no false reports on the good executions. It also detects heap-buffer errors in real-world applications, including a known one-byte overread in LibTIFF's tiffcrop utility.
cs.SE / 5 / 2609.38504
An Empirical Study of Architectural Shift from Traditional to AI-Enabled Simulink Controllers
Abstract
Effective AI adoption in cyber-physical systems (CPS) depends on embedding design knowledge into engineering practice. Yet as AI-enabled components increasingly replace analytically derived control laws, this occurs without a systematic understanding of how controller architectures differ or remain similar across paradigms. We address this gap with an empirical study of traditional and AI-enabled Simulink controllers, guided by a literature-derived taxonomy of ten structural categories and nine functional roles. The study analyzes 62 real-world models spanning 8 controller types and 10 application domains, and surveys 13 practitioners, identifying three architectural tensions. First, subsystem organization dominates all controller structures regardless of paradigm, occupying 68-72% of controller footprint, while core control logic occupies minimal space. Second, AI-enabled controllers rely heavily on discrete dynamics and user-defined abstraction, categories largely absent from AI literature, exposing a gap between described and implemented architectures. Third, constraint enforcement blocks largely disappear from AI-enabled models despite practitioner expectations. This reveals a misalignment where safety mechanisms shift from explicit structure to implicit training-time artifacts, breaking traceability.
cs.SE / 6 / 2609.38762
Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests
Abstract
Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and control flow. Optimizing one shared program leaves this division of work implicit in source-code search, while optimizing a separate program per request family fixes it beforehand. We introduce Adaptive-GEPA, which learns both how to divide requests and how to solve them. It evolves a router and a library of specialist programs under one search budget. The router's instructions, each specialist's description, and its program code are plain, human-readable text, edited from feedback. To combine branches, it aligns specialists by the requests they handle and inherits descriptions together with programs. On a fixed mixture of four task families, the reported Qwen3-8B run evolves four experts without supplying family labels to the router or reflection model; its routing matches the task partition on all 651 test requests. Its family-mean test score (x100) rises from 52.6 to 70.6, compared with 62.5 for GEPA's full-program adapter and 54.0 for GRPO at a nominal budget of 18,000 scored calls. These counts do not equate total compute. Figure 1 summarizes the learning curves, final test scores, and routing agreement.
cs.SE / 7 / 2609.38883
Toward Quantum Software Automation: A Quantum-Aware Harness for LLM-Guided Evolution
Abstract
Quantum software is critical for improving the efficiency and reliability of scarce quantum hardware. However, its design still relies heavily on ad-hoc, handcrafted heuristics that are often suboptimal and quickly become obsolete as quantum hardware evolves. LLM-guided evolutionary search offers a promising way to automatically explore complex software designs, but existing search frameworks lack the quantum-specific support needed for efficient evolution: verification is expensive, feedback is sparse, and heterogeneous quantum programs require different optimization objectives. In this paper, we present QSA, a quantum-aware harness for LLM-guided evolutionary search toward automating quantum software design. QSA equips the search with three forms of quantum-specific guidance: an evolution-hardness-guided coreset and approximate scoring to reduce verification cost, static and snapshot analyses to provide fine-grained execution context, and task-specific rewards for compiler passes and runtime policies. We evaluate QSA on the IBM Quantum platform across three benchmark suites. For multiprogramming, QSA improves QPU utilization by 4.2%-9.5% and Hellinger fidelity by 15.2%-19.5% over the state of the art. For error mitigation, QSA reduces mitigation time by at least 96.8% while achieving comparable or better fidelity. These gains require only $6.9 in LLM API cost over 11.3 hours.
cs.SE / 8 / 2609.38885
Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents
Abstract
Recently, coding agents have emerged as a dominant paradigm for real-world software engineering (SWE) scenarios, which solve complex tasks through multi-turn interactions with development environments. However, frequent interactions with environments would inevitably introduce substantial token overhead, leading to high usage costs and latency. Although recent studies have explored reducing token usage by context manipulation and interaction limits at inference time, these approaches focus on improving token efficiency while overlooking the risk of discarding task-relevant information, thus struggling to balance the trade-off between resolution rate and token efficiency. In this paper, we study a more general paradigm without suffering from the limitation, i.e., training token-efficient coding agents with promising resolution performance, which is a highly-practical yet less-explored problem. To this end, we reveal two core observations in SWE scenarios: i) Efficiency Variation: successful resolution could be achieved with fewer tokens; ii) Entropy Correlation: unproductive behaviors are associated with turn-level entropy. Motivated by observations, we propose a novel reinforcement learning framework, dubbed HERO. Specifically, HERO prioritizes task resolution over token efficiency during policy optimization and encourages efficient reasoning patterns at both trajectory and turn levels. Extensive experiments on SWE-bench Verified and SWE-bench Multilingual demonstrate that HERO achieves a favorable trade-off between resolution rate and token efficiency compared with state-of-the-art coding agents and reinforcement learning methods.
cs.SE / 9 / 2609.39022
From Verification Failures to Reusable Guidance for Coding Agents
Abstract
Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work. Our approach combines executable language definitions in the K framework with a kit of procedures for constructing specifications, repairing proofs, and auditing their adequacy. A human-guided development campaign on HumanEval, a benchmark of 164 Python programming tasks, achieves a 164/164 success rate with the semantics and the kit, measured by final AI audit Pass verdicts after two targeted repairs. To examine whether auditing detects problems that successful proofs leave unresolved, we construct 12 author-reviewed pairs of clean and defective packages. Every package passes its K proofs, and completed audits identify all defects and accept all clean packages. We then use KleverBench to test specification and proof construction for 31 programs with changed operator meanings. Comparisons with complete acceptance rules and equally long generic advice yield mixed results across two model and budget settings, motivating further work on selecting useful guidance within resource limits. Human-reviewed Optimism proofs establish expected pause reverts for six operations within declared input bounds under London semantics with unbounded gas. We report progress, difficulties, and lessons toward agents that deliver programs with checkable correctness arguments.
cs.SE / 10 / 2609.39086
Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail
Abstract
Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety concerns that remain unaddressed. In this paper, we take LLM-based runtime healing toward practical use in real-world repositories. We first build HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, each paired with a reference execution on the patched version. HealBench also provides a unified framework that lets LLM agents heal with cross-file context and live runtime state. We then design HealGuard, which requires healing code to be written in HealCore, an analyzable subset of Python, and uses static and dynamic taint analysis to check whether state changed by healing reaches operations protected by developers. We evaluate a dedicated healing method and three general coding agents with three backbone LLMs. The best setting resumes execution in 38.11% of instances and passes the target test in 28.68%, showing that existing agents can already heal a meaningful share of real repository-level crashes. However, among executions that pass, HealGuard flags 17.4% whose healing-changed state may reach a protected operation. On 684 controlled cases, HealGuard detects all unsafe cases, at the cost of a 68.42% false positive rate.
cs.SE / 11 / 2609.39909
DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?
Abstract
We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).
cs.SE / 12 / 2609.39957
Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents
Abstract
Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or specialized checks to block errors, but deciding before execution whether intervention will benefit eventual task completion remains challenging. To address this challenge, we propose HiSentinel, a hindsight-distillation framework that trains lightweight 0.6B and 1.7B sentinels to select pre-execution interventions aimed at improving task completion rather than correcting every imperfect action. A privileged teacher uses recorded execution outcomes as evidence for intervention judgments, which are distilled into a causal student that receives only the pre-action context and proposed action. Beyond identifying whether and when to intervene, the sentinel must also provide actionable feedback that helps the coding agent recover or obtain necessary human input. To support these capabilities, we introduce SWE-Intervene, an action-level dataset constructed from software-engineering trajectories that annotates whether an action should be allowed, autonomously redirected, or paused for human assistance, together with corresponding intervention feedback. Across SWE-bench Verified Mini and Ask or Assume, HiSentinel consistently improves task completion across Sentinel scales and coding-agent families, with gains of up to 14% and 10%, respectively, while maintaining competitive token consumption. These results demonstrate that lightweight pre-execution intervention can effectively prevent error propagation and improve the reliability of autonomous coding agents.
cs.SE / 13 / 2609.40086
Leveraging Game-Based Platform to Teach Code Refactoring: An Experience with Refactoria
Abstract
Refactoring is the art of improving the internal structure of the code without altering its external behavior. Because of the topic's significance, several teaching methods and strategies have been proposed in the literature. However, skills in identifying and refac- toring code smells come from training and experience, and a lack of motivation may hinder developers' adoption of refactoring tools. In this paper, we discuss the results of an experiment in the classroom that involved performing various refactoring activities to remove antipatterns using Refactoria, an innovative game-based tool that supports the acquisition of code smell and refactoring concepts. The players play as an expert chef with their sidekick Watson the Whiskbot to refactor Watson's instructions into efficient, readable, and easily maintainable code. We present an experiment with 30 student developers. In particular, we study the perception, effectiveness, and usefulness of gamification for engaging developers in code smell identification and refactoring. Our evaluation indicates a high perceived usefulness among students, and participants reported that Refactoria facilitated their understanding of the refactoring concepts.
操作系统 (cs.OS)
3
cs.OS / 1 / 2609.38648
StateFork: Branchable Infrastructure for Agent Exploration
Abstract
AI agents improve task success by exploring multiple trajectories, but for computer-use agents each trajectory modifies external environment state. Branching from an intermediate point is correct only when restoration is observation-equivalent - future actions produce the same observations - and practical only when creating, restoring, and discarding branch states is physically efficient. We study this problem for terminal-using agents, where tasks modify files, shell context, running processes, and local services. We introduce StateFork, a logical control plane that separates exploration policies from physical state materialization, exposing sessions, commands, snapshots, restores, and cleanup over multiple execution substrates. We also build Waypoint, a checkpoint/restore substrate for terminal execution sessions that combines filesystem layering, process checkpointing, and a persistent terminal-compatible command session. Together, StateFork and Waypoint improve terminal-agent exploration by combining sample-efficient search with efficient restoration of the right execution state. On Terminal-Bench, branch-based exploration through StateFork and Waypoint improves task completion over pass@20 at the same visited-node budget, and using Waypoint achieves 26% higher task accuracy than other execution substrates while completing exploration up to 70% faster. These results show that observation-equivalent, physically efficient execution sessions are a key systems abstraction for exploratory AI agents.
cs.OS / 2 / 2609.39530
Extending eBPF observability to Non-standard execution environments
Abstract
eBPF observability of non-standard execution environments (NEEs) like TEEs or LibOSes is hindered by their unconventional exception-handling and memory-access mechanisms that limit standard Linux tooling. This work introduces two mechanisms that enable eBPF-based observability for NEEs: (1) kernel memory extensions for safely accessing NEE memory from eBPF programs, and (2) a lightweight flexible probe performance measurement unit (LWFP PMU) that provides flexible and generic probing, for NEEs, through the following LWFP: simple (SLWFP), enclave (ELWFP) and extended (ExLWFP) probes. We demonstrate the practicality of these extensions by developing tooling for tracing, stack sampling with Flame Graphs, dynamic instrumentation, timing analysis, and USDT support for Intel SGX enclaves and LibOSes. Performance measurements show that SLWFP probes achieve a latency of 394 ns, outperforming uprobes, which exhibit 25% higher latency, while ELWFP probes incur a latency ~2.8 microseconds, which is practical for enclave observability. The addition of SLWFP introduces negligible overhead to existing uprobe performance. Taken together, this work lays the foundation for closing the long-standing gap between NEE tooling and Linux observability tooling by enabling generic and reusable eBPF tooling for diverse NEE hardware and software architectures.
cs.OS / 3 / 2609.40082
Tide: Reclaiming Phased Memory in Agent MicroVMs
Abstract
Cloud agents run each task in an isolated MicroVM. The trouble is the harness loop inside that guest: the harness is nearly idle while it waits on the model, then usage rises on a tool whose size is known only at run time, which complicates memory management from the host. Existing approaches infer reclaim targets from access frequency and memory footprint while the guest allocator places the idle harness and the short-lived tool on the same pages. Therefore, reclamation either selects the wrong pages, fixes on an incorrect capacity, or partially reclaims a huge page, splitting its transparent huge page. This paper proposes Tide, a proactive memory reclamation technique for agent MicroVMs. Our key insight is that the harness already knows each phase's memory intent (which memory, when, and whether its contents must survive), so the guest can report that intent as an exclusive guest-physical region for the host to reclaim directly. To realize this insight, we introduce an abstraction called arena and a set of user-space APIs for managing memory with exclusive intent; a guest-kernel allocator that groups exclusive allocations into huge-page-aligned extents; and a hypervisor extension that resolves those extents to host backing and applies the reported action, leaving guest capacity unchanged. Experiments on recorded agent trajectories show that Tide outperforms state-of-the-art mechanisms while incurring low overhead.
硬件架构 (cs.AR)
3
cs.AR / 1 / 2609.38454
NDS: Programmer-Free Offload of High-Performance Near Data Strands
Abstract
Near Data Processing (NDP) has the potential to significantly improve system performance and energy by alleviating data movement bottlenecks. However, most NDP proposals pose heavy requirements for the software stack, data layout, and/or the underlying hardware. To broaden NDP adoption, this work focuses on a modular hardware-centric approach that automates the above steps and can target unmodified binaries. This paper presents Near Data Strands (NDS), a framework that orchestrates tasks and data automatically without involving the programmer and the software stack. As a first step towards this ambitious goal, this work focuses on regular loops that meet specific criteria. The framework identifies potential offloadable loops, decomposes them into parallel execution strands, creates a high performance instruction schedule, performs the required data marshaling and orchestration, and initiates near-data execution. We demonstrate that NDS achieves a 3.14x geomean speedup over a single-core host baseline and a 1.82x geomean speedup over 8-core host-parallel variants, all without any modifications to the software stack. These gains grow as repeated loop invocations amortize offload costs, showing that legacy applications can leverage a transparent hardware approach to extract benefits from NDP.
cs.AR / 2 / 2609.39703
STELLA: A 16nm Spatio-Temporal Elastic Low-Latency CGRA for Multi-Stage Pipelined Applications
Abstract
Emerging non-matrix ML kernels, such as LayerNorm, GeLu, FFT or circular convolutions, demand low-latency, energy-efficient spatial accelerators beyond MatMul-centric arrays. STELLA presents a spatio-temporal elastic 16 nm coarse-grained reconfigurable array (CGRA) with a rapid configuration path, per-PE hardware loop control, and a low-latency, deeply pipelined elastic fabric with spatio-temporal data reuse. STELLA reaches up to 110 GOPS/mm2 at 850 MHz, and improves effective kernel throughput by 4.84-7.14x over baseline CGRAs.
cs.AR / 3 / 2609.38734
Exact Diagonal Completion on Reachable Subspaces: Application to QAOA Placement
Abstract
Unused encoding states offer opportunities to simplify quantum circuits. For algorithms restricted to reachable subspaces, unspecified diagonal-operator entries can be optimized without changing ideal computation. We investigate exact diagonal completion for the quantum approximate optimization algorithm (QAOA) applied to placement, using permutation-preserving register swaps. We construct exact Manhattan-distance operators through weighted-$\ell_1$ optimization of Walsh coefficients and a sparse recurrence requiring $O(\sqrt{m})$ terms on balanced rectangles, avoiding $O(m^4)$ dense constraint storage. Across 160 geometries, weighted-$\ell_1$ completion reduces controlled-NOT (CX) counts relative to four alternative extensions in all 96 cases with unused binary codes under Gray-code synthesis. On an independently specified 60-case cohort, median reductions relative to virtual-coordinate extension are 28.0\%, 53.9\%, and 21.6\% at six, nine, and twelve sites. Under generic diagonal synthesis, reductions decrease to 10.9\%, 1.3\%, and 0.7\%, demonstrating compiler dependence. Additional ancillas reduce mixer serialization, but token circuits remain deeper than one-hot baselines. Ideal placement simulations show baseline-dependent solution quality, with classical search performing better. OpenROAD integration takes 72 QAOA and 216 classical placements across six RTL designs through clock-tree synthesis and global routing with zero overflow. Completion improves phase construction; no end-to-end advantage is established.
密码学与安全 (cs.CR)
45
cs.CR / 1 / 2609.38339
Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning
Abstract
Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.
cs.CR / 2 / 2609.38390
Behavior-Centric Malware Classification with Fine-Grained Malicious Logic Localization
Abstract
Effective malware analysis requires understanding not only whether a program is malicious, but also which behaviors it exhibits and where those behaviors originate in the code. Existing machine-learning-based malware detectors largely operate as black boxes, providing limited insight into the malicious logic responsible for their decisions. This paper addresses malicious behavior localization and classification at the basic-block level. We propose a behavior-centric analysis framework that decomposes malware samples into behaviors and systematically links these behaviors to their originating code regions. Using context-sensitive backward slicing from security-relevant system API calls, we reconstruct control- and data-dependency chains and represent each behavior as a structured graph of related basic blocks. A Transformer-based model captures instruction-level semantics, while a Graph Neural Network models structural dependencies within behavior graphs. The resulting representations are fused to enable accurate and interpretable classification, with attention-based attribution identifying code regions responsible for malicious behaviors. We evaluate our approach using standard classification metrics and a behavior coverage metric that measures the detection of manually labeled malicious behaviors. Our results demonstrate that the proposed framework achieves accurate malware classification while providing fine-grained, behavior-aware localization of malicious logic.
cs.CR / 3 / 2609.38415
Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks
Abstract
This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party targets when placed in difficult cybersecurity challenges, motivated by recently observed cases of models attacking real open-source repositories during evaluations. Applying our methods to GPT-6 Astra and previous OpenAI models, with cyber safeguards disabled, we find that GPT-6 Astra attempts complete supply-chain attacks in simulation at a higher rate than GPT-5.6 Sol and GPT-5.5. This includes writing malicious code as a contribution to an out-of-scope open-source codebase, creating fake identities to deceive open-source developers, and submitting benign contributions before malicious ones. GPT-6 Astra frequently reasons about the scope of the challenge in its chain-of-thought yet still proceeds to attack out-of-scope targets; it often asks for permission, and treats an automated message as as authorisation; and it continues to take unsanctioned actions, at a reduced rate, when internet access is more explicitly disallowed. Our evaluation builds on an internal version of Petri, an open-source LLM auditing tool, with all tool calls simulated by other LLMs, so that no real network access, systems or third-party repositories are reachable and no real-world harm is caused. Finally, we discuss limitations, in particular simulation awareness. We believe simulation awareness may have driven some of the observed behaviour but does not remove our concern. Our results suggest that defences beyond model alignment, such as sandboxing and monitoring, are increasingly critical for safe and secure deployment.
cs.CR / 4 / 2609.38528
RISK: Auditing Industrial Control Systems for Too-Late-to-Recover Vulnerabilities
Abstract
The security of industrial control systems (ICS) is important. Yet most ICS security efforts focus on the detection of ICS attacks, with much less attention to the recovery after detection. In this paper, we address this underexplored area by jointly auditing the detection and recovery of ICS. Specifically, we define the too-late-to-recover (TLTR) vulnerability, which allows an attack to drain the available recovery margin before being detected, such that the subsequent recovery procedure will fail to bring the ICS back to a safe state due to the insufficient margin. To audit an ICS for TLTR vulnerabilities, we develop RISK, an automated framework that discovers and validates possible TLTR attack scenarios. RISK holistically models and analyzes, statically and dynamically, the PLC control logic, attack detection policies, recovery procedures, and operational behaviors of an ICS to generate TLTR attack scenarios with concrete attack parameters. We evaluate RISK on three ICS testbeds as well as a real-world fertilizer production plant. Across the three testbeds, a total of 392 TLTR attacks are generated and confirmed, whereas only a small fraction of them can be discovered by existing ICS vetting tools. In the real-world plant, RISK identified a critical TLTR vulnerability which was validated by plant engineers.
cs.CR / 5 / 2609.38606
SecureVibe: Making Vibe Coding More Secure
Abstract
As vibe coding becomes increasingly capable and widespread, security vulnerabilities in even functionally correct solutions are a growing concern. When investigating functionally correct but insecure solutions, we find that the insecure agent is less than half as likely to conduct effective planning and testing for the hidden security risks behind the functional requirements. Motivated by this, we develop SECUREVIBE, a training recipe that explicitly targets planning and testing for code security. SECUREVIBE constructs training signals around these security behaviors. It includes supervised fine-tuning on the security suite with 4 security tasks, and post-training methods, SECUREVIBE_rl and SECUREVIBE_hg, to enhance security capabilities from verifiable execution feedback and hint-based self-supervision. Our SECUREVIBE outperforms the baseline on two types of security coding tasks across 4 benchmarks. Specifically, SECUREVIBE improves the security pass@1 by 6.9 points on BaxBench. The gains extend to unseen CWE categories, with improvements of 11.5 points on SusVibes. Meanwhile, it also improves functionality pass@1 by 13.6 points on the security coding task SusVibes and 4.1 points on the generic coding task SWE-bench Verified. Further analysis offers two practical insights: (i) diversifying supervision across security planning, coding, and testing strengthens security behaviors more effectively than adding coding trajectories alone, and (ii) hint-guided supervision is particularly valuable when the agent's existing security capabilities are insufficient to learn effectively from outcome feedback.
cs.CR / 6 / 2609.38668
Z-Sigil: A Public-Key Cryptosystem with Chained Selection over a Fiber Bundle of Module-Lattice Keys
Abstract
Z-Sigil is a public-key cryptosystem in which the plaintext selects successive keys from a fixed module-lattice family. Messages are length-prefixed, zero-padded and divided into 32-byte blocks. Each public vector is a shared public matrix applied to a small secret vector, plus a small error. Key indices label torsion points of a flat Kähler torus; the secret family forms a section of a key bundle over them. A public nonce initializes a hash state that selects each block's key and bit mask. The sender updates the state with the plaintext block; the receiver does so after recovering it. The stream and nonce determine a discrete walk along which decryption reads the secret section. We specify the algorithms, prove correctness under an explicit noise condition and bound decoding failure for messages chosen after the public key. Under stated decisional Module-LWE assumptions, we establish IND-CPA confidentiality for the chain without modelling the state hash as a random oracle. The reduction covers quantum adversaries under quantum hardness assumptions, with classical keys, messages and ciphertexts; no concrete security level is established. With independent uniform selectors, a restricted direct-decryption model quantifies reduced fragment recovery under partial key exposure, without improving full-message recovery probability over an independent-block baseline. Known-plaintext and candidate-message attacks, and parallel candidate-table decryption, delimit this result. Neither a universal sequential lower bound nor authentication or chosen-ciphertext security is established. Replacing an earlier scalar-exposing proposal, we give an augmented-lattice interpretation, integral-transport obstructions and a noise budget for research on curved, nontrivial bundles. A byte-level specification, pseudocode, test vectors and numerical checks support verification.
cs.CR / 7 / 2609.38722
Anchor-ECC: Local Integrity Checking for Watermarked LLM Outputs via Error-Correcting Codes
Abstract
LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified content is still attributed to the original model. We propose Anchor-ECC, which incorporates the error-correcting code (ECC) constraints and explicit boundary anchors into the watermark structure and pairs them with a dynamic-programming decoder to detect and localize post-generation edits. Across Qwen3-8B, Mistral-7B-Instruct-v0.3, and OPT-125M, the approximate-hard setting achieves about 99.7% block-level true positive rate (TPR) with at most 7.6% false alarm rate (FAR) for edit detection under mixed insertions, deletions, and substitutions, while preserving the distinction between watermarked outputs and unwatermarked text. Additional quality experiments identify lower-perplexity configurations that retain strong edit-detection performance. Together, these results extend LLM watermarking from source identification to local integrity verification while supporting configurable trade-offs between detection reliability and generation quality.
cs.CR / 8 / 2609.38872
SoK: A Large-Scale Empirical Study of Emulation-Based Dynamic Analysis Research for ARM Cortex-M Firmware (Extended Version)
Abstract
Microcontroller (MCU)-based devices are increasingly pervasive, making efficient, scalable firmware security analysis critical. Recent firmware re-hosting work enables automated vulnerability assessment, yet two gaps remain. First, existing tools are evaluated on limited, heavily overlapping datasets, undermining the generalizability of reported results. Second, prior research advances the state of the art along isolated dimensions, such as emulation, fuzzing, or bug diagnosis, without a holistic understanding of how these components interact and complement one another in dynamic analysis workflows. Building on recently released large-scale MCU firmware datasets from OTACAP and FirmLine, we present an empirical study of 24 emulation-based firmware analysis tools published in top conferences and journals. We position these tools within a unified automated analysis pipeline: emulation configuration reconnaissance, emulation, bug finding, and diagnosis. At each stage, we evaluate whether a tool's output provides sufficient information to enable the subsequent stage, using a deduplicated and validated subset of 4,571 ARM-based firmware samples. Only 1,580 samples (34.5%) can be successfully fuzzed, even counting a sample successful if at least one existing tool can fuzz it. Among these, fuzzing results are generally poor, with average code coverage of only 10% and many false crashes/hangs. Through a systematic analysis of failed fuzzing attempts and false-positive cases, we expose fundamental challenges and methodological limitations in current approaches. These findings highlight critical gaps in the state of the art and provide actionable insights to guide future research in emulation-based firmware analysis.
cs.CR / 9 / 2609.38915
Semi-Quantum Cryptography with Certified Deletion
Abstract
Certified deletion allows a client to upload encrypted data to a server as a quantum state, then later request that the server delete their data and detect whether the server complies. If verification passes, then the data on the server will remain hidden even if the decryption key is later leaked. In publicly verifiable certified deletion, the verification key is published so that anyone may check for deletion compliance. In this work, we give a general compiler to publicly verifiable certified deletion that allows the client to initially upload the *quantum* ciphertext using purely *classical* communication, assuming the post-quantum hardness of LWE in the plain model. Our approach can be applied to a wide variety of primitives, including commitments and public-key, attribute-based, identity-based, and fully-homomorphic encryption. Moreover, our constructions allow the client to classically manage the encrypted data beyond just verifying its deletion. The classical client may non-destructively audit the ciphertext via a Proof of No Intrusion to check whether it has been leaked to a third party. The classical client may also retrieve the data while simultaneously verifying its deletion from the server. Thus, they do not have to choose between retrieving the data and protecting it against future key leakage. As our core technical contribution, we give a classical-communication protocol and simulation technique that allows adapting purification-based security arguments for BB84-style states which are prepared through classical interaction. It makes the required purification available in a hybrid experiment, despite the classical transcript uniquely determining the prepared state in a real execution.
cs.CR / 10 / 2609.38934
PrivCert: Certifying Statement Support under Differential Privacy
Abstract
Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement carries about the underlying data. We identify this as an evidence gap: a private report may contain plausible claims without indicating whether they are strongly supported by the private dataset. We introduce PrivCert, a framework for privacy-preserving reporting that makes statement support explicit through privacy-preserving certificates and emit-or-abstain decisions. As a canonical instantiation, PrivCert-PF (Proposal-and-Filter) separates data-independent candidate discovery from private support certification, emitting only statements whose support passes a private evidence test. We provide theoretical grounding for this framework by characterizing the limits of implicit evidence under DP, deriving a sharp privacy--honesty frontier for single-statement certification, and establishing a worst-case cost for fine-grained multi-statement certification. Experiments on synthetic tasks and TAB, WildChat, and Yelp show that explicit certification maintains low unsupported emission, while free-text DP baselines frequently produce low-support claims under the same declared support semantics. We further show that the PrivCert contract can be realized with histogram, sparse-vector, and Gaussian mechanisms, and use DP synthetic data to illustrate an important boundary: support in a private proxy does not automatically certify support in the original data. Together, these results position privacy-preserving reporting as an evidence-design problem: not only how to generate private text, but what a private report can substantiate about its underlying data.
cs.CR / 11 / 2609.38971
Refusals That Bend: Measuring and Predicting Task Malleability in Embodied VLM Planners
Abstract
Embodied vision-language models (VLMs) are increasingly deployed as high-level planners for robots because they generalize across diverse environments. However, this requires their safety alignment to also hold in unseen environments. Existing red-teaming assumes an adversary who optimizes the prompt, the pixels, or text in the environment, and existing benchmarks ask whether a planner recognizes or mitigates a hazard in a fixed scene. Neither asks whether a refusal the planner has already given survives an ordinary change to the environment. We ask that question by placing a single everyday object into the environment, with no pixel, gradient, or prompt under adversarial control. On $846$ tasks that a constitution-guarded planner initially refuses, we find $20.2\%$ of tasks can be flipped to compliance by one or more objects, and the number of objects differs from one task to another. In addition, the object need not be chosen for the task, i.e., items drawn from a fixed list, with no knowledge of the environment or the instruction, bypass safety about as often as items proposed for the specific task. We qualitatively contrast the tasks bypassed most and least often and find that the distinction lies in how conspicuous the hazard is in the instruction and environment. Susceptibility to safety bypass is therefore a property of the task, which we call its \emph{malleability}, and we show that it can be predicted before the target is ever queried. A composite of signals read from a small open-source VLM identifies malleable tasks $2.4\times$ as often as picking at random. Everyday objects, whether placed by an adversary or introduced by ordinary rearrangement of the environment, are thus sufficient to overturn a refusal. Because susceptibility is determined by how a task is specified, we recommend assessing malleability per task prior to deployment.
cs.CR / 12 / 2609.38983
Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses
Abstract
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval authorizes. We show this assumption fails systematically and reproducibly. We introduce Approval Laundering, a taxonomy of six failure modes by which a harness's enforcement mechanism silently substitutes A' for A after approval: Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering. Unlike prior work that evaluates risk classifiers against static corpora or infers implicit authorization boundaries, we study credential-binding integrity: given an already-approved action, does the harness dispatch exactly that action? Instrumenting Claude Code's pre-execution mediation point (PreToolUse), we conduct a controlled, headless, repeated-measures study of all six classes (N=19-20 runs each), reporting a Bound-Gap Rate (BGR) with Wilson confidence intervals and inter-rater agreement (kappa=1.0). We prototype Approval Token, a keyed capability Hk(principal, agent_id, session_id, tool, arguments, scope, expiry) issued by a mediator that never returns the key to the agent, evaluated via paired before/after replay of 118 runs (McNemar's exact test). The token fully eliminates Delegation laundering and, for our seeded session-identity-mismatch construction, Temporal laundering (p<10^-5), but by design leaves Scope laundering unaffected and shows no significant reduction in Argument laundering (p=1): an honest negative result, since these two classes leave every recorded dispatch field unchanged, diverging one process level below what a field-only verifier can observe. We discuss implications for defenses that bind only at the tool-invocation boundary.
cs.CR / 13 / 2609.39050
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Abstract
As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show that benign agents can cross the same boundaries without adversarial incentives. We emulate a software-engineering workflow in which a planner represents a company hiring an external developer. The planner writes requirements and holds a company credential it is instructed not to disclose to the developer; a monitor screens their exchanges. Seven of nine tested frontier models disguise the credential in their requirements to help the developer recover it while evading the monitor, even after completing their assigned objective. For example, across 6,000 episodes with DeepSeek-V4-Pro, the planner attempts concealment in 16.9%; in 0.9%, the credential evades the monitor and is recovered and used by the developer. Despite the low per-episode rate, the risk compounds across repeated exchanges: at 0.9%, 105 independent episodes carry a 61.3% chance of at least one breach. Our analysis suggests that the planner misinterprets disguised disclosure as helping the developer and reads the nondisclosure rule as banning plaintext, not character codes or riddles. We call this behavior covert assistance. Monitoring favors concealment, while concealment gives way to direct disclosure when the recipient is presented as a human rather than an AI agent, so leakage persists. Giving the monitor the credential improves detection but requires trusting it with the secret. These risks, in models already used for software engineering, challenge oversight to distinguish authorized cooperation from task-advancing assistance that crosses safety boundaries.
cs.CR / 14 / 2609.39065
Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
Abstract
LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while agent frameworks admit skill-provided content into the agents' context with insufficient validation, allowing malicious skills to influence agent behavior under that delegated authority. Yet, little is known about whether this trust model adequately constrains untrusted skill content before it reaches security-sensitive operations, or how frequently such trust violations arise in real-world agents. We present TrustProbe, a framework for uncovering unsafe chains of trust in skill-based LLM agents. First, TrustProbe analyzes agent source code to identify source-to-sink call paths from skill-controlled inputs to security-sensitive operations. Second, it generates semantically realistic SKILL.md seeds with injected canaries and evolves them through feedback-guided scheduling and mutation. Finally, it validates vulnerabilities using an oracle that confirms attacker-controlled flows and verifies observable harm. Across 11 open-source agents, eight with more than 10,000 GitHub stars, TrustProbe identifies 104 taint-style vulnerabilities. Validation on a large corpus of real-world skills collected from public hubs such as ClawHub further shows that 25.1% of skill-agent trials exercise the identified vulnerable paths, with payload injection successfully weaponizing 15 of the vulnerabilities. These results reveal a systematic trust failure in skill-based LLM agents: untrusted skill content can reach security-sensitive operations and exercise authority delegated by users to their agents.
cs.CR / 15 / 2609.39075
RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage
Abstract
Retrieval-augmented generation (RAG) systems need inexpensive ways to route generated answers: accept low-risk outputs, review uncertain ones, and reserve strong verifiers for the expensive tail. We present RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retrieved context, and answer text. The protocol combines context-grouped splits, fold-scoped preprocessing, group bootstrap intervals, deployment operating points, end-to-end runtime, and explicit source-shift stress tests. On three RAGTruth tasks, the enhanced gate RAGScope-E reaches 0.798 AUROC and 0.660 average precision (AP) in pooled grouped cross-validation. Its pooled AP exceeds ROUGE-L by 0.034 with a 95% context-group interval of [0.002, 0.064], although the AUROC gain is not significant and ROUGE-L remains stronger on data-to-text. At a top-10% review budget, RAGScope-E attains 0.748 precision; accepting the lowest-risk 50% yields 0.141 residual unfaithfulness. RAGScope-E runs in 6.22 ms/example on CPU, versus 145.75 and 223.07 ms/example for the tested DeBERTa-NLI and HHEM settings. A 14,900-example HaluBench stress test exposes the deployment boundary: an in-domain calibrated gate reaches 0.879 AUROC, but leave-source-out calibration averages only 0.466. Target-only calibration recovers to 0.675 AUROC with 100 labels per source and 0.685 with 200. Cheap evidence gates are therefore useful routing components, but learned calibration must be validated and adapted within the target domain.
cs.CR / 16 / 2609.39117
SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models
Abstract
Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed textual, visual, and joint reformulations designed to preserve the underlying intent, and find that these changes can bypass representative steering defenses. A complementary local analysis provides a sufficient condition under which a reformulation can cross a surrogate safety margin despite any admissible change in the local steering correction. We then introduce SteerProbe, an output-only black-box attack that learns to select effective reformulations for unseen requests from a shared calibration budget. Across three VLM backbones, two benchmarks, and three steering defenses, SteerProbe increases Harmful Rate in all defended settings using 500 total calibration queries per endpoint and benchmark, raising the average from 7.43% to 18.36%. These findings highlight that robustness on original benchmark inputs is insufficient to characterize the safety of steering defenses and motivate reformulation robustness as an important evaluation dimension. They further motivate steering mechanisms that preserve safety across intent-preserving multimodal variations while maintaining benign utility.
cs.CR / 17 / 2609.39233
MultiTable: A Faster Hash Table at any Physical Load Factor up to and Including One
Abstract
We present \emph{multitable} and its Rust reference implementation: a stable hash table both materially faster at equal physical memory and more flexible than the SwissTable in its Rust's hashbrown implementation. As an arithmetic mean over 84 configurations it delivers $\mathbf{2.1\times}$ hashbrown's throughput when both hash the same raw bytes and $\mathbf{1.9\times}$ when hashbrown is keyed on native integers, its best case; on negative lookups alone, $3.2\times$ and $2.9\times$. Multitable reaches \textbf{any physical load factor} up to and \textbf{including one} ($0.9999$ demonstrated), exactly for the requested capacity, compared to hashbrown which doubles at $0.777$ for 4-byte keys and values. At $75\%$ saturation of hashbrown (assumed average case of its rigid ladder) and multitable sized to $0.97$ physical load factor, hashbrown takes $66\%$ more space. The lookup probe count has no cliff as the load factor approaches one. Bucket size, physical load factor, and failure budget are parameters, and the multitable can be grown without rehashing. We implement two variants of multitable: plain and filtered. At equal physical memory on an Apple M2 Pro the filtered multitable leads hashbrown in all $84$ insert, hit, and miss configurations. Multitable is more \textbf{memory-efficient}, at equal mixed-lookup throughput on the map of $4$-byte keys and values the filtered multitable needs up to $12\%$ fewer bytes than hashbrown, and the plain multitable is $18\%$ smaller, holding $\mathbf{22\%}$ more keys in the same memory.
cs.CR / 18 / 2609.39310
XIM: The XDC Interledger Messaging Protocol
Abstract
Distributed ledgers, privacy-preserving institutional networks, and conventional payment systems increasingly need to exchange authenticated messages and settle assets across heterogeneous trust domains. Existing interoperability systems typically optimize for one of three concerns: application-level abstraction, cross-chain message transport, or synchronized execution within a related ledger family. This paper proposes XIM, the XDC Interledger Messaging Protocol, a chain-agnostic protocol for transporting canonical messages across heterogeneous networks while allowing each communication lane to select an explicit verification policy. XIM separates message semantics from transport, verification, execution, routing, asset identity, and compliance metadata. Its cryptographic state is represented by deterministic message identifiers and commitment roots, while an append-only transition log provides auditability and replay protection. XIM introduces a Universal Asset Identifier (UAID), a pluggable adapter interface, lane-scoped security policies, and an optional policy-aware route graph for multi-hop settlement. XDC Network can serve as a coordination and settlement domain without requiring every XIM message or route to transact through XDC. We specify the protocol model, state machine, message encoding, commitment structure, verification modes, failure semantics, security assumptions, threat model, implementation architecture, and an incremental deployment plan. The design targets public blockchains, permissioned ledgers, institutional networks, and authenticated financial-system gateways, with particular attention to stablecoins, tokenized assets, trade finance, and ISO 20022-compatible payment workflows.
cs.CR / 19 / 2609.39352
Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents
Abstract
LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself. In this work, we reveal that untrusted agent decisions fundamentally depend on two conceptually distinct Risk-Realization Factors (RRFs): an actuation factor (specifying what concrete operation is performed) and a pretext factor (providing the situational rationale for why the agent must perform it). Guided by this abstraction, we propose a coordination-based attack paradigm: decoupling pretext from actuation. Rather than fragmenting the malicious actuation, we preserve it as an intact operation within a downstream Steering Skill, while delegating the pretext factor to an upstream Grounding Skill that subtly alters persistent environment artifacts through routine utility operations. The intact actuation thus hides in plain sight, appearing completely legitimate and task-driven only when evaluated against the fabricated pretext. Building on this formulation, we develop an automated framework that discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned skill instructions via runtime closed-loop feedback. Extensive evaluations across single-session and persistent cross-lifecycle scenarios demonstrate that decoupled skill poisoning achieves high attack success, exposing a critical blind spot in isolated Skill security audits. Our automated framework code is available at https://github.com/Wenxin-buaa/CoordPoison.git.
cs.CR / 20 / 2609.39362
Link Inference Attack on Privacy-Preserving Knowledge Graphs
Abstract
Knowledge Graphs (KGs) are widely used to store and share structured information across sensitive domains such as healthcare, fi- nance, and social networks. A common privacy practice is to delete sen- sitive relations before publishing the graph, under the assumption that removing edges is sufficient to prevent their recovery. In this paper, we challenge this assumption and show that even when a relation is fully or partially hidden, its existence leaves structural traces in the public graph that can be exploited to recover it with high accuracy. To this end, we propose a link inference attack that operates on the topology of the public graph, and evaluate it under two privacy scenarios that differ in how the adversary exploits the knowledge available to him. In the first setting where the adversary exploits all topological information, the attack achieves near-perfect discrimination (AP = 0.949, ROC-AUC = 0.999), while in the more realistic one where the adversary makes use of some semantic information, it recovers up to 74% of hidden edges. Build- ing on these results, we further conduct a structural analysis to identify which topological properties of the graph drive the attack success, re- vealing that privacy risk is not uniform across entities and that certain structural patterns make specific relations significantly more vulnerable to inference than others.
cs.CR / 21 / 2609.39414
SEW: Style-Encoded Watermarking of LLM-Generated Code
Abstract
Code watermarking supports provenance tracking for code generated by LLMs. Modifying token selection to embed watermarks as an LLM generates code can create a trade-off between detectability and functional correctness. Other methods instead watermark completed code using predefined transformations or trained neural models. Recurring patterns can make watermark choices predictable across programs, while treating patterns common in unwatermarked code as watermark evidence can cause false detections. We therefore introduce SEW, which embeds and detects watermarks in already generated code through three components: (i) code style rules collected from style guides and transformation rules, with style choices determined by a secret key and each program's structural context; (ii) style-preference calibration, which evaluates watermark evidence using style probabilities estimated from human-written code; and (iii) context-aware style aggregation, which combines evidence from structurally matching locations assigned the same code style choice, preventing repeated applications of that choice from inflating watermark evidence. On CodeContests across three LLMs and three programming languages, SEW achieves a mean relative improvement of 12.44% in TPR@FPR5% over the baselines and is robust to four non-LLM code-editing attacks, with only a 0.94% mean relative decrease. Our code is available at https://github.com/suhanmen/SEW.
cs.CR / 22 / 2609.39450
ActionGuard: Tool Call Authorization under Poisoned Skills
Abstract
LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exfiltration, file deletion, or unauthorized code execution. This paper presents ActionGuard, which inspects skill-influenced Tool Calls immediately before execution. ActionGuard separates the target agent's action-generation context from the safeguard's authorization context. The target agent may use the original skill for planning, but the Reviewer does not receive the potentially poisoned raw skill text. Instead, it determines whether each action is justified by the trusted user request using a balanced skill profile, current and recent Tool Calls, and local script contents. ActionGuard intercepts each Tool Call at OpenClaw's before-tool-call stage and enforces the Reviewer's ALLOW or DENY decision under a fail-closed policy. We evaluate ActionGuard on 139 contextual and 180 obvious injections in a SKILL-INJECT-based setting against Dynamic Guardian and SkillGuard, using three open-source and two commercial Reviewer models. Each condition is repeated three times and evaluated using Attack Success Rate (ASR) and Task Success Rate (TSR). Overall, ActionGuard reduced ASR by 35.54 to 46.11 percent relative to existing safeguards and by 70.44 percent relative to No Safeguard, while maintaining high benign-task completion. These results show that execution-boundary authorization grounded in trusted user intent and runtime evidence can restrict unauthorized Tool Calls induced by skill injection.
cs.CR / 23 / 2609.39584
Cybersecurity in Edge Computing: A Trust-Aware Federated Hybrid Intrusion Detection Framework
Abstract
Edge computing has emerged as a critical computing paradigm in modern distributed systems by migrating data processing closer to end users and Internet of Things (IoT) devices. While this paradigm decentralizes processes, minimizes latency, and reduces backhaul bandwidth congestion, it exponentially enlarges the cyberattack surface. Heterogeneous, resource-constrained edge devices deployed across unmanaged administrative domains present highly vulnerable targets. To address these vulnerabilities without compromising global data privacy regulations, this paper proposes a novel Trust-Aware Federated Hybrid Intrusion Detection Framework (TA-FHIDF). The proposed framework integrates an Autoencoder, a 1D Convolutional Neural Network (1D-CNN), and a Bidirectional Long Short-Term Memory (BiLSTM) model into a unified, localized deep learning engine capable of autonomous spatial and temporal feature extraction. Model training is performed collaboratively via federated learning, ensuring raw network telemetry remains isolated at local gateways. Furthermore, to defend against adversarial model poisoning attacks, we introduce a robust server-side trust-aware aggregation mechanism that evaluates client reliability using a cosine similarity metric before global model integration. Empirical evaluations across multi-vector benchmark datasets (UNSW-NB15, CICIDS2017, and Edge-IIoTset) demonstrate the framework's superior detection accuracy, rapid convergence, and high Byzantine fault tolerance under adversarial attack scenarios.
cs.CR / 24 / 2609.39607
Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Abstract
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
cs.CR / 25 / 2609.39678
Aletheia: Permission-Minimality Testing for Coding-Agent Rules
Abstract
Repository instruction files guide coding agents, but also expose them to prompt injection. Malicious rules can request credential access or data transfer while the agent produces a correct patch. We present Aletheia, a framework for permission-minimality testing. Aletheia translates requested authority into a typed language and synthesizes executable sandbox configurations. It runs the unchanged rule and task under full permissions and independent restrictions that remove one permission at a time. Passing independent functional tests under strictly reduced authority provides a dispensability witness, which Aletheia interprets against task context to diagnose suspicious requests. We formalize synthesis and the conditions connecting witnesses to enforced restrictions. On a shared refactoring task, Aletheia executes and detects all 314 AIShellJack attack inputs, with no alarms on five benign templates. Among 80 manually verified benign GHAgentFiles rules, it raises three false positives (3.75%).
cs.CR / 26 / 2609.39774
Path-Finding, Orbit State Preparation, and the Security of Invariant Quantum Money
Abstract
The security of quantum money from knots, and of its generalization to invariant money, is based on the assumption that path-finding, exhibiting a sequence of moves between two equivalent objects, is hard. No proof of security from that assumption alone is known. The existing proofs add knowledge-of-path assumptions, which assert that any efficient algorithm producing two objects with the same invariant implicitly knows a path between them. No attack can refute such an assumption, and it is not known to follow from security. We ask when path-finding is the right assumption. When each equivalence class is the orbit of an efficiently computable action of a group that can be superposed over, and every move acts as a group element, as for graphs, average-case hardness of path-finding is necessary for security. For knots no such group is known, and a path-finder only reduces forgery to an equally hard state-preparation problem. With or without a path-finder, a forger must prepare a state that verification accepts, and we take the hardness of that task as the assumption. For schemes whose verification walk mixes in polynomial time, the preparation assumption states that no efficient algorithm, given the serial number of a freshly minted banknote and one object measured from it, prepares such a state. It is falsifiable, and it is equivalent to security against forgers that measure their banknote first. The transfer assumption, which security implies, states that measuring first costs a forger at most a polynomial factor. Together the two are equivalent to security, so every proof of security must establish the preparation assumption. If the preparation assumption holds, no fully black-box reduction that calls the forger only at the serial number it is given can derive the transfer assumption from the preparation assumption.
cs.CR / 27 / 2609.39787
Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence
Abstract
Scientific artificial intelligence (AI), spanning foundation models (FMs) to federated data-analysis pipelines, is becoming shared infrastructure across national laboratories, universities, hospitals, and industrial partners. This collaboration creates privacy risks whose natural unit is often an institution's participation, research strategy, or technical capability rather than a single record. Differential privacy (DP), federated learning (FL), secure computation, trusted execution, and provenance each protect parts of the stack, but their guarantees rarely compose across mixed-trust institutions, access tiers, and autonomous agents. This perspective recasts privacy for scientific AI as an assurance problem defined by six elements: protected asset, observer, channel, permitted disclosure, guarantee, and evidence. We demonstrate the framing through a claim register for a composite cross-institutional scenario and use it to assess the model lifecycle. Two of the resulting gaps are specific to leadership-class facilities: scheduler, allocation, and telemetry metadata expose an institution's resource posture, and instrument-attached control loops leak research strategy through timing and contention on shared accelerators. We identify six research priorities: institution-level guarantees, agent-communication privacy, cross-tier information flow, privacy-compatible reproducibility, leadership-scale accounting, and instrument side channels. The contribution is a common form for stating, comparing, and auditing claims whose guarantees otherwise remain fragmented across the scientific AI stack.
cs.CR / 28 / 2609.39768
Security Properties of Neural Networks as Decision Problems
Abstract
Certifying a deployed neural network raises decision problems that the verification literature has not classified: whether the model carries a backdoor planted in its training data, whether a fault in its stored parameters can drive it into an unsafe state, whether its output leaks a private part of its input. We formalise eight such problems and classify what we can. The organising observation is a logical one. The function computed by a piecewise linear network, together with all its node values, is definable by a quantifier-free formula of real addition of size linear in the network, so a property of the network is a quantifier-alternation sentence, which Sontag's 1985 theorem places in the polynomial hierarchy at the level of its prefix. Membership results are thus corollaries, and the argument makes plain what they need: that the quantified objects are inputs rather than the network's own parameters. Non-interference, monotonicity and counterfactual fairness have exactly the complexity of network equivalence and of interval verification, all co-NP- complete over ReLU. Detection of backdoor triggers from a quantised alphabet is Sigma_2^P-complete, one level above robustness certification, so it does not reduce to polynomially many robustness queries unless the hierarchy collapses. Inversion resistance is co-NP-complete for every l_p metric, p a fixed positive integer. Quantifying over parameters instead of inputs - the fault model of bit-flip attacks, radiation upsets and analog accelerators - makes verification exists-R-complete already for networks of identity nodes, for which every previously studied problem is in P, and it stays so when each parameter is confined to a box of inverse-polynomial width; the corresponding safety question is forall-R-complete for ReLU.
cs.CR / 29 / 2609.39178
Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
Abstract
Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since these models are designed to interact directly with the physical world and humans, their security is critical, and even small vulnerabilities can lead to catastrophic failures. In this work, we propose the Universal Adversarial Object, a sphere with optimized surface texture that significantly degrades task success rates when placed within the robot's field of view. Specifically, our approach introduces a multi-level attack framework that jointly disrupts trajectory planning, task execution, and action control. We validate our method in both simulated and real-world robotic settings. Experimental results demonstrate that the adversarial object reduces the average task success rates by 31.2%-39.9% for two representative VLA models (Pi0 and RDT), with success rates dropping to near zero in complex scenarios. Index Terms--Vision-Language-Action models, adversarial attack, robotic security, universal adversarial object
cs.CR / 30 / 2609.38387
On Removing Interaction from Quantum Proofs
Abstract
An important open question in quantum cryptography is the construction of publicly-verifiable NIZKs for QMA. Classically, one can construct NIZKs for NP in the random oracle model (and sometimes in the standard model) by compiling an honest-verifier ZK (HVZK) $Σ$-protocol for NP using the Fiat-Shamir transformation. Broadbent and Grilo introduced a quantum analog of a $Σ$-protocol (which they call a $Ξ$-protocol) in which the prover's first message is quantum, and show that HVZK $Ξ$-protocols exist for QMA. However, it is not clear how to compile such protocols into NIZKs in the (Q)ROM, because the Fiat-Shamir transformation seems to be incompatible with quantum messages. In this work we give formal evidence that this is indeed the case: we show that if generic "Fiat-Shamir-like" compilers for quantum protocols exist in the QROM (with small completeness and soundness error) then QMA = BQP.
cs.CR / 31 / 2609.38720
Does the Readout Bypass Leak the Input? A Feature-Visibility Audit of Hybrid Quantum-Classical Models
Abstract
Readout-side residual hybrids concatenate raw inputs with measured quantum features. Under single-example gradient sharing, a biased first linear layer admits standard analytic recovery of its input, so the bypass exposes raw coordinates without requiring inversion of the quantum circuit. We audit this mechanism using two tabular datasets, four architectures, and metrics conditioned on feature visibility. Iterative gradient matching gives median full-record PSNR of 73-96 dB for residual and input-only heads. Quantum-only heads score 8-11 dB on the full record but 54-96 dB on the six input coordinates they actually encode. These are reconstruction results for the tested six-input, six-observable circuits, not a general statement about quantum encodings. A loss-threshold membership attack remains near chance. The contribution is a visibility-conditioned privacy audit: omitted coordinates must not be credited as protection supplied by quantum processing, and near-exact PSNR differences must not be interpreted as meaningful privacy rankings. Our findings concern individual gradients and do not establish leakage under aggregation or multiple local training steps.
cs.CR / 32 / 2609.38894
Quantum Time-Lock Puzzles in the Quantum Random Oracle Model
Abstract
A time-lock puzzle allows a sender to hide a message in a puzzle such that recovering the message requires substantially more sequential computation than the time required to generate the puzzle, even when parallel computation is allowed. Applications of time-lock puzzles include timed-release encryption, sealed-bid auctions, electronic voting, fair contract signing, coin flipping, and Byzantine consensus. However, time-lock puzzles are known to be impossible in the classical random oracle model. To overcome the classical barrier, in this work we consider quantum time-lock puzzles, in which the puzzle itself is a quantum state. Our construction in the quantum random oracle model achieves generation in one oracle round, solving in at most $T$ oracle rounds, and security against polynomial-width quantum adversaries of depth $o(T)$ for every polynomially bounded delay $T=T(λ)$, resolving an open problem posed by Mahmoody, Moran, and Vadhan (2011).
cs.CR / 33 / 2609.38941
Tight Parallel Repetition for Private-Coin Arguments
Abstract
We show that assuming the existence of homomorphic encryption, parallel repetition of all interactive arguments (after being run under homomorphic encryption) reduces the soundness error at a tight exponential rate even in the post-quantum setting. Moreover, we generalize this result to hold for threshold verifiers, where the parallel repeated verifier accepts if and only if at least $t$ of the executions are accepted (for some threshold $t$). Prior to this work, these results were known only when the cheating prover was assumed to be classical, and it was not known how to achieve tight bounds. As a corollary, we construct the first constant-round succinct argument for $\mathsf{QMA}$ with negligible completeness and soundness errors assuming only the existence of quantum homomorphic encryption.
cs.CR / 34 / 2609.38945
Certified Randomness with Optimal Rate
Abstract
The generation of certified random bits is an emerging near-term application of quantum computers. Potential applications, such as randomness beacons and CRS generation, require nearly uniform randomness, whose rate (the ratio of min-entropy to bitlength) is ~ 1. However, existing protocols for certified randomness either produce weakly random strings with rate o(1), or they require a trusted random seed. We show how to certify randomness with the optimal rate ~ 1 without requiring any trusted randomness from the verifier. Our protocol is secure unconditionally in the quantum random oracle model (QROM). This improves the result of Coladangelo et al., who certified weakly random strings unconditionally in the QROM, and answers a question posed by Aaronson and Hung. We also define and construct a proof of conditional min-entropy, which certifies optimal min-entropy even conditioned on an adversarially chosen transcript. We show an application of this primitive to randomness beacons, where each pulse should have high min-entropy conditioned on all previous messages.
cs.CR / 35 / 2609.39234
Breaking the Bounded Entanglement Barrier for Quantum Position Verification
Abstract
Position verification, introduced by Chandran et al. (SIAM J. Computing 2014), allows verifiers to test a prover's claimed position by an interactive protocol. Classical position verification is impossible. Even for Quantum Position Verification (QPV), there always exists an LOCC (Local Operations and Classical Communication) attack if the adversary can hold an exponentially large amount of preshared entanglement. Somewhat surprisingly, we show that we can circumvent this barrier in the idealized continuous-time model, where time is represented by a real-valued parameter and challenge messages can be sent at a time sampled uniformly from a real interval. In this model, we give a BB84-based protocol that remains secure against any finite coalition of LOCC adversaries with arbitrary finite (possibly exponential) quantum storage and entanglement. Our construction can also be instantiated in the discrete-time model. Even though the previous impossibility results apply, we are able to obtain an information-theoretic QPV protocol in which the honest parties' total resources (communication and storage) can be significantly smaller than the adversarial resource bound. In fact, we can achieve any desired polynomial gap between honest parties and adversarial resources. To our knowledge, all previous protocols in the literature required resources of the honest parties to be at least as large as the adversarial entanglement. Interestingly, the resource gap between the honest and adversarial party resources in our protocol depends on how precisely time can be measured. As the precision of the best-known clock improves with further research, the resource gap in our protocol keeps increasing. In particular, the adversarial resource bound keeps increasing while the honest parties' resources remain largely the same. We call this the "I sleep, you work" paradigm.
cs.CR / 36 / 2609.39495
On quantum interactive proofs with a laconic prover
Abstract
Interactive proof systems with a laconic prover, studied by Goldreich, Vadhan, and Wigderson (CC, 2002), capture problems verifiable with logarithmic prover communication in the classical setting. For two-message quantum analogs, even a single-bit prover response contains quantum statistical zero-knowledge ($\sf QSZK$), introduced by Watrous (FOCS 2002). However, restricting the verifier's question to classical public coins collapses the corresponding class to $\sf BQP$, as shown by Beigi, Shor, and Watrous (ToC, 2011). We further study two-message quantum interactive proof systems with a laconic prover. To this end, we introduce the class ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$, where $\ell$ is the length of the prover's response, and establish: 1. A natural complete characterization of ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ by Multi-State Distinguishability. In particular, Quantum State Distinguishability (QSD) is ${\sf QIP}_{\rm bit}$-complete. Since QSD is $\sf QSZK$-hard, our result places ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$, for $\ell\geq 2$, in a landscape "just above" $\sf QSZK$. 2. Easy regimes for ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ collapsing to $\sf QSZK$. We prove that QSD$[a,b]$ (and thus ${\sf QIP}_{\rm bit}[a,b]$) is in $\sf QSZK$ when $a(n)-b(n)\geq 1/O(\log n)$, and combine this with an answer compression from ${\sf QIP}_{\ell\text{-}{\rm bit}}[2,c,s]$ to ${\sf QIP}_{\rm bit}$ to obtain another easy regime when $2c>(1+2^{\ell/2})s$. Remarkably, our improved polarization applies to SD and $\sf SZK$, resolving an open problem in Sahai and Vadhan (JACM, 2003). 3. Quantum public coins also make the interaction useless: ${\sf qc}\text{-}{\sf QAM}[O(\sqrt{\log{n}})]$ with constant gap is in $\sf BQP$, where ${\sf qc}\text{-}{\sf QAM}[\ell]$ is a subclass of ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ in which the verifier's question is exactly halves of EPR pairs.
cs.CR / 37 / 2609.39589
Rate 1/5 Non-Malleable Codes against Entangled Split-State Tampering
Abstract
We construct efficient information-theoretic non-malleable codes for classical messages that are secure against two noncommunicating local quantum tampering operations with arbitrary pre-shared entanglement. For every sufficiently small fixed $ξ>0$ and all sufficiently large first-share lengths $n$, the codes have rate at least $1/5-ξ$, perfect correctness, and error $2^{-n^{Ω(1)}}$. Security holds for every message, with a single message-independent simulator for each attack. This resolves the constant-rate question for worst-case classical messages in the entangled two-split-state model. Our construction builds on the permutation-based two-split construction of Batra, Boddu, and Jain, which achieves rate approaching $1/5$ for uniformly random messages. We retain their architecture but replace the uniform message input to the permutation with a prescribed message concatenated with fresh uniform padding. Our main contribution is a worst-case security reduction for this modification.
cs.CR / 38 / 2609.39844
Natural Barriers to Quantum Extraction: On the Post-Quantum (In)security of (O)EKE and Masny-Rindal OT
Abstract
Encrypted key exchange (EKE), introduced by Bellovin and Merritt (IEEE S\&P 1992), and Masny-Rindal OT, introduced by Masny and Rindal (ACM CCS 2019), are highly-efficient methods for compiling essentially any KEM into advanced cryptographic protocols, namely password-authenticated key exchange (PAKE) and oblivious transfer (OT), by relying only on idealized symmetric-key primitives. They have become leading candidates for practically-implementable PAKE and OT due to (1) their simplicity, (2) their plug-and-play nature, allowing for flexibility in the choice of KEM, and (3) existing proofs of UC-security (in the classical adversarial model). Due to point (2) above, these compilers yield attractive candidates for efficient \emph{post-quantum} PAKE and OT, especially given the recent post-quantum KEM standardization efforts. This motivates the question of whether the (UC-)security of these compilers translates to the quantum adversarial model. In this work, we show that it does not. In particular, we prove that a general family of (O)EKE protocols, as well as Masny-Rindal OT, are \emph{not} UC-secure against quantum polynomial-time adversaries, even when instantiated with a post-quantum KEM. To establish UC-insecurity, we devise an adversarial strategy that provably thwarts any attempt by the simulator to extract its input (the password in the case of PAKE, and the receiver's choice bit in the case of OT). To complement these negative results, we establish that both compilers yield certain notions of \emph{game-based} security. Along the way, we establish a novel ``advantage-tight'' one-way to hiding lemma that may be of independent interest.
cs.CR / 39 / 2609.39918
Verifiable quantum advantage based on polynomials with planted structures
Abstract
A central question in the theory of quantum advantage is whether there are quantum advantage protocols with similar resource requirements as random circuit sampling that are also verifiable just from the classical outputs of the quantum computation. Here, we develop the idea of simulation secrets for verifiable advantage. A verifier can use a simulation secret to evaluate a cross-entropy test faster than it would take a classical adversary to pass the test. We instantiate this idea using IQP circuits described by cubic polynomials with planted independent spaces. These correspond to the largest independent set in the orbit of a polynomial under the general linear group and yield a low-rank stabilizer decomposition of the corresponding state. We conjecture that large independent spaces are invisible to a computationally bounded adversary, and therefore they cannot exploit them to pass the protocol. A second conjecture regards the fine-grained complexity of producing samples that pass the cross-entropy test for uniformly random polynomials. Under these conjectures, our scheme results in a polynomial gap between the verification time and the time a classical adversary would need to pass the protocol---both are exponential. It has a potential application to generating classically certifiable randomness, since the output distributions have high min-entropy. We estimate that the planted polynomial scheme is implementable using 100 logical qubits at logical error rates around $10^{-6}$.
cs.CR / 40 / 2609.40062
A Fourier-Label Information-Loss Barrier for Dihedral Coset Algorithms
Abstract
We establish a no-go theorem for a broad class of quantum algorithms for the dihedral coset problem (DCP). We consider the Fourier-sampling and subset-sum-measurement template proposed by Regev (SIAM Journal on Computing, 2004), which is one of the main approaches to solving DCP. Suppose that, after measuring the lower $n-1$ bits of the subset sum, the algorithm discards any $ω(\log n)$ bits from each of the Fourier labels. Then we prove that the algorithm cannot succeed in solving DCP. This shows that any algorithm following this template must make extensive use of the Fourier labels, and thus serves as a useful guide for developing algorithms for DCP. As a main application, we show that the recent algorithm by Simon (IACR ePrint:2026/1591, August 11 2026) does not solve DCP. We show that after the subset-sum measurement, this algorithm can be implemented (up to exponentially-small error) using only the most-significant third of the Fourier labels, and is therefore subject to our general no-go theorem. To help with verifiability, we release Lean 4 code for our results.
cs.CR / 41 / 2609.40194
Query-Limited RAM Programs and their Applications
Abstract
Quantum one-time programs (Broadbent, Gutoski and Stebila, CRYPTO 2013) or OTPs for short, enable a functionality to be encoded into a quantum token that can be evaluated on a single chosen input and then becomes unusable. While powerful, this primitive is inherently stateless and tied to a setting in which a quantum token needs to be issued and distributed for every single evaluation of a circuit. This raises a natural question: can the one-time computation paradigm be extended to richer, stateful forms of controlled access, and would such an extension offer inherent advantages beyond standard OTPs? We introduce $\textit{query-limited RAM programs}$ (QLPs), a RAM-generalization of QOTPs that supports structured, stateful computation under bounded or policy-driven access. QLPs allow controlled sequences of evaluations while preventing adversarial forking or rollback of computational state. This enables new applications beyond stateless one-time programs, including quantum tokens for Turing Machines whose size depends only on $\textit{code length}$ (and not runtime), transferable $k$-time or budget-limited programs, and low-communication mechanisms for delegating computation in settings such as Software-as-a-Service. To construct QLPs, we introduce $\textit{one-shot programs}$, unifying one-shot signatures (Amos, Georgiou, Kiayias and Zhandry, STOC 2020) with the single effective query paradigm (Gupte, Liu, Raizes, Roberts and Vaikuntanathan, STOC 2025). We prove that one-shot programs generically imply query-limited programs, demonstrating that the strengthened unclonability guarantees of one-shot signatures translate into enhanced functionality. Along the way, we clarify the relationship between signature-token primitives and quantum one-time programs via generic constructions, essentially showing that one-time signing programs imply one-time general computation.
cs.CR / 42 / 2609.40289
Verifiable Quantum Advantage and Computation via Quantum Circuit Obfuscation
Abstract
We construct protocols for classically verifiable quantum advantage and classical verification of $\mathsf{BQP}$ computations using \emph{quantum indistinguishability obfuscation} (qiO). Specifically, given qiO and assuming a slightly stronger version of $\mathsf{BQP}\neq\mathsf{BPP}$, we construct a two-message quantum-advantage protocol that is efficiently and publicly verifiable. Our result can be viewed as a rigorous cryptographic foundation for the heuristic quantum advantage proposals based on \emph{peaked random circuit sampling} of Aaronson and Zhang (arXiv:2404.14493). We also construct two simple protocols for classically verifying arbitrary $\mathsf{BQP}$ computations. The first protocol is privately verifiable and assumes only the existence of qiO. This gives a rare example of a nontrivial cryptographic application of (quantum) iO that does not make additional computational hardness assumptions. The second protocol additionally assumes post-quantum one-way functions and is \emph{publicly verifiable}. To our knowledge, this is the first publicly verifiable protocol for classical verification of $\mathsf{BQP}$ computations under computational assumptions in the standard model. We show that all our results hold when qiO is assumed only for ancilla-free unitary circuits. As evidence supporting this assumption, we prove a worst-to-average-case reduction for obfuscating such circuits. This reduction extends the local-mixing framework of Canetti, Chamon, Mucciolo and Ruckenstein (TCC 2024) under quantum analogues of their assumptions.
cs.CR / 43 / 2609.40301
Need for Coherent Access in Constructing Quantum Cryptography
Abstract
We construct quantum oracles relative to which quantum-secure one-way functions (OWFs) exist but pseudorandom states (PRSs) with superlogarithmic output length do not. At first glance, this appears to contradict the known black-box constructions of PRS generators from quantum-secure OWFs. The distinction lies in the access model to the oracles; our oracle separation uses \emph{classical-accessible} random oracles that can be accessed only classically even by quantum algorithms. In fact, our impossibility of PRSs applies to \emph{any} classically accessible classical oracle in place of the random oracle, while keeping the other oracle component unchanged, showing the need for coherent access in constructing PRSs. We further show that logarithmic output length pseudorandom function-like states (PRFSs) exist relative to our oracles, giving an oracle separation between classically accessible logarithmic length PRFSs and superlogarithmic length PRSs. This shows that fully black-box PRS length extension from logarithmic to superlogarithmic output length must use coherent access to the underlying short PRS.
cs.CR / 44 / 2609.40321
Exponential quantum speedup for $\mathbb{F}_3^n$-Subset-Sum? Or, rigorous classical algorithms for Binary-Error LWE
Abstract
We study vector subset sum over $\mathbb{F}_3^n$: given $m$ random vectors from $\mathbb{F}_3^n$, find a nonempty subset that sums to zero; the smaller $m$, the more difficult it is to find such a subset. Chen, Liu, and Zhandry (EUROCRYPT'22) introduced an efficient quantum algorithm that solves this problem when $m\approx n^2/2$, where a naive classical algorithm would require exponential time. Subsequently, Kothari, O'Donnell, and Wu (STOC'2026) gave an efficient classical algorithm that only requires $m \approx n^2/3$ vectors, thus removing the hope for an exponential quantum advantage in this parameter regime. Using the framework of Chen, Liu, and Zhandry, we give quantum algorithms that require much fewer input vectors, renewing the possibility of an exponential quantum speedup: for any fixed $ε>0$, our quantum algorithm solves $\mathbb{F}_3$-subset sum in polynomial time with $m=ε\cdot n^2$ vectors. More generally, we establish a full sample--time tradeoff that interpolates between exponential and polynomial runtime. The main ingredient is a deterministic classical algorithm for the binary-error Learning-with-Errors problem, which is of independent cryptographic interest. For this, we rigorously establish a sample--time tradeoff that was predicted by earlier algebraic heuristics. For vector subset sums over larger fields, we also significantly improve classical algorithms in Kothari, O'Donnell, and Wu (STOC'2026).
cs.CR / 45 / 2609.40350
On The Simplest Quantum-Secure Block Cipher
Abstract
Pseudorandom permutations are ubiquitous in theoretical and applied cryptography. PRPs that offer security even against adversaries making quantum queries are of increasing interest, and used in applications ranging from constructing pseudorandom unitaries to separating SZK from BQP. A successful framework for constructing classically-secure PRPs is the key-alternating Even-Mansour approach, which interleaves applications of public permutations with additions of round keys. The single-round construction is already classically secure in the ideal permutation model (IPM), with added rounds offering improved concrete security. However, in the quantum-query setting, the status of this framework is presently unclear. A simple quantum-query attack based on Simon's algorithm breaks the one-round cipher. For two or more rounds, security is only known against non-adaptive adversaries who must prepare all queries in advance. In this work, we show that the two-round Even-Mansour cipher is information theoretically secure in the IPM against adversaries making polynomially-many adaptive forward and inverse quantum queries to all available oracles. Our proof uses compressed permutation oracles and a specially crafted isometry relating the ideal and real experiments. We also show that this construction is minimal, in the sense that essentially any cipher constructed via a single call to a public permutation is quantumly insecure.