← Back to Index
Daily Research Digest

arXiv Papers

2026-09-29
874
Papers
9
Categories
210
Translated
收藏清单 0
精选 · Favorites
210
cs.AI / 1 / 2609.33878
Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
使用两个置信度门控的本地 LLM 评审器策管商户匹配训练数据
Donghao Huang, Jinling Pei, Zhaoxia Wang
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.
Chinese Translation
商户匹配将一个含噪的支付描述符解析为一个检索到的商户实体,或者返回无匹配。在策管训练标签时,一个关键挑战是区分教师模型的弃答与不存在可接受实体的证据:错误的无匹配标签会污染伪标注数据,而保守的标注则会降低覆盖率。我们研究两个本地大语言模型评审器之间的一致性是否能够提高伪标签的可靠性。仅当评审器一致时才保留该标签,并为选择与弃答分别设置有序阈值,从而保证正标签集与负标签集互不相交。在 2,000 条专家标注查询上的回溯式重放表明,较高的选择阈值可以提高正标签纯度,而较高的弃答阈值则会增加错误的无匹配标签。在阈值 (0.86, 0.80) 下,Muse Glimmer 30B 与 Gemma 4 31B 共同标注了 1,633 条查询(覆盖率 81.7%),纯度达 96.88%;正标签纯度与负标签纯度分别为 99.47% 和 93.38%。这比任一组成模型在相同阈值下的表现高出两个百分点以上,但覆盖率更低。折半检验发现阈值选择带来的乐观偏差仅为 0.14 个百分点。对称阈值 0.86 会额外引入 40 个错误的无匹配标签,而即使不设置信度阈值,仍有 46 个错误的弃答持续存在。在五项匹配的模型内对比中,更高的推理投入并未带来明显的 F0.5 增益,却使中位延迟增加 1.8–5.0 倍。这些结果说明有必要对正、负伪标签分别进行阈值设定与审核。本研究确立的是标签纯度,而非学生模型的效用;要证明下游价值,仍需进行新数据的策管和学生模型的微调。
cs.AI / 2 / 2609.33920
HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents
HyperMCTS:面向长时程 LLM 智能体的超图增强 MCTS
Tingsong Xiao, Nithish Balachandar Moudhgalya, Chandrayee Basu, Lichao Wang, Luyang Kong, Benjamin Z. Yao, Zhe Jiang, Jie Hao
cs.AI
large language model
大语言模型相关
Abstract
Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment interaction make search costly. Efficient search therefore requires effective reuse of trajectory feedback. Standard MCTS maintains prefix-specific statistics, without explicitly accumulating outcomes for decision groups that recur across different paths. To fill this gap, we propose HyperMCTS, a training-free method that augments an ordered MCTS tree with a cross-trajectory hypergraph. Hyperedges represent groups of canonical decisions and accumulate their observed returns within the current task. Our hypergraph-guided HyperUCT selection rule aggregates evidence from overlapping hyperedges into an action prior, allowing outcomes collected under one prefix to inform selection under another while preserving execution histories in the tree. On DeepPlanning, HyperMCTS improves average planning accuracy by 2.3--7.3 percentage points over the strongest baseline for each of three backbone models. It enables Qwen3.6-27B to outperform Claude Opus 4.6 (max) on Shopping Planning, while achieving higher accuracy with fewer LLM calls and output tokens than the evaluated MCTS-based baselines. SealQA experiments further demonstrate improvements in question answering.
Chinese Translation
长时程任务要求大语言模型(LLM)智能体在跨越整个解决方案的约束下协调决策。蒙特卡洛树搜索(MCTS)通过探索替代动作轨迹,为测试时扩展提供了一种有前景的方法,但模型计算和环境交互使搜索成本高昂。因此,高效搜索需要有效复用轨迹反馈。标准 MCTS 维护前缀特定的统计量,而没有显式地累积在不同路径中重复出现的决策组的结果。为填补这一空白,我们提出 HyperMCTS,这是一种无需训练的方法,它用跨轨迹超图增强有序 MCTS 树。超边表示规范决策的组,并在当前任务内累积其观测到的回报。我们的超图引导的 HyperUCT 选择规则将来自重叠超边的证据聚合为动作先验,允许在一个前缀下收集的结果为另一个前缀下的选择提供信息,同时在树中保留执行历史。在 DeepPlanning 上,对于三个骨干模型中的每一个,HyperMCTS 将平均规划准确率相比最强基线提高了 2.3--7.3 个百分点。它使 Qwen3.6-27B 在 Shopping Planning 上优于 Claude Opus 4.6 (max),同时与所评估的基于 MCTS 的基线相比,以更少的 LLM 调用和输出 token 实现了更高的准确率。SealQA 实验进一步证明了在问答方面的改进。
cs.AI / 3 / 2609.34039
Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
面向结构化临床数据分析的大语言模型:双智能体接地与验证
Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten, P. Ellen Grant, Yangming Ou
cs.AI
large language model
大语言模型相关
Abstract
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
Chinese Translation
目的:开发并刻画 CLEAR-Med,这是一个用于结构化临床数据自然语言分析的双智能体框架,其将基于 SQL 的调用与独立验证相分离。方法:CLEAR-Med 使用一个智能体将问题翻译为可执行的结构化查询语言(SQL),保留已执行的查询与数据库结果,并生成一份草稿。随后,确定性检查与一个单独调用的跨供应商验证智能体(Validation Agent)会接受该草稿、请求一次有界修复,或选择弃权。我们将该系统形式化为一个有界选择性流水线,并在一个包含 25 个查询的开发基准上评估了 CLEAR-Med 的配置与可扩展性,以及调用智能体(Invocation Agent)的准确性与一致性;该评估使用一张经过协调统一的、涵盖 21 个中心的新生儿缺氧缺血性脑病数据表,其中包含 532 条去标识化婴儿记录和约 1,300 个变量。结果:CLEAR-Med 完成了全部六种标称可扩展性配置,其中包含 500x1300。在重复五次的 25 个开发基准查询中,调用智能体在 125 次回答中正确回答了 83 次(66.4%;查询簇自助法 95% CI,48.0-83.2%),而未接地的 ChatGPT 基线为 125 次中的 15 次(12.0%;95% CI,3.2-22.4%),配对改善为 54.4 个百分点(95% CI,36.8-72.0%)。结论:CLEAR-Med 为结构化临床数据的可追溯分析提供了一种通用架构:数值性断言始终与已执行的 SQL 保持关联,而未解决的案例可以以失败关闭(fail closed)的方式处理。所报告的各项实验刻画了 CLEAR-Med 的配置与可扩展性以及调用智能体的准确性,而形式化分析则确立了完整控制流的编码属性保证;对验证与弃权阶段开展前瞻性的全流水线评估是这项工作的下一阶段。
cs.AI / 4 / 2609.34069
Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
迈向证书驱动的软件移植:一种用于科学程序优化的自我改进智能体框架
Piyush Jha, Aishik Ghosh, Vijay Ganesh
cs.AI · cs.PL · cs.SE
large language model
大语言模型相关
Abstract
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.
Chinese Translation
大型科学代码库的升级与重写历来是一项重大挑战。尽管基于大语言模型(LLM)的演化搜索可以对遗留代码进行移植与加速,但仅在提示中进行修复反馈并不能阻止后续候选方案重复相同的错误。我们提出了证书驱动演化搜索(Certificate-Driven Evolutionary Search, CDES),它通过从失败候选中导出的可强制执行约束来扩展演化搜索,这些约束被记录为假设证书、检查器证据和经论证的限制。其控制逻辑通过拒绝、回溯和有针对性的修复来强制执行这些限制,同时保留兼容的编辑。我们将 CDES 应用于 Geant4 工具包中两个粒子模拟函数的 CPU 到 GPU 转换,并使用一套超越单元测试的评估框架进行评测,该框架结合了形式化检查、数值比较、物理检查和 GPU 安全性测试。生成的实现相比 CPU 代码在函数级别分别实现了 13.78 倍和 23.54 倍的加速,其中包括数据转换与传输;对于其中一个函数,GPU 吞吐量比专家实现高出 14.9%,当互补组件结合时达到 16.1%。在一项针对执行设置的消融实验中,证书反馈使通过所需正确性检查的候选比例从 55% 提高到 90%。
cs.AI / 5 / 2609.34072
PhysFieldBench: Can Multimodal Models Understand Physical Fields?
PhysFieldBench:多模态模型能否理解物理场?
Yuezhou Ma, Huikun Weng, Jialong Wu, Chenyi Zhao, Hang Zhou, Haonan Shangguan, Jianmin Wang, Mingsheng Long
cs.AI
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.
Chinese Translation
多模态大语言模型(MLLMs)正日益被设想为科学与工程智能体的核心组件,然而它们解释物理场的能力仍然鲜为人知。现有的物理基准在很大程度上强调教科书式问题求解或直观物理推理,尚未回答 MLLMs 能否从连续场观测中推断出具有物理意义的信息。我们提出 PhysFieldBench,这是一个包含 24 个任务和 1,160 个评估样例的基准,涵盖受控方程场、模拟物理场和观测物理场。这些任务评估三种推理形式:识别物理机制、比较潜在控制变量以及预测结果属性。在具有代表性的开源和专有 MLLMs 上,零样本性能很低:最佳模型取得 29.3 的随机归一化得分,而若干开源模型仍接近随机水平。相比之下,一个任务特定的监督视觉 Transformer 表现显著更好,表明输入中包含可学习的物理信息。为诊断这些失败,一项结构化的自我解释分析将大多数错误归因于遗漏的视觉模式和错误的视觉到物理映射。进一步,为了探索后训练能否提升物理推理并泛化到未见任务,我们比较了使用最终答案或思维链监督的监督微调,以及强化学习。最终答案监督总体上表现最好,但迁移效果较差,而思维链监督之后的强化学习取得了最佳泛化。总之,这些发现突显了需要改进视觉到物理的 grounding 和跨任务泛化,以使 MLLMs 能够在科学与工程工作流中可靠地解释物理场。
cs.AI / 6 / 2609.34079
GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation
GenoMorph:通过自适应潜在计算实现的基于通路的基因组疾病推理
Tanmoy Kanti Halder, Akash Ghosh, Arijit Roy, Sriparna Saha
cs.AI · q-bio.GN
large language model
大语言模型相关
Abstract
Large language models (LLMs) have demonstrated strong capabilities in biological reasoning; however, genomic disease inference remains largely dependent on memorized gene-disease associations rather than understanding biological pathways. This shortcut learning undermines robustness and generalization, and breaks down when molecular identifiers are unavailable. We present GenoMorph, a multimodal genomic reasoning framework that shifts disease prediction from associative gene-disease mapping toward pathway-grounded reasoning. GenoMorph couples a frozen DNA foundation model with question-conditioned cross-attention fusion, self-adaptive latent reasoning (LatentSp), a residual reasoning gate for iterative genomic evidence reinjection, and rejection sampling fine-tuning regularized by hierarchical optimal transport (OT). Rather than learning direct gene-disease mappings, GenoMorph aligns genomic sequence representations with latent pathway dynamics, enabling reasoning trajectories that follow molecular interactions before producing disease predictions. LatentSp dynamically allocates computation according to reasoning confidence, reducing unnecessary reasoning steps and improving inference efficiency. We further construct an anonymized benchmark from the Kyoto Encyclopedia of Genes and Genomes (KEGG), replacing every gene and molecular identifier with anonymous symbols while preserving sequences and pathway topology, thereby removing memorization shortcuts. GenoMorph raises the weighted F1 from 0.7863 (BioReason) to 0.9412, and rejection sampling fine-tuning with self-adaptive latent reasoning pushes it to 0.9725 while cutting latency nearly 60%. On the anonymized benchmark it reaches 0.9465 F1, substantially outperforming prior systems and confirming that accurate disease prediction can arise from pathway reasoning rather than memorized gene-disease associations.
Chinese Translation
大型语言模型(LLMs)已在生物推理中展现出强大能力;然而,基因组疾病推断仍在很大程度上依赖于记忆的基因-疾病关联,而非理解生物通路。这种捷径学习削弱了鲁棒性和泛化能力,并且在分子标识符不可用时失效。我们提出 GenoMorph,一个多模态基因组推理框架,它将疾病预测从关联性基因-疾病映射转向基于通路的推理。GenoMorph 将冻结的 DNA 基础模型与问题条件化的交叉注意力融合、自适应潜在推理(LatentSp)、用于迭代基因组证据重注入的残差推理门控,以及由分层最优传输(OT)正则化的拒绝采样微调相结合。GenoMorph 不是学习直接的基因-疾病映射,而是将基因组序列表示与潜在通路动态对齐,使推理轨迹在产生疾病预测之前遵循分子相互作用。LatentSp 根据推理置信度动态分配计算,减少不必要的推理步骤并提高推理效率。我们进一步从京都基因与基因组百科全书(KEGG)构建了一个匿名化基准,将每个基因和分子标识符替换为匿名符号,同时保留序列和通路拓扑,从而消除记忆捷径。GenoMorph 将加权 F1 从 0.7863(BioReason)提升至 0.9412,而结合自适应潜在推理的拒绝采样微调将其推高至 0.9725,同时将延迟降低近 60%。在匿名化基准上,它达到 0.9465 F1,显著优于先前系统,并证实准确的疾病预测可以源自通路推理,而非记忆的基因-疾病关联。
cs.AI / 7 / 2609.34082
K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings
K-OPSD:用于在 AEC 图纸上对视觉语言模型进行后训练的可验证在策略自蒸馏
Yunfei Bai, Enrico Chionna, Akash Amol, Kawaljit Singh KC, Joern Tinnemeyer
cs.AI · cs.CV
large language model
大语言模型相关
Abstract
Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillation (OPSD) with verifiable supervision, we construct a teacher from the model's own best-of-N generations, certified by a process-level verifier, and rescue failed prompts by resampling under a hint that exposes the verified answer. We then perform an on-policy model update by training on verified completions with a cross-entropy inner-loss, outperforming the bounded token-wise generalized Jensen-Shannon divergence (JSD) used by on-policy distillation. Using K-OPSD, we fine-tune Qwen3-VL models on the AECV-Bench dataset. The resulting models attain the top average judge score (0.819) and combined accuracy (0.738), achieving competitive results against open-source baseline models. The recipe transfers to the out-of-domain ArchCAD dataset, where the 8B model gains most. We present the verifier suite and the continual learning and self-improving pipeline, our results provide preliminary evidence that verifier-guided self-distillation is a promising route toward more reliable machine reading of architecture drawings.
Chinese Translation
解读建筑、工程与施工(AEC)图纸对于通用多模态大语言模型(MLLMs)和视觉语言模型(VLMs)而言很困难。我们提出 K-OPSD,一种用于提升 AEC 图纸理解的 VLM 后训练方法。基于带有可验证监督的在策略自蒸馏(OPSD),我们从模型自身的最佳 N(best-of-N)生成中构建教师,并由过程级验证器进行认证,同时通过在暴露已验证答案的线索下重新采样来挽救失败的提示。然后,我们通过在经过验证的补全上以交叉熵内部损失进行训练,来执行在策略模型更新,其表现优于在策略蒸馏所使用的有界逐 token 广义 Jensen-Shannon 散度(JSD)。使用 K-OPSD,我们在 AECV-Bench 数据集上微调 Qwen3-VL 模型。所得模型取得了最高的平均评判分数(0.819)和综合准确率(0.738),相对于开源基线模型取得了有竞争力的结果。该配方迁移到域外的 ArchCAD 数据集,其中 8B 模型获益最多。我们提出了验证器套件以及持续学习与自我改进流水线;我们的结果为以下观点提供了初步证据:验证器引导的自蒸馏是通往更可靠机器阅读建筑图纸的一条有前景的路径。
cs.AI / 8 / 2609.34137
You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning
鱼与熊掌不可兼得:概念纠缠限制了扩散模型的遗忘
Yian Wang, Ali Ebrahimpour-Boroojeny, Hari Sundaram, Varun Chandrasekaran
cs.AI
diffusion
扩散模型相关
Abstract
Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75\%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the $κ$-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.
Chinese Translation
文本到图像扩散模型中的概念遗忘旨在抑制一个目标概念(例如 \texttt{horse}),同时保留相关但不同的内容(例如 \texttt{donkey}),然而现有方法要么在间接提示下发生泄漏,要么明显损害其他概念。我们表明,这些失败模式源于概念表示的几何结构,而非源于任何特定的算法。我们将概念形式化为激活空间中的区域,并证明目标概念与其他概念之间的重叠为任何稳健的擦除必然对它们造成的损害给出了下界,且这种权衡随重叠程度线性缩放。在十三种遗忘方法(包括为保留非目标概念而设计的方法)中,没有任何一种方法能同时实现强擦除与强邻居保留:STEREO 几乎消除了间接泄漏,却使邻居生成减少超过 75\%,而稀疏的推理时方法保留了邻居,却发生泄漏。损害随我们的重叠度量而增加,对 STEREO 而言是单调增加;$κ$-缩放规律在 SDXL 上重现,而邻居选择性损害在 FLUX 上再次出现。对于纠缠的概念而言,完美的遗忘是错误的追求目标;方法应在我们定理所确立的帕累托前沿上进行评估。
cs.AI / 9 / 2609.34259
QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models
QuantaSpike:面向大语言模型的短窗口脉冲驱动量化
Bang Hu, Guowei Zhu, Changze Lv, Xiaoqing Zheng, Fengzhe Zhang, Fan Zhang, Wei Cao
cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
Chinese Translation
大语言模型(LLMs)在许多任务上取得了强大性能,但在推理过程中依赖稠密的乘累加(MAC)运算,导致高能耗。脉冲神经网络(SNNs)提供了一种事件驱动的替代方案,其中突触整合使用轻量级累加。然而,脉冲驱动的 LLM 推理仍然困难,因为包含大量离群值的激活通常需要长发放窗口或辅助非脉冲路径。我们提出 QuantaSpike,一个围绕对数三值积分-发放(LTIF)神经元构建的、面向 LLMs 的短窗口脉冲驱动量化框架。LTIF 使用带有 2 的幂膜响应量子的三值事件,提高了每个发放步骤所表示的信息,同时保留了与 shift-ACC 兼容的计算。QuantaSpike 将该神经元与组自适应增益和选择性离群值准入相结合:正常值使用残差 LTIF 步骤,而被准入的离群值在进入相同的残差动力学之前接收一个额外的起始脉冲。在 OPT 和 Llama-2 上,QuantaSpike 在脉冲驱动的 LLM 量化方法中实现了最先进或有竞争力的困惑度和零样本准确率。它也能迁移到更新的稠密 LLM 上,在相同的四步发放窗口下,在 Llama-3-8B 和 Qwen3-8B 上保持接近 FP16 参考。分析性线性能耗预测表明,相对于 SpikeQuant,QuantaSpike 在 OPT 模型上将一次线性变换的能耗降低约 $80.0\%$,在 Llama-2 模型上降低 $67.1\%$,为 LLM 推理提供了一条准确且高能效的脉冲驱动路径。
cs.AI / 10 / 2609.34359
Improving Large Language Models for Code through Runtime Program-State Reasoning
通过运行时程序状态推理提升大型语言模型的代码能力
Hongwei Li, Spandan Garg, Yufan Huang
cs.AI
large language model
大语言模型相关
Abstract
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
Chinese Translation
大型语言模型在推理运行时程序状态方面接受的显式训练有限。我们研究训练模型推理运行时程序状态是否能够提升下游软件工程能力。我们引入两个互补的程序状态推理任务。有缺陷的输入-输出推理要求模型生成一个具体输入,该输入能暴露有缺陷程序与隐藏的正确实现之间的行为差异,并预测由此产生的执行行为。前置条件-后置条件推理要求智能体以符号方式刻画触发缺陷的前置条件,预测预期的后置条件,解释它们之间的因果联系,并将这种推理实例化为可执行的回归测试。通过将这两个任务纳入分阶段后训练流程,我们开发了 Comet-9B,一个基于 Qwen3.5-9B Base 的 9B 语言模型。我们在仓库级补丁生成、回归测试生成和安全 PoC 生成上评估所得检查点。将两个程序状态推理任务加入针对问题解决的监督微调(SFT)后,在 SWE-bench Pro 上成功率提升了 7.25 个百分点,在 SWT-Bench Verified 上提升了 9.70 个百分点。对这两个任务进行顺序强化学习,分别在 SWE-bench Pro、SWT-Bench Verified 和 CyberGym 上进一步带来 7.25、26.79 和 4.67 个百分点的增益。尽管仅有 9B 参数,Comet-9B 在 SWE-bench Pro 上取得了与报告的 GPT-5.2 结果相当的分数,并在 SWT-Bench Verified 上匹配了报告的基于 GPT-4o 的智能体的成功率。
cs.AI / 11 / 2609.34397
SkillFocus: Evolving Agent Skills via Capability Decomposition
SkillFocus:通过能力分解演化智能体技能
Ning Wang, Zhiren Gong, Bingdong Li, Peng Yang, Aimin Zhou
cs.AI
large language model
大语言模型相关
Abstract
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24\% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.
Chinese Translation
智能体技能演化旨在通过迭代修订,改进面向大语言模型(LLM)智能体的可复用过程性指导。现有方法主要基于执行轨迹或反馈来进行每次修订,这使得跨任务反复出现的行为需求保持隐含,并将修订与当前技能的行为绑定在一起。我们提出 SkillFocus,它将反复出现的任务需求分解为一个在技能演化过程中保持固定的能力空间,从而将任务需要什么与当前技能如何表现分离开来。SkillFocus 将当前任务结果映射到该空间,以识别出导致最多任务未能解决的能力,然后利用该能力来确定需要修订什么以及使用哪些证据。在涵盖异构任务的四个基准上,SkillFocus 在所有四个基准上都取得了最佳的留出准确率,平均比最强的竞争结果高出 5.7 个点,同时平均比最接近的迭代基线少用 24\% 的演化 token。对照研究进一步表明,由反复出现的任务需求推导出的能力优于任务语义和由执行推导出的替代方案,而将任务--能力分配随机化会使最终准确率下降最多 20.2 个点。在优先修订下,将证据与所选能力相匹配可使候选增益提高 4.4 个点。
cs.AI / 12 / 2609.34449
CORTEX: Learning to Share and Specialize in Dense Language Models
CORTEX:在稠密语言模型中学习共享与专精
Chuiyang Meng, Ming Tang, Vincent W. S. Wong
cs.AI
large language model
大语言模型相关
Abstract
Large language models are trained on heterogeneous data mixtures, where different knowledge domains require both shared knowledge and specialization. Existing modular approaches typically impose explicit components or discover modules through interpretability analysis after training. In this work, we propose CORTEX, a learning dynamics-inspired framework that learns internal modularization within dense language models. CORTEX partitions trainable matrices into parameter groups and learns module assignments from domain-conditioned gradient and cross-domain gradient similarity. We introduce the selective lesion score and module-domain mutual information to characterize the target-domain lesion effects and alignment, and analyze how module assignment affects the trade-off between assignment bias and update magnitude. Experiments with 160M, Qwen3-8B, and Qwen3-32B backbone models show that CORTEX achieves the highest synthetic-domain exact match and largest average perplexity reduction, while remaining competitive on real-domain evaluations and forming identifiable modules.
Chinese Translation
大型语言模型在异构数据混合上进行训练,其中不同的知识领域既需要共享知识,也需要专精化。现有的模块化方法通常强加显式组件,或在训练后通过可解释性分析来发现模块。在这项工作中,我们提出 CORTEX,一个受学习动力学启发的框架,它在稠密语言模型内部学习内部模块化。CORTEX 将可训练矩阵划分为参数组,并从领域条件梯度与跨领域梯度相似度中学习模块分配。我们引入选择性损伤评分(selective lesion score)与模块-领域互信息,以刻画目标领域的损伤效应与对齐程度,并分析模块分配如何影响分配偏差与更新幅度之间的权衡。在 160M、Qwen3-8B 和 Qwen3-32B 骨干模型上的实验表明,CORTEX 取得了最高的合成领域精确匹配和最大的平均困惑度降低,同时在真实领域评估上保持竞争力,并形成可识别的模块。
cs.AI / 13 / 2609.34492
PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
PowerBench:电力系统中智能体式检索与推理的基准
Xijing Wang, Yinsheng Yao, Jinru Ding, Yidong Jiang, Ziwen Xu, Yiwen Jiang, Jie Xu, Dawei Cheng
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.
Chinese Translation
大型语言模型(LLM)智能体为工业中的自动化分析提供了新的机遇。然而,对此类智能体——例如在电力系统场景中——的严格评估仍然受到阻碍:真实运行数据是保密的,而现有公共资源未能充分捕捉链式依赖关系与异构证据。为填补这一空白,我们提出 PowerBench,其包括(1)一个生成框架,该框架通过共同的依赖链推导出相互关联的异构运行数据,以及(2)一个由该框架生成的合成数据集。该数据集覆盖 100 种设备类型下的 761 台设备,包含跨越两年的 1335 万条逐小时遥测记录以及 24,939 份运行文档。在此数据集基础上,我们构建了跨三个任务族的 300 个问题,用于评估前沿 LLM 在受限的工具调用次数和时间预算下,完成需要在相互关联且异构的数据中进行自主证据检索与推理的分析任务的能力。结果表明,所评估的前沿 LLM 在这些任务上仍面临挑战:最佳模型仅达到 74.2% 的联合准确率。我们的轨迹分析进一步揭示,模型性能在证据发现、内容检索、工具使用、证据推理和答案提交等环节存在差异。这些发现为评估 LLM 智能体并指导其在工业中的可靠部署提供了细致见解。该框架、数据集和基准任务可在 https://github.com/open-compass/PowerBench 获取。
cs.AI / 14 / 2609.34510
Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
AI 能在加密货币中赚钱吗?衡量从回测到真实市场的差距
Xingtong Yu, Jiarun Zhou, Guanlin Ding, Wenkang Wei, Jiarui Liu, Chang Zhou, Fangzhou Ge, Chenyi Xu, Xikun Zhang, Renqiang Luo, Jie Zhang, Hong Cheng, Xinming Zhang, Hui Zhang, Yuan Fang
cs.AI
large language model
大语言模型相关
Abstract
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at https://github.com/Starlien95/Awesome-TradingAI.
Chinese Translation
基于 AI 的交易方法已迅速从机器学习和强化学习演进到大型语言模型(LLMs)和交易智能体,然而它们的表现仍然主要通过历史回测来评估。此类评估只能提供有限证据,说明一种方法能否泛化到未见过的未来市场,或其回测表现能否在现实交易摩擦(例如延迟、滑点、流动性约束和市场冲击)下持续。我们提出了一个统一基准,通过三个逐步更加真实的阶段来评估加密货币市场中具有代表性的机器学习、强化学习、基于 LLM 以及基于智能体的交易方法:历史回测、前瞻性的基于交易所的模拟交易和真实资金实盘交易。这些阶段共同提高了时间真实性,即从历史市场转向未见过的未来市场,并通过从离线模拟转向实盘交易提高了执行真实性。该协议使我们能够量化回测到实际的差距,识别性能何时开始恶化,并比较这一差距在主要 AI 交易方法类别之间有何不同。我们进一步提供一个统一的开源系统,支持所有三个评估阶段,并提供一个持续更新基准结果的公共平台。代码可在 https://github.com/Starlien95/Awesome-TradingAI 获取。
cs.AI / 15 / 2609.34537
The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions
科学推理的马拉松:科学智能体在多轮交互中对扰动的鲁棒性
Xiaoting Lyu, Xinbo Ma, Yufei Han, Hangwei Qian, Ziyang Lin, Bin Wang, Bin Wang, Wei Wang
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations), a benchmark for evaluating scientific agents under scientifically plausible perturbations throughout multi-turn problem solving. \textsc{SciARP} transforms 620 scientific problems into interdependent tasks of 3--13 turns and defines 13 perturbation types spanning problem understanding, evidence processing, reasoning, and conclusion formation. Clean and perturbed versions of each task are independently executed under matched settings, producing paired live trajectories for evaluating both task success and process reliability. Experiments across eight LLMs from four model families reveal three key robustness characteristics. First, different classes of scientific perturbations exhibit distinct robustness profiles and can decouple task progression from scientific reliability: agents may continue advancing through the task even after their information or reasoning has become unreliable. Second, stronger clean-task performance does not necessarily translate into stronger robustness, as models with higher clean-task accuracy can exhibit larger degradation under perturbation. Third, perturbation effects exhibit strong temporal dynamics: they may remain latent for multiple turns before emerging and subsequently propagate through downstream dependencies. Together, these findings show that current scientific agents remain insufficiently robust to scientifically plausible perturbations, with failures often remaining undetected, propagating, and resisting recovery.
Chinese Translation
基于大语言模型(LLM)的科学智能体正越来越多地被用于科学问题求解,然而它们对多轮交互过程中出现的不完美之处的鲁棒性仍然未得到充分理解。我们提出了 \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations),这是一个基准,用于在整个多轮问题求解过程中在科学上合理的扰动下评估科学智能体。\textsc{SciARP} 将 620 个科学问题转化为 3--13 轮的相互依赖任务,并定义了 13 种扰动类型,涵盖问题理解、证据处理、推理和结论形成。每个任务的干净版本和扰动版本在匹配设置下独立执行,产生成对的实时轨迹,用于评估任务成功和过程可靠性。来自四个模型家族的八个 LLM 的实验揭示了三个关键鲁棒性特征。首先,不同类别的科学扰动表现出不同的鲁棒性特征,并能够将任务进展与科学可靠性解耦:即使智能体的信息或推理已变得不可靠,它们仍可能继续推进任务。其次,更强的干净任务性能并不一定转化为更强的鲁棒性,因为具有更高干净任务准确率的模型在扰动下可能表现出更大的性能退化。第三,扰动效应表现出强烈的时间动态:它们可能在多个轮次中保持潜伏,之后才显现,并随后通过下游依赖关系传播。总之,这些发现表明,当前的科学智能体对科学上合理的扰动仍然不够鲁棒,失败往往未被检测到、会传播,并且难以恢复。
cs.AI / 16 / 2609.34540
APOLO: Automatic Prompt Optimization for Ontology Learning
APOLO:面向本体学习的自动提示优化
Huu Tan Mai, Roman Kochnev, Cuong Xuan Chu, Lukas Lange, Heiko Paulheim, Daria Stepanova
cs.AI
large language model
大语言模型相关
Abstract
Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimization for Ontology Learning, by casting OL as an explicit prompt optimization problem over LLM modules. To obtain training data, we employ a multi-agent system that generates text-ontology pairs from existing expert-curated ontologies. We then propose two ontology learner architectures: a greedy and an autoregressive learner, and optimize both using GEPA, a greedy evolutionary prompt optimizer built on DSPy. Experiments on two ontologies - a biomedical (DOID) and a plant ontology (PO) show consistent improvements after optimization across nearly all model and mode combinations, with autoregressive learners achieving the largest gains. Our results demonstrate that prompt optimization is a viable and lightweight alternative to fine-tuning for OL, and that the autoregressive formulation better captures ontological structure than the greedy approach.
Chinese Translation
随着大型语言模型(LLMs)的出现,从文本中进行的本体学习(OL)已取得进展,但由于标注训练数据的可用性有限以及使 LLMs 适应以有效执行 OL 的困难,它仍然具有挑战性。我们通过 APOLO——面向本体学习的自动提示优化(Automatic Prompt Optimization for Ontology Learning)来解决这一问题,通过将 OL 表述为在 LLM 模块上的显式提示优化问题。为了获得训练数据,我们采用一个多智能体系统,该系统从现有的专家整理的本体中生成文本-本体对。然后,我们提出两种本体学习器架构:一个贪心学习器和一个自回归学习器,并使用 GEPA(一种构建在 DSPy 之上的贪心进化式提示优化器)对二者进行优化。在两个本体——一个生物医学本体(DOID)和一个植物本体(PO)——上的实验表明,优化后在几乎所有模型和模式组合中都取得了一致的改进,其中自回归学习器取得了最大的增益。我们的结果表明,对于 OL 而言,提示优化是微调的一种可行且轻量的替代方案,并且自回归表述比贪心方法更好地捕获了本体结构。
cs.AI / 17 / 2609.34548
SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning
SGG-ReflAct:子目标引导的 ReflAct 与结构化规划,用于可靠的长时程推理
Jaeho Jung, Sung Hoon Jung
cs.AI
large language model
大语言模型相关
Abstract
Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLACT, which reflects only on the end-goal at each step without explicitly considering intermediate sub goals. To address this problem, we propose SGG-ReflAct (Sub-Goal Guided Re flAct), a reasoning backbone that integrates sub-goals generated through a single path LLM planner into the reflection process. We further extend this framework to BeamSGG-ReflAct, which replaces the single-path planner with a beam search based LLM planner for structured plan exploration. We run experiments on ALF World, ScienceWorld, and Jericho with multiple LLM models. SGG-ReflAct out performs REFLACT in nearly all settings, achieving best success rate gains of 14.9 percentage points on ALFWorld and 8.0 percentage points on ScienceWorld with Llama-3.1-8B-Instruct. Our experimental analysis shows that SGG-ReflAct re duces hallucinated actions and achieves its largest gains on procedurally ordered tasks. Furthermore, experimental results with BeamSGG-ReflAct show that the backbone's effectiveness depends on plan quality: explicitly specifying the re quired operations recovers gains that plan searching alone cannot achieve. These results demonstrate that SGG-ReflAct offers a practical and highly effective rea soning backbone, enabling LLM agents to achieve reliable performance in com plex, long-horizon tasks through easy integration.
Chinese Translation
近期,推理骨干的进展使大语言模型(LLM)智能体能够处理复杂的多步任务。然而,随着推理时程的增长,不一致的内部信念会诱发中间错误,导致智能体偏离其目标。这一局限在 REFLACT 中同样存在,它在每一步仅对最终目标进行反思,而没有显式考虑中间子目标。为解决这一问题,我们提出 SGG-ReflAct(子目标引导的 ReflAct),这是一种推理骨干,它将通过单路径 LLM 规划器生成的子目标整合到反思过程中。我们进一步将该框架扩展为 BeamSGG-ReflAct,它用基于束搜索的 LLM 规划器替换单路径规划器,以进行结构化规划探索。我们在 ALFWorld、ScienceWorld 和 Jericho 上使用多个 LLM 模型进行了实验。SGG-ReflAct 在几乎所有设置中都优于 REFLACT,在使用 Llama-3.1-8B-Instruct 时,在 ALFWorld 上取得了 14.9 个百分点的最高成功率提升,在 ScienceWorld 上取得了 8.0 个百分点的最高成功率提升。我们的实验分析表明,SGG-ReflAct 减少了幻觉动作,并在程序化顺序任务上取得了最大的增益。此外,BeamSGG-ReflAct 的实验结果表明,该骨干的有效性取决于规划质量:显式指定所需操作可以恢复仅靠规划搜索无法实现的增益。这些结果表明,SGG-ReflAct 提供了一种实用且高效的推理骨干,使 LLM 智能体能够通过轻松集成,在复杂的长时程任务中实现可靠的性能。
cs.AI / 18 / 2609.34571
PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations
PersonaManifold:揭示并利用LLM人格表征中的弯曲几何
Rui Xu, Yinghui Xu, Libo Wu
cs.AI
large language model
大语言模型相关
Abstract
Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations---addition, scaling, and linear interpolation---under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold's intrinsic geometry---local metric tensors, geodesic distances, and Ollivier-Ricci curvature---and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.
Chinese Translation
在推理时控制大语言模型(LLM)中的人格对于角色扮演、个性化对话和社会模拟很重要。近期方法从模型激活空间中提取人格向量,并在线性表示假设下施加欧几里得操作——加法、缩放和线性插值。然而,这些方法本身报告了系统性失败:非正交的特质维度、不对称的天花板效应和阻力效应,以及多特质组合中的显著偏差,这表明线性各向同性假设并不成立。我们提出PersonaManifold,一个将人格表征建模为激活空间中弯曲、低维黎曼子流形上的点的框架。我们估计该流形的内在几何——局部度量张量、测地距离和Ollivier-Ricci曲率——并引入测地引导,它沿流形测地线而非欧几里得直线在人格之间插值。我们还提出行为相似性三元组(BST)基准,它自动生成基于六个既定心理学构念的情境问题,并通过行为反应而非自我报告问卷来定义人格相似性。在三个开源LLM上的实验表明,人格激活形成一个具有异质曲率的流形,测地距离比欧几里得替代方案更准确地预测行为相似性,其中各向异性和曲率具有独立贡献,并且测地引导在我们的BST基准和外部评估上都能产生更连贯的中间人格,其优势集中在流形偏离平坦程度最大的高偏差区域。
cs.AI / 19 / 2609.34575
Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
用于长时程离线目标条件强化学习的扩散子目标规划
Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang
cs.AI
diffusion
扩散模型相关
Abstract
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
Chinese Translation
离线目标条件强化学习(GCRL)从无奖励数据中学习目标导向的策略,但在长时程任务中,由于奖励稀疏和折扣,目标条件价值函数往往提供不稳定的指导。分层方法通过子目标分解部分缓解了这一问题;然而,高层决策仍然依赖对噪声敏感的价值估计,导致在复杂环境中行为不稳定。我们通过提出\textbf{扩}散\textbf{子}目标\textbf{规}划(\textbf{DSP})来解决这一局限,这是一个用于高层子目标生成的基于扩散的框架。DSP 将高层规划表述为对目标条件子目标的引导式生成推断,并同时学习条件流和无条件流,从而能够在推理时借助无分类器引导引入目标导向的偏置。通过从高层规划中移除显式的基于价值的引导,DSP 借助生成模型生成可达且目标导向的子目标,同时保留分层执行。在离线 GCRL 基准上的实验表明,DSP 在一系列导航和操作任务上优于先前的方法,在需要多步子目标规划的迷宫环境中表现尤为突出。
cs.AI / 20 / 2609.34619
LLMs for Executable Multi-Agent System Specification Generation
用于可执行多智能体系统规范生成的大语言模型
Andreas Kouvaras, Periklis Mantenoglou, Alexander Artikis
cs.AI
large language model
大语言模型相关
Abstract
MAS specifications express the effects of the actions of the agents and their environment, as well as other temporal phenomena, such as the intervals during which an agent may perform an action. The specification of a MAS should also be executable in order to allow for run-time monitoring. Constructing the specification of a MAS requires formal language expertise, while machine learning techniques depend on labelled data which are rarely available. To address these issues, we propose `genRTEC', a method that leverages pre-trained Large Language Models (LLMs) to generate executable MAS specifications, in the language of the `Run-Time Event Calculus' (RTEC), from natural language descriptions. genRTEC constructs MAS specifications with complex hierarchical and cyclic dependencies based only on short natural language descriptions of the concepts involved. We present an extensive empirical evaluation of genRTEC, spanning various MAS specifications, including both a qualitative and a quantitative assessment. Our results demonstrate that genRTEC constructs executable MAS specifications of high predictive accuracy without compromising reasoning efficiency.
Chinese Translation
MAS 规范表达智能体及其环境动作的效果,以及其他时间现象,例如智能体可以执行动作的区间。MAS 的规范还应当是可执行的,以便允许运行时监控。构建 MAS 的规范需要形式语言专业知识,而机器学习技术依赖于很少可获得的标注数据。为了解决这些问题,我们提出了 `genRTEC`,一种利用预训练大语言模型(LLMs)从自然语言描述中生成以 `Run-Time Event Calculus`(RTEC)语言表示的可执行 MAS 规范的方法。genRTEC 仅基于对所涉及概念的简短自然语言描述,构建具有复杂层次和循环依赖的 MAS 规范。我们展示了对 genRTEC 的广泛实证评估,涵盖各种 MAS 规范,包括定性评估和定量评估。我们的结果表明,genRTEC 构建了具有高预测准确率的可执行 MAS 规范,同时不损害推理效率。
cs.AI / 21 / 2609.34636
MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics
MechReasoner:一个用于定性物理学中机制推理的模拟器与基准
Danilo Gusicuma, André Freitas
cs.AI
large language model
大语言模型相关
Abstract
This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying structural and causal constraints. The benchmark tests whether answers preserve simulator-licensed ambiguity, quantified claims, episode-graph transition evidence, repairs, and trace-support judgments. Its 1,120 items are generated deterministically from admissible interpretation sets, component states, scenario restrictions, confluence constraints, and derivation steps across 18 catalog mechanisms and six task families. Each mechanism undergoes converter checks of structure and topology and behavioral checks against quantitative simulations. GPT-5.5 accuracy decreases as family-specific mechanistic complexity increases, from 76.1% in the lowest-complexity bucket (B1) to 38.0% in the highest-complexity bucket (B4). The negative association remains after controls for rendered-prompt and expected-answer length. These results show that qualitative simulators can support auditable NLP benchmarks for mechanistic inference.
Chinese Translation
本文介绍了 MechReasoner,一个以基于汇流的定性物理学为基础构建的机制性定性模拟器,以及一个用于机制性推理的基准。当前的大语言模型(LLM)生成的机制性描述虽然流畅,却并不能可靠地从底层的结构与因果约束中推出。该基准测试答案是否保留了模拟器所许可的歧义、量化断言、事件图转移证据、修复以及轨迹支持判断。其 1,120 个题目由可接受的解释集、组件状态、场景限制、汇流约束和推导步骤确定性地生成,覆盖 18 种目录机制和六个任务族。每个机制都要接受针对结构与拓扑的转换器检查,以及针对定量模拟的行为检查。随着任务族特定机制复杂度的增加,GPT-5.5 的准确率下降,从最低复杂度分桶(B1)的 76.1% 降至最高复杂度分桶(B4)的 38.0%。在控制渲染提示与预期答案长度之后,这种负向关联依然存在。这些结果表明,定性模拟器能够支撑用于机制性推理的可审计 NLP 基准。
cs.AI / 22 / 2609.34686
Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents
越狱上下文挥之不去:工具智能体中分歧的安全路由及其跨任务可预测性
Xi Wang, Songlei Jian, Yiming Zhang, Bin Ji, Zhaoye Li, Ma Jun, Baosheng Wang, Jie Yu
cs.AI
large language model
大语言模型相关
Abstract
As large language models increasingly operate as tool-using agents, post-jailbreak safety feedback is often assumed to serve as a reliable safeguard; however, how lingering jailbreak context shapes subsequent agent behavior remains largely unexplored. To systematically examine this dynamic, we introduce a paired continuation framework across 192 parent tasks spanning 42 domains, evaluating 12,148 analyzed continuation pairs (curated from a 12,288-pair initially design) across eight diverse agents. We find that identical safety feedback induces sharply model-dependent behavioral routing rather than uniform protection: redirecting unsafe trajectories toward legitimate completion (\emph{rescue}), sustaining unauthorized execution (\emph{persistent unsafe}), or triggering over-refusal on benign tasks (\emph{collateral loss}). Through layer-wise activation patching, we discover a shared \emph{late-commit pattern} where causal intervention effects surge sharply near the final layers (relative depths of 0.958--0.984) despite an over 30-fold variation in peak magnitude across architectures. Crucially, critical-layer representations correlate with macroscopic routing outcomes, and intervening at these layers causally alters concrete next-step tool actions. Building on this causal foundation, we test whether localized intervention-derived features can serve as predictive proxies for full-trajectory routing outcomes on unseen parent tasks under leave-one-parent-task-out evaluation, finding that they provide viable predictive signals in responsive agents with peak ROC AUCs reaching 0.675 for \emph{rescue}, 0.777 for \emph{collateral loss}, and 0.702 for \emph{persistent unsafe}. These findings establish a mechanistic lens and a predictive baseline for anticipating the safety and utility trade-offs of post-jailbreak feedback in autonomous agents.
Chinese Translation
随着大型语言模型越来越多地作为使用工具的智能体运行,越狱后的安全反馈通常被假定为一种可靠的保障措施;然而,挥之不去的越狱上下文如何塑造后续智能体行为,在很大程度上仍未得到探索。为了系统地考察这一动态,我们引入了一个配对延续框架,涵盖跨42个领域的192个父任务,在八个不同的智能体上评估了12,148个已分析的延续对(从最初设计的12,288对中筛选而来)。我们发现,相同的安全反馈会诱发截然不同、强烈依赖模型的行为路由,而非统一的保护:将不安全轨迹重新导向合法完成(\emph{救援})、维持未授权执行(\emph{持续不安全}),或在良性任务上触发过度拒绝(\emph{附带损失})。通过逐层激活修补,我们发现了一个共享的\emph{晚期承诺模式},其中因果干预效应在最后几层附近急剧激增(相对深度为0.958--0.984),尽管不同架构之间的峰值幅度变化超过30倍。至关重要的是,关键层表征与宏观路由结果相关,并且在这些层进行干预会因果性地改变具体的下一步工具动作。在这一因果基础上,我们测试了局部干预衍生特征能否在留一父任务评估下,作为未见父任务上全轨迹路由结果的预测代理,发现它们在响应性智能体中提供了可行的预测信号,峰值ROC AUC在\emph{救援}上达到0.675,在\emph{附带损失}上达到0.777,在\emph{持续不安全}上达到0.702。这些发现为预判自主智能体中越狱后反馈的安全性与效用权衡,建立了一个机制性视角和一个预测基线。
cs.AI / 23 / 2609.34712
RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents
RSI-Router:为成本高效智能体演化子任务级 LLM 路由与技能
Hao Li, Hangfan Zhang, Zhiyao Cui, Chunjiang Mu, Yiqun Zhang, Bo Zhang, Danyang Jia, Shuyue Hu
cs.AI
large language model
大语言模型相关
Abstract
Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance-cost trade-off can be improved through within-task large-small model collaboration, as smaller models can handle some stages even when they cannot solve the full task. In this paper, we introduce RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience. Each iteration consists of four stages: Subtask Mining derives subtask definitions and identification rules from training trajectories; Routing Strategy Evolution proposes and evaluates diverse model assignments; Model-Specific Skill Evolution compares routed and large-model-only trajectories to diagnose failures and develop reusable execution skills; and Pareto-Optimal Router Selection updates the Pareto population using historical and newly generated routers while retaining dominated routers as experience for subsequent evolution. Routing between DeepSeek-V4.1-Flash and Qwen3.5-9B, RSI-router consistently surpasses the DeepSeek-only baseline at roughly half the inference cost (48.3%) across five agentic benchmarks. In particular, on ALFWorld, ScienceWorld, and WebShop, it cuts inference cost by 74.7-82.2% while simultaneously improving performance; on Terminal-Bench 2.0, it achieves a 16.7% relative performance gain at 18.0% lower cost. Moreover, RSI-router establishes a stronger performance--cost Pareto frontier than 9 routing methods.
Chinese Translation
大型语言模型(LLM)智能体的实际部署要求在可负担的推理成本下具备强大的任务性能。对于长时程智能体任务,这种性能—成本权衡可以通过任务内的大模型—小模型协作得到改善,因为即使小模型无法解决完整任务,它们也能处理其中某些阶段。在本文中,我们提出 RSI-router,一个通过在累积经验上进行递归自我改进来构建子任务级模型分配与模型专用技能的路由框架。每一次迭代由四个阶段组成:子任务挖掘(Subtask Mining)从训练轨迹中推导出子任务定义与识别规则;路由策略演化(Routing Strategy Evolution)提出并评估多样化的模型分配方案;模型专用技能演化(Model-Specific Skill Evolution)比较经路由的轨迹与仅使用大模型的轨迹,以诊断失败并开发可复用的执行技能;而帕累托最优路由选择(Pareto-Optimal Router Selection)利用历史路由与新生成的路由更新帕累托种群,同时保留被支配的路由作为后续演化的经验。在 DeepSeek-V4.1-Flash 与 Qwen3.5-9B 之间进行路由,RSI-router 在五个智能体基准上以约一半的推理成本(48.3%)持续超越仅使用 DeepSeek 的基线。特别地,在 ALFWorld、ScienceWorld 和 WebShop 上,它将推理成本降低 74.7–82.2%,同时提升性能;在 Terminal-Bench 2.0 上,它在成本降低 18.0% 的情况下实现了 16.7% 的相对性能提升。此外,RSI-router 建立了比 9 种路由方法更强的性能—成本帕累托前沿。
cs.AI / 24 / 2609.34736
SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing
SeLMRoute:面向大型语言模型路由的概率语义证据
Vasilis Perifanis, Nikolaos Pavlidis, Symeon Symeonidis
cs.AI · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing rarely states what a query actually requires. We introduce SeLMRoute, a routing framework that separates the extraction of candidate-independent semantic evidence from the learning of candidate performance and the application of deployment objectives. A decision model first evaluates a set of interpretable questions about the query, such as its reasoning requirements and use of external knowledge, with each judgment retained as a probability distribution. The resulting probabilistic semantic state is used by a lightweight supervised router to estimate candidate model performance. Routing objectives are applied after performance estimation, which allows the same semantic state to support performance-oriented and cost-aware decisions. On the LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries), SeLMRoute achieves an average accuracy of $72.08\% \pm 0.45$, while grouped five-fold out-of-fold evaluation reaches $72.64\%$, compared with $69.23\%$ for the strongest fixed candidate. The representation achieves the highest mean performance among the evaluated semantic, dense, lexical, and domain-level representations. In a separate 13-model performance-cost setting, SeLMRoute improves performance in all five grouped splits, with a mean PerfGain of $2.66\%$. Our code is available at https://github.com/Indigma-Innovations/SeLMRoute.
Chinese Translation
大语言模型(LLM)路由旨在为每个到来的查询选择最合适的模型。大多数现有路由器直接从查询嵌入、模型表示、偏好数据或相似示例的聚类中学习这一决策。此类方法可能有效,但用于路由的表示很少说明查询实际需要什么。我们提出 SeLMRoute,一个路由框架,它将与候选模型无关的语义证据的提取,与候选模型性能的学习以及部署目标的应用分离开来。一个决策模型首先评估一组关于查询的可解释问题,例如其推理需求以及对外部知识的使用,并将每个判断保留为概率分布。由此得到的概率语义状态被一个轻量级监督路由器用于估计候选模型的性能。路由目标在性能估计之后应用,这使得相同的语义状态能够支持面向性能的决策和成本感知决策。在 LLMRouterBench(15 个数据集、20 个候选模型、11,481 个查询)上,SeLMRoute 达到了 $72.08\% \pm 0.45$ 的平均准确率,而分组五折折外评估达到 $72.64\%$,相比之下最强固定候选模型为 $69.23\%$。在被评估的语义、稠密、词汇和领域级表示中,该表示取得了最高的平均性能。在一个单独的 13 模型性能-成本设置中,SeLMRoute 在所有五个分组划分中都提升了性能,平均 PerfGain 为 $2.66\%$。我们的代码可在 https://github.com/Indigma-Innovations/SeLMRoute 获取。
cs.AI / 25 / 2609.34772
Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs
在 Token 承诺之前:扩散 VLM 中视觉幻觉的轨迹级基准测试
Yadong Wang, Siping Yue, Yu Tian, Chuanxing Geng, Xiang Chen
cs.AI
diffusion
扩散模型相关
Abstract
Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.
Chinese Translation
多模态扩散语言模型通过迭代地解掩码 token 来生成响应,使每个答案成为多步轨迹的终点,而不是一次即时的承诺。为自回归模型构建的幻觉基准测试只评估最终输出,因此无法判断扩散 VLM 中一个无依据的断言是出现得较晚,还是在任何答案 token 被揭示之前就已经稳定下来。我们提出 DynaHall,一个轨迹级基准,由带标注的二元视觉命题构成,涵盖物体存在性、计数、属性和关系,并配有按视觉先验分级的受控难负样本。DynaHall 搭配一个承诺感知协议,该协议在每一个解掩码步骤记录中间答案倾向,并与已承诺的输出一并记录。在来自三个架构家族的五个扩散 VLM 上,视觉幻觉在承诺之前就已确定:当答案位置仍被掩码时,一个无依据的答案就已经是首选状态,而后续的解掩码步骤很少将其逆转,因此该失败并非在写入步骤被引入。这一点在解码调度、答案格式和开放式生成中均成立。DynaHall 还揭示了被最终输出指标所掩盖的失败,包括计数与关系坍塌、先验驱动的假阳性,以及方向随类型而变的属性错误。在此诊断的指导下,PGS(Pre-commitment Gradient Steering,承诺前梯度引导)对仍处于掩码状态的答案状态进行编辑以减少假阳性,使肯定率接近平衡,并可迁移到另一种架构而不降低通用能力。DynaHall 和 PGS 表明,幻觉应沿着扩散 VLM 的生成轨迹进行度量与缓解,而不仅仅是在最终答案处。
cs.AI / 26 / 2609.34780
Applying Language Models in medical Medicine: Recent Trends and Perspectives
在医学中应用语言模型:近期趋势与展望
Erik Aerts
cs.AI
large language model
大语言模型相关
Abstract
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.
Chinese Translation
近年来,人工智能(AI)在医学研究和临床实践中的使用及适用性在文献中受到越来越多的关注。大语言模型(LLMs)的出现扩展了关于AI在医疗保健中应用的讨论。虽然传统的基于深度学习的医学AI应用通常聚焦于特定且明确的任务,但LLMs在处理可用数据时提供了更广泛的能力和灵活性。在撰写本文时,将LLMs整合到医疗环境中引发了关于其可靠性、准确性、透明度、安全性以及在医疗环境中适当角色的重要问题。本文呈现并讨论了近期关于LLMs在医学中应用的演讲和文章,特别强调其在研究和临床实践中的潜在效用。它既考虑了这些技术所提供的机会,也考虑了与其实现相关的挑战,旨在为LLMs在医学领域当前和新兴的角色提供一个视角。
cs.AI / 27 / 2609.34828
Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models
模拟受访者,而非单个问题:基于大语言模型的连贯调查生成
Ji Huang, Mengfei Li, Shuai Shao
cs.AI
large language model
大语言模型相关
Abstract
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
Chinese Translation
大语言模型正越来越多地被用于模拟社会调查中的回答分布。先前的研究已经在单个问题上实现了准确的群体层面模拟。然而,真实问卷会向每位受访者询问一系列相关问题。一个模拟受访者应在整个问卷中表现出连贯的偏好,而不仅仅是对孤立题项具有准确的分布。现有的单题方法无法准确复现同一个人如何回答一份完整的调查。我们提出 FullRespondent-LLM (FR-LLM),它微调两个专用大语言模型:一个用于每个题项回答分布的边缘模型,以及一个用于答案之间依赖关系的受访者级自回归模型。然后,边缘约束联合投影 (MCJP) 将自回归联合分布投影到满足第一个模型学到的题项级边缘分布的集合上。这能生成具有真实跨题项关系的完整问卷,同时保持很强的题项级准确性。在两个真实世界社会调查数据集上,FR-LLM 更准确地复现多问题回答模式,保持有竞争力的单题准确性,并且更好地泛化到未见过的群体和问题。在一个小型商业调查数据集中,我们使用模拟回答来制定定价和备货决策;FR-LLM 实现了最高的已实现利润。
cs.AI / 28 / 2609.34832
BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
BV 损失:面向块扩散投机解码的块验证感知损失
Suyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang, Seunghyun Lee, Hyunjae Oh, Baeseong Park, Dongsoo Lee
cs.AI · cs.CL
diffusion
扩散模型相关
Abstract
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0\% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.
Chinese Translation
扩散草稿模型通过并行提出多个 token 来加速投机解码。尽管近期通过序列级起草与验证在投机解码方面取得了进展,但现有的训练目标在很大程度上仍围绕 token 级验证来设计。为解决这一不匹配问题,我们提出块验证感知损失(Block Verification-aware loss,BV 损失),这是一种旨在最大化草稿序列期望接受长度的训练目标。BV 损失直接从块验证的接受规则推导而来,在序列层面上为草稿模型训练目标与推理时验证机制之间提供了有原则的关联。在数学、代码和聊天基准上,对于使用 Qwen3-4B 和 Qwen3-8B 的 DFlash 与 DSpark,在不改变推理流程的情况下,BV 损失相较于交叉熵损失训练,将块验证下每次验证调用所接受 token 的平均数量提高了 13.0--21.0\%。BV 损失也优于 TV 损失和 LK 损失等逐 token 接受目标,其增益还可延伸至 token 验证和贪心解码。这些结果表明,使用与序列级验证相对齐的目标来训练块扩散草稿模型,而非独立优化每个 token,具有明显优势。
cs.AI / 29 / 2609.34879
One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair
一次读出,多次修复:面向工具智能体修复的扩散引导分层搜索
Xiang Xia, Cheng Yan, Fan Xu, Zhijun Fan, Shuyuan Zhang, Wuyang Zhang
cs.AI · cs.CL
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both operation choices and their concrete realizations, making complete-sequence regeneration costly. Moreover, regeneration repeats operation selection even when failure arises from how those operations are realized. The resulting challenge is to reduce this repetition while preserving exploration of alternative operations and realizations. Therefore, we formulate repair as hierarchical search over operation supports, which we introduce as sets of permitted operation types that define reusable search regions for concrete tool-call sequences. We propose ReCommit, a training-free, diffusion-guided framework for improving tool-agent failure recovery while reducing repair computation. ReCommit amortizes operation-level proposal computation across repair trials by reusing operation-type scores from a single parallel readout of a masked diffusion language model. These scores guide search across supports, while realization search explores alternative entity bindings, arguments, and action composition within each support. Experiments on real failures across four enterprise services in the Agent-Diff benchmark show 75.9\% and 63.2\% relative recovery gains with 61.3\% and 51.3\% reductions in mean full-budget repair time at repair budgets $B=3$ and $B=13$, respectively, over the strongest evaluated 8B comparison method. ReCommit achieves a favorable recovery--cost trade-off, including in comparisons with the evaluated 32B models.
Chinese Translation
工具智能体使用大语言模型通过外部工具行动,然而成功执行的调用仍可能无法满足用户请求。工具智能体修复寻求能够成功执行并满足原始请求的替代调用序列。然而,修复需要同时探索操作选择及其具体实现,这使得完整序列重新生成代价高昂。此外,即便失败源于这些操作如何被实现,重新生成仍会重复操作选择。由此产生的挑战是减少这种重复,同时保留对替代操作和实现的探索。因此,我们将修复表述为对操作支持集的层次化搜索;操作支持集是我们引入的允许的操作类型集合,这些集合为具体工具调用序列定义了可复用的搜索区域。我们提出 ReCommit,一个无需训练的、扩散引导的框架,用于改进工具智能体故障恢复,同时减少修复计算。ReCommit 通过复用来自掩码扩散语言模型单次并行读出的操作类型分数,将操作级提议计算摊销到多次修复试验中。这些分数引导跨操作支持集的搜索,而实现搜索则在每个操作支持集内探索替代的实体绑定、参数和动作组合。在 Agent-Diff 基准中四个企业服务的真实故障上进行的实验表明,在修复预算 $B=3$ 和 $B=13$ 下,相较于所评估的最强 8B 对比方法,分别取得了 75.9\% 和 63.2\% 的相对恢复提升,以及 61.3\% 和 51.3\% 的平均全预算修复时间减少。ReCommit 实现了有利的恢复--成本权衡,包括在与所评估的 32B 模型比较时也是如此。
cs.AI / 30 / 2609.34930
PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents
PDEU-Bench:对工具调用LLM代理的个性化规划生命周期进行基准测试
Huayi Lai, Shichao Song, Qingchen Yu, Simin Niu, Mengwei Wang, Hanyu Wang, Xun Liang
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive execution, leaving unclear whether agents can formulate, execute, and revise an explicit plan while preserving user preferences throughout long-term interaction. To address this gap, we introduce \textbf{PDEU-Bench} (\textbf{P}ersonalized plan \textbf{D}efinition, plan \textbf{E}xecution, and plan \textbf{U}pdate \textbf{Bench}mark), a benchmark for evaluating the complete planning lifecycle of personalized tool-using agents. PDEU-Bench comprises 214 long-horizon interaction tasks spanning 12 everyday domains and 94 tools, with stage-specific assessments of preference adherence and plan quality. Extensive evaluations of 15 representative open-source and closed-source LLMs reveal a pronounced gap between local tool execution and dynamic planning: LLMs can often instantiate preferences in individual calls, yet struggle to construct coherent plan definition and plan update. We further evaluate mainstream personalization and memory-augmentation methods. Although these methods improve particular stages, none of the evaluated methods reliably propagates user preferences throughout the complete lifecycle, and their gains frequently fail to transfer to subsequent execution. Fine-grained error analysis further reveals that preference omissions and conflicts persist throughout the planning lifecycle, highlighting the need for future research to parameterize LLMs with preference-aware information retrieval and memory capabilities. We provide the relevant code and data in the appendix to support future research.
Chinese Translation
大型语言模型(LLM)代理正在从执行孤立指令的工具调用系统演变为通过持续的多步交互追求用户目标的面向任务的代理。然而,现有的个性化工具使用基准主要评估孤立的调用或反应式执行,使得代理能否在长期交互中在保留用户偏好的同时制定、执行和修订显式计划仍不清楚。为弥补这一空白,我们引入了 \textbf{PDEU-Bench}(\textbf{P}ersonalized plan \textbf{D}efinition, plan \textbf{E}xecution, and plan \textbf{U}pdate \textbf{Bench}mark),这是一个用于评估个性化工具使用代理完整规划生命周期的基准。PDEU-Bench 包含 214 个长时程交互任务,跨越 12 个日常领域和 94 个工具,并针对偏好遵循和计划质量进行阶段特定的评估。对 15 个具有代表性的开源和闭源 LLM 的广泛评估揭示了局部工具执行与动态规划之间的显著差距:LLM 通常可以在单个调用中实例化偏好,但难以构建连贯的计划定义和计划更新。我们进一步评估了主流的个性化和记忆增强方法。尽管这些方法改进了特定阶段,但所评估的方法均无法在整个完整生命周期中可靠地传播用户偏好,并且它们的增益往往无法迁移到后续执行。细粒度错误分析进一步揭示,偏好遗漏和冲突在整个规划生命周期中持续存在,凸显了未来研究需要以偏好感知的信息检索和记忆能力来参数化 LLM。我们在附录中提供了相关代码和数据以支持未来研究。
cs.AI / 31 / 2609.34971
Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias
面向LLM智能体的动作空间塑造:测量与缓解工具模式偏差
Yinhong Liu, Zhili Tan, Zilin Wang, Zhijiang Guo
cs.AI
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.
Chinese Translation
大型语言模型(LLM)在给定固定工具模式时,在工具使用智能体任务上表现出强大性能。然而,工具模式并不是智能体的动作空间;它仅仅是其一种接口表示。同一个可执行动作可以通过许多不同且功能等价的工具定义来暴露,而一个真正学会任务的智能体应在这些定义之间表现一致。我们表明,当前的智能体往往并非如此,我们将这一现象称为模式偏差。为了系统地研究这一点,我们引入了一个可执行变换框架,该框架使用九种算子重写原生工具模式,包括合并和拆分工具、改变单个工具的表达方式,以及将一个动作分布到多个依赖调用上。任务、可执行动作和可达状态保持固定,因此成功率的任何变化都仅可归因于接口。我们在多达32种模式变体上评估了十一个LLM(包括两个闭源模型),并追问模式偏差有多大、它如何表现、是否可以在不进行完整评估的情况下预测某个模式变体的难度,以及训练是否能消除它。我们发现,即使对于最新的模型,模式偏差也相当大:成功率仅取决于模式,范围从完全失败到97%。为了可靠地估计模式难度,需要运行一小组目标查询样本。只有当某个模式变体出现在训练数据中时,训练才能修复该变体。
cs.AI / 32 / 2609.34994
From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing
从一次性生成到增量式音乐作曲:将通用指令大语言模型适配用于持久化符号编辑
André Ricardo Ducca Fernandes, Jean-Pierre Briot, Simone Diniz Junqueira Barbosa1, Hélio Côrtes Vieira Lopes
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations -- chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.
Chinese Translation
大多数音乐生成系统仍然主要被界定和评估为完整输出的生产者,而作曲往往通过对同一音乐制品进行连续修订来推进。本文研究一种通用指令遵循大语言模型的不同用法:不是作为一次性音乐生成器,而是作为作用于不断演化的符号乐谱的可复用算子。我们将增量作曲形式化为在持久化 ABC 记谱上的一系列操作感知的状态转移,并明确规定每个操作可以改变什么、必须保留什么。该交互包括两种制品初始化变体以及三种编辑操作——和弦添加、修复补全(inpainting)与移调。我们通过在来自爱尔兰传统音乐的 496,038 条操作感知对话记录上,使用低秩适配(LoRA)对 Llama 3.1 8B Instruct 进行适配,来实现该形式化方案。与未适配模型的比较用于检验学习这一交互契约的可行性,而非为微调本身主张新颖性。在每个模型 500 段对话(1,750 个尝试的输出状态)中,检查器准入率从 29.37% 上升至 99.37%,而在准入条件下的合规率从 0.7205 上升至 0.9798。用于相对于参考的音乐特征分析的严格合格输出从 14 个增加到 1,548 个,并且在指定协议下,最长公共子序列分析相对于留出基线并未显示高重叠序列的系统性增加。这些结果支持使用通用指令大语言模型进行持久化、操作感知的符号编辑在技术上的可行性。它们并未确立更优的音乐质量或人机共创性,这些仍是以音乐家为中心的评估所要回答的问题。
cs.AI / 33 / 2609.35036
Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation
人格遵循并非选择性控制:LLM 用户模拟中的中立性缺口
Jiashen Ren, Wenlin Zhang, Bohan Zhang, Xiaopeng Li, Zichuan Fu, Wanyu Wang, Junyi Li, Xiangyu Zhao
cs.AI
large language model
大语言模型相关
Abstract
Persona prompting is widely used to construct user simulations with large language models (LLMs), yet it relies on a largely untested assumption: specifying one user attribute should change that attribute alone. We test this assumption and identify a systematic failure of selective control: across all eight black-box LLMs we audit, changing a target attribute also shifts responses on unspecified, non-target attributes. For example, describing a user as more risk-seeking shifts color choices, even though the prompt never mentions color; we term this cross-attribute influence. Semantic, contextual, and internal analyses collectively suggest that models treat a persona prompt as evidence about the user and extend the inferred profile to unspecified preferences, a process we call trait-conditioned completion. We next ask whether explicitly specifying non-target attributes restores selective control. When a non-target attribute is assigned a clear direction, models generally follow the declaration and suppress the target attribute's influence. However, when the same attribute is declared neutral, the target continues to affect choices across all five open-weight checkpoints, even when the model correctly reports the declared state. This disparity, the neutrality gap, demonstrates that successful persona following does not imply selective persona control, which additionally requires keeping non-target attributes stable. We operationalize this distinction with a three-state diagnostic that leaves the non-target attribute unspecified or declares it directional or neutral; because directional tests can be passed by simply following the stated persona, the neutral state reveals failures they miss. In a post hoc analysis of independent items, neutral declarations leave 51-81% of items target-sensitive, against at most 1 of 320 item-pole comparisons under directional ones.
Chinese Translation
人格提示被广泛用于构建基于大语言模型(LLM)的用户模拟,但它依赖于一个在很大程度上未经检验的假设:指定某一个用户属性应当只改变该属性本身。我们检验了这一假设,并发现了一种系统性的选择性控制失效:在我们审计的全部八个黑盒 LLM 中,改变目标属性也会使未指定的非目标属性上的回应发生偏移。例如,将用户描述为更偏好风险会使颜色选择发生偏移,尽管提示中从未提及颜色;我们将此称为跨属性影响。语义、语境与内部表征分析共同表明,模型将人格提示视为关于用户的证据,并将推断出的画像扩展到未指定的偏好上,我们将这一过程称为特质条件化补全。接下来我们追问,显式指定非目标属性是否能够恢复选择性控制。当非目标属性被赋予明确的方向时,模型通常遵循该声明,并抑制目标属性的影响。然而,当同一属性被声明为中立时,目标属性依然会在全部五个开放权重检查点上影响选择,即便模型正确地报告了所声明的状态。这种差异,即中立性缺口,表明成功的人格遵循并不意味着选择性人格控制,后者还额外地要求保持非目标属性稳定。我们用一个三态诊断来操作化这一区分,该诊断将非目标属性留作未指定,或将其声明为有方向的或中立的;由于仅仅遵循所陈述的人格就能通过方向性测试,中立状态揭示了它们所遗漏的失效。在对独立题项的事后分析中,中立声明使 51–81% 的题项对目标属性敏感,而在方向性声明下,320 个题项–极点比较中至多只有 1 个如此。
cs.AI / 34 / 2609.35077
What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages
什么驱动了生产环境大型语言模型中的引用?一项跨越一万个网页、针对两百万条AI引用的观察性多方法研究
Ben Moore, Liam Dunne
cs.AI
large language model
大语言模型相关
Abstract
Production large language models retrieve and cite web pages alongside generated answers, yet the page-level features that predict citation frequency remain poorly characterised. We present an observational study of approximately 2 million LLM citations from four commercial engines (ChatGPT, Claude, Google AI, Gemini) over six months, joined to 10,000 crawled pages from nineteen B2B SaaS workspaces. Sixty-plus features are tested using a nine-method consensus framework combining mixed-effects regression with domain fixed effects, FDR correction, stability-selection Lasso, double machine learning, generalised additive models, and temporal hold-out replication. Four findings survive all checks. First, prompt-content alignment (Jaccard overlap between page tokens and the full workspace prompt corpus, including non-citing prompts) is the dominant page-level predictor (beta = +0.37, 95% CI [+0.33, +0.41], q ~ 10^-73). Second, the standard AEO checklist (FAQ blocks, structured data, Core Web Vitals) shows positive effects in pooled data that reverse or collapse to zero once domain fixed effects are applied: Simpson's paradox with practical consequences for the AEO literature. Third, domain-level AI authority exceeds the strongest non-alignment page-level feature by a factor of six in mean absolute SHAP value. We release the analytic pipeline as a methodological contribution.
Chinese Translation
生产环境大型语言模型在生成答案的同时会检索并引用网页,然而,预测引用频率的页面级特征仍然未得到充分刻画。我们呈现了一项观察性研究,涵盖来自四个商业引擎(ChatGPT、Claude、Google AI、Gemini)在六个月内的约200万条LLM引用,并将其与来自十九个B2B SaaS工作区的10,000个已抓取页面相连接。使用一个九方法共识框架检验了六十多个特征,该框架结合了带领域固定效应的混合效应回归、FDR校正、稳定性选择Lasso、双重机器学习、广义加性模型以及时间留出复现。有四项发现通过了所有检验。首先,提示词-内容对齐(页面词元与完整工作区提示词语料库之间的Jaccard重叠,包括非引用提示词)是占主导地位的页面级预测因子(beta = +0.37,95% CI [+0.33, +0.41],q ~ 10^-73)。其次,标准AEO清单(FAQ模块、结构化数据、Core Web Vitals)在合并数据中显示出正向效应,但一旦应用领域固定效应,这些效应就会反转或坍缩为零:这是辛普森悖论,并对AEO文献具有实际后果。第三,在平均绝对SHAP值方面,领域级AI权威超过最强的非对齐页面级特征,达到其六倍。我们将分析流程作为方法学贡献予以发布。
cs.AI / 35 / 2609.35109
Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling
仅使用上下文是不够的:面向个性化奖励建模的测试时训练
Bohao Wang, Xiaoyan Zhao, Yang Zhang, Jinghang Guo, Chun Chen, Can Wang, Jiawei Chen
cs.AI
large language model
大语言模型相关
Abstract
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.
Chinese Translation
基于人类反馈的强化学习(RLHF)使大语言模型(LLMs)与人类偏好对齐,然而大多数流程所学习的只是一个单一的奖励模型,忽略了偏好的个体差异。个性化奖励模型(PRMs)通过将奖励建立在用户特定反馈的条件之上来解决这一问题,最常见的方式是借助上下文学习(ICL),即将用户的历史比较作为上下文偏好对提供出来。然而,我们发现了基于 ICL 的 PRMs 的一个关键局限:它们未能捕捉上下文对所传达的偏好关系。为解决这一问题,我们提出了偏好对齐的测试时训练(P-TTT),它将这些关系显式地编码为用户特定的快速权重,以用于个性化奖励预测。P-TTT 引入了序列级的更新与应用操作,以匹配偏好反馈的响应级粒度,并配合一个偏好对齐的目标函数,直接利用成对偏好关系来指导快速权重适配。值得注意的是,P-TTT 实现简单且计算高效,可在单次前向传播中更新快速权重,而无需在推理时进行反向传播。大量实验表明,P-TTT 能够更有效地捕捉历史偏好关系,并以较大幅度优于当前最先进的方法。
cs.AI / 36 / 2609.35110
Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
Sol-H3:用于跨云端与边缘的 Sol-Engine 上 MiniMax-H3 推理加速的递归自我改进
Yitong Li, Jincheng Yu, Junsong Chen, Haopeng Li, Shuchen Xue, Haozhe Liu, Ping Luo, Song Han, Enze Xie
cs.AI · cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.
Chinese Translation
视频扩散模型正在快速扩展,并展现出增强的生成能力。在这些近期进展中,MiniMax-H3 作为一个高度强大、达到生产级水平的开源模型而脱颖而出。然而,其 330 亿参数和多步迭代去噪过程带来了巨大的计算开销。因此,其实际生产受到云部署(如 NVIDIA-GB200)中生成延迟的阻碍,同时还面临严格的内存限制,这在边缘设备(如 DGX-Spark)上进一步构成挑战。为了解决从云端到边缘设备的这些多样化硬件瓶颈,我们提出了一种全栈推理流水线,将高效算法设计与优化后的算子实现相结合。在算法层面,我们引入了一种跨分辨率两阶段生成调度器,它利用了扩散的逐步特性:早期的低分辨率步骤快速建立全局布局,而后期的更高分辨率步骤则专注于局部细节和感知细节的细化。这些阶段通过一个学习到的潜变量到潜变量映射模块连接,完全消除了在不同分辨率之间进行分辨率转换时计算开销高昂的 VAE 解码-重编码循环。在算子实现方面,我们部署了一个递归自我改进(RSI)循环,它搜索内核融合和内存布局,并同时评估延迟与数值一致性。这些优化共同带来了最高 30 倍的端到端加速和降低 20% 的内存占用:在 8xGB200 节点上,一段带音频的 5 秒 1344x768 视频的生成速度比实时快 3.5 倍,并且在单台 DGX Spark 上以完全驻留内存的方式在不到一分钟内生成。
cs.AI / 37 / 2609.35117
Tool Mediation Alters Refusal Mechanisms in Large Language Models
工具中介改变大型语言模型中的拒绝机制
Abel Rodríguez, Giuseppe Garofalo, Lieven Desmet, Vera Rimmer
cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms across a diverse set of open-weight language models. We find that information about the harmfulness of a request remains strongly encoded in the model's representations and transfers across conversational and tool-mediated inputs. Evidence from representation geometry and neuron-level analysis further indicates that the two interaction modes systematically distribute harm-related computation differently. Crucially, while conversational inputs can be refused at relatively low levels of perceived harmfulness, tool-mediated inputs remain permissive until harmfulness crosses a substantially higher effective refusal threshold. Moreover, tool-mediated refusal is also more brittle: progressively weakening the refusal computation disrupts tool-mediated refusal at lower intervention strengths than conversational refusal, even when benign capabilities remain intact. Together, our findings indicate that tool mediation does not simply reduce the internal perception of harm, but instead impacts its conversion into refusal. Overall, this suggests tool-mediated environments may intrinsically reduce robustness of models to harmful requests, and that conventional safety evaluations may not fully transfer to LLM agents.
Chinese Translation
大型语言模型(LLMs)越来越多地被部署为可访问外部工具,然而与常规对话式交互相比,有害的工具中介交互更不容易被拒绝。由于这种拒绝行为的变化仍未得到充分探索,我们在多样化的开放权重语言模型集合中研究其潜在机制。我们发现,关于请求有害性的信息仍然强编码在模型的表示中,并且能够在对话式输入和工具中介输入之间迁移。来自表示几何和神经元层面分析的证据进一步表明,这两种交互模式系统性地以不同方式分配与危害相关的计算。关键的是,尽管对话式输入可以在相对较低的感知有害性水平下被拒绝,但工具中介输入会保持放行,直到有害性跨过一个显著更高的有效拒绝阈值。此外,工具中介的拒绝也更脆弱:逐步削弱拒绝计算时,工具中介拒绝会在比对话式拒绝更低的干预强度下被破坏,即使良性能力仍保持完好。总之,我们的发现表明,工具中介并非仅仅降低了内部对危害的感知,而是影响了其向拒绝的转化。总体而言,这表明工具中介环境可能本质上降低模型对有害请求的鲁棒性,并且传统安全评估可能无法完全迁移到 LLM 智能体。
cs.AI / 38 / 2609.35188
Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
词元之下:GPU 加速大语言模型推理中多词元预测的性能工程研究
Suwesh Prasad Sah
cs.AI · cs.PF
large language model
大语言模型相关
Abstract
Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by \(1.91\times\) to \(2.19\times\) across all prompts and reduced time to first output by 10.0--14.2\%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4--78.1\% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.
Chinese Translation
自回归大语言模型推理反复调用目标模型以一次生成一个词元,使生成对 GPU 内存移动和顺序执行敏感。本研究在 NVIDIA A10G GPU 上受控的单请求部署中,评估两词元多词元预测(MTP)相对于自回归解码的表现。一个 360 请求的基准测试覆盖了纯文本、推理密集型和工具调用工作负载,同时使用运行时遥测、Nsight Systems、PyTorch Profiler 和选定的 Nsight Compute 测量来解释观察到的性能。MTP 在所有提示上将输出吞吐量提高了 \(1.91\times\) 到 \(2.19\times\),并将首次输出时间降低了 10.0--14.2\%。每个验证迭代的平均接受长度中位数范围为 2.370 到 2.595 个词元。性能剖析显示,MTP 引入了更长且更复杂的执行路径,包括提议、采样、注意力、收集和归约操作。然而,每生成一个词元,它所需的选定重复 CUDA Graph 执行次数减少了 56.4--78.1\%。占主导地位的 MTP GEMM 内核并不比占主导地位的自回归 GEMV 内核更快,并且两者的选定实例都接近 A10G 的内存带宽限制。这些结果表明,MTP 通过摊销改善了推理:更大的词元推进量充分减少了重复的 GPU 执行,从而抵消了额外的推测执行成本。
cs.AI / 39 / 2609.35233
EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents
EP-Mem:用于社交关系感知型 LLM 智能体的弹性隐私记忆
Fengzhou Sun, Yuan Zhang, Xintong Yu, Jinyao Yan
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users' social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus on instantaneous interactions, leaving the long-term relational disclosure problem unexplored. In this paper, we propose EP-Mem, an Elastic Privacy Memory architecture that reframes privacy as user-owned boundary control across social roles. EP-Mem introduces (1) token-level memory driven by user-configurable a privacy policy that stratifies persons and events, combining domain-level default circulation rules with fact-level whitelist/blacklist exceptions; and (2) a pluggable sidecar with a privacy engine that aligns disclosure controls with memory across summary, detail, and boundary granularities, enforced throughout generation, storage, and retrieval. We construct EP-Bench, to our knowledge the first long-term multi-party benchmark with cross-session correlated events for policy-conditioned relational disclosure. Experiments show that EP-Mem achieves 94.0% privacy classification accuracy, improves disclosure-permission judgment from 22% to 68%, and reduces privacy leakage by 75.6%, while maintaining retrieval performance and cross-benchmark generalization.
Chinese Translation
大型语言模型(LLM)智能体在人类-智能体-人类通信中充当代表时面临严重的隐私风险。为防止此类泄露,智能体必须理解用户的社会关系,并遵守情境依赖的社会信息披露边界。当前关于智能体记忆隐私的研究聚焦于瞬时交互,未探索长期关系披露问题。在本文中,我们提出 EP-Mem,一种弹性隐私记忆架构,它将隐私重新定义为跨社会角色的用户自有边界控制。EP-Mem 引入了(1)由用户可配置的隐私策略驱动的 token 级记忆,该策略对人物和事件进行分层,将领域级默认流通规则与事实级白名单/黑名单例外相结合;以及(2)一个可插拔的 sidecar,其带有隐私引擎,在摘要、细节和边界粒度上将披露控制与记忆对齐,并在生成、存储和检索全过程中强制执行。我们构建了 EP-Bench,据我们所知,这是首个带有跨会话相关事件的长期多方基准,用于策略条件下的关系披露。实验表明,EP-Mem 实现了 94.0% 的隐私分类准确率,将披露许可判断从 22% 提升至 68%,并将隐私泄露降低了 75.6%,同时保持了检索性能和跨基准泛化能力。
cs.AI / 40 / 2609.35255
Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
迈向可靠的 AI 数据科学家:配备工作流约束框架的数据智能体
Huachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai, Zijin Hong, Chuang Zhou, Zheng Yuan, Qinggang Zhang, Qing Li, Xiao Huang
cs.AI
large language model
大语言模型相关
Abstract
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.
Chinese Translation
大语言模型智能体正越来越多地部署于数据密集型工作,然而可靠的数据分析所需要的不仅仅是通用推理和临时性的工具增强。配备工作流约束框架的数据智能体为自动化端到端数据科学生命周期提供了一种有前景的范式。本文从以约束框架为中心的视角考察数据智能体。首先,我们提出数据智能体及相关数据环境的分类体系,并围绕五个功能阶段组织文献:感知、规划、执行、验证和修复。其次,我们分析每个阶段内的关键技术路线,识别出 15 种不同的方法,涵盖从数据结构探测到数据状态重构。第三,我们识别出四个尚未解决的可靠性问题:语义校准不活跃、澄清缺失、经验迁移缺失,以及验证-修复仓库缺失。这些问题解释了为什么即使各个组件都正确运行,静默失败仍可能持续存在,并凸显了对严格的工作流约束框架和共享可靠性资源的需求。最后,我们总结数据智能体的横向任务族,考察其纵向应用场景以及用于评估的基准,同时维护一个配套仓库:https://github.com/DEEP-PolyU/Awesome-Data-Agents。
cs.AI / 41 / 2609.35285
Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale
文本化用户品味:面向大规模基础模型推荐系统的自然语言用户上下文
Ghazal Fazelnia, Paul Gigioli, Eliza Klyce, Sharon Zheng, Katie Zelvin, Ye Myat Thein, Anurag Deshpande, Seda Davtyan, Kate Remeika, Maya Hristakeva, Erik Franco, Karen Banzon, Peng Ge, Jacqueline Wood, Nandini Singh, David Murgatroyd, Mounia Lalmas, Yves Raimond, Andreas Damianou
cs.AI
large language model
大语言模型相关
Abstract
Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embedding vectors remain highly effective for retrieval and ranking, but they are opaque to users and not natively expressed for language model workflows. We present Textual User Taste, a system that generates structured natural-language taste profiles from listening behavior, interaction signals, content metadata, and optional user feedback, and deploys them to millions of Spotify users. We describe the end-to-end production lifecycle required to generate, evaluate, optimize, and maintain these representations at industrial scale, including prompt development and compression, user steering, and integration with downstream personalization systems. Because no unique ground-truth taste profile exists, we introduce a multi-faceted evaluation framework to evaluate taste profiles as a production representation: they carry user-specific predictive signal independently, and when integrated with behavioral embeddings, improve MRR by 0.6% for future-track prediction and NDCG@7 by 2.2% for search ranking. Our evaluation also reveals that taste profiles support positive natural-language steering, while exposing important limitations, including challenges with negation and short-term temporal adaptation. These findings position taste profiles not as replacements for behavioral embeddings, but as an interpretable and steerable interface between evolving user context and foundation-model recommender systems.
Chinese Translation
基础模型推荐系统需要一种能够被大语言模型消费、在其上进行推理,并通过自然语言交互加以细化的用户上下文。传统行为嵌入向量在检索和排序方面仍然非常有效,但它们对用户而言不透明,并且并非为语言模型工作流原生表达。我们提出“文本化用户品味”(Textual User Taste),这是一个从收听行为、交互信号、内容元数据和可选用户反馈中生成结构化自然语言品味画像,并将其部署给数百万 Spotify 用户的系统。我们描述了在工业规模下生成、评估、优化和维护这些表示所需的端到端生产生命周期,包括提示开发与压缩、用户引导,以及与下游个性化系统的集成。由于不存在唯一的真值品味画像,我们引入一个多层面的评估框架,以将品味画像作为一种生产表示进行评估:它们独立地携带用户特定的预测信号,并且当与行为嵌入集成时,在未来曲目预测中使 MRR 提升 0.6%,在搜索排序中使 NDCG@7 提升 2.2%。我们的评估还揭示,品味画像支持正向自然语言引导,同时暴露出重要局限,包括在否定和短期时序适应方面的挑战。这些发现将品味画像定位为并非行为嵌入的替代品,而是不断演化的用户上下文与基础模型推荐系统之间的一种可解释且可引导的接口。
cs.AI / 42 / 2609.35302
Narrowing the Horizon: Quantifying Topic Saliency Shifts in Generative Monoculture
收窄视野:量化生成式单一文化中的主题显著性偏移
Oriane Peter, Elena Simperl, Kate Devlin
cs.AI
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) become central to how we access and share information, they play an increasingly powerful role in shaping global knowledge. However, as these models evolve, their outputs risk converging into a \textit{generative monoculture}, where the diversity of perspectives they represent narrows over time. Studies at the model level often fail to pinpoint which specific topics or viewpoints are being marginalised or amplified in this process. In this paper, we introduce a method to measure shifts in topic saliency across model families, tracking what gains or loses prominence during post-training. Applying this approach to a case study of climate change discourse, we demonstrate how homogenisation affects the representation of diverse solutions across different models. We also test interventions to counter this trend, showing that specialised models can help preserve a broader range of perspectives. This underscores the importance of monitoring topic saliency to diagnose the risks of monoculture and to ensure AI systems reflect a pluralism of ideas. Data and Code are accessible \href{https://github.com/oriane/topic_saliency_shift}{here}.
Chinese Translation
随着大型语言模型(LLMs)成为我们获取和分享信息的核心,它们在塑造全球知识方面发挥着越来越强大的作用。然而,随着这些模型的发展,它们的输出有趋同于\textit{生成式单一文化}的风险,在这种文化中,它们所代表的观点多样性会随着时间推移而收窄。模型层面的研究往往无法准确指出,在这一过程中哪些具体主题或观点正在被边缘化或被放大。在本文中,我们提出了一种衡量不同模型家族之间主题显著性偏移的方法,追踪在后训练期间哪些内容获得或失去突出地位。我们将这一方法应用于气候变化话语的案例研究,展示同质化如何影响不同模型对多样化解决方案的表征。我们还测试了用以对抗这一趋势的干预措施,表明专用模型可以帮助保留更广泛的视角。这凸显了监测主题显著性的重要性,以便诊断单一文化的风险,并确保人工智能系统反映思想的多元性。数据和代码可在\href{https://github.com/oriane/topic_saliency_shift}{此处}获取。
cs.AI / 43 / 2609.35336
TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving
TMCS:面向组合化学问题求解的工具接地多智能体推理
Shengqin Wang, Jie Jin, Yu Cheng, Yihang Chen, Weilin Luo, Yuan Xie, Zhizhong Zhang
cs.AI · cs.CV
large language model
大语言模型相关
Abstract
Despite the promise of Large Language Models (LLMs) in computational chemistry, rigorous combinatorial chemistry problems remain difficult because they require quantitatively constrained molecular modification, candidate validation, and systematic revision after failed attempts. Existing tool-augmented chemical agents demonstrate useful planning and tool use, but they rarely provide a unified loop for property-driven molecular optimization and workflow-level composition. To bridge this gap, we propose Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving (TMCS), a step-by-step multi-agent framework that formalizes chemical problem solving as an interpretable, tool-augmented workflow. At the task level, specialized agents leverage external tools, few-shot trajectory memory, and structured reflection to iteratively refine solutions. At the workflow level, TMCS chains generation, understanding, editing, description, and optimization into a closed-loop pipeline. Evaluations across multiple chemical tasks demonstrate that TMCS consistently enhances chemical reasoning across both open- and closed-source base models, achieving state-of-the-art performance.
Chinese Translation
尽管大语言模型(LLMs)在计算化学中展现出前景,但严格的组合化学问题仍然困难,因为它们需要定量约束下的分子修饰、候选物验证,以及在尝试失败后的系统性修正。现有的工具增强化学智能体展示了有用的规划与工具使用能力,但它们很少为性质驱动的分子优化和工作流层面的组合提供一个统一的循环。为弥合这一差距,我们提出面向组合化学问题求解的工具接地多智能体推理(TMCS),这是一个逐步式的多智能体框架,将化学问题求解形式化为一个可解释的、工具增强的工作流。在任务层面,专门的智能体利用外部工具、少样本轨迹记忆和结构化反思来迭代地精炼解决方案。在工作流层面,TMCS 将生成、理解、编辑、描述和优化链接为一个闭环流水线。跨多个化学任务的评估表明,TMCS 在开源和闭源基础模型上均持续增强了化学推理能力,达到了最先进的性能。
cs.AI / 44 / 2609.35472
Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
为什么确定性 PRM 指导在离散扩散推理中表现不佳
Yan Zhan, Shaobo Liu, Zhijun Gao
cs.AI
diffusion
扩散模型相关
Abstract
Discrete diffusion language models (dLLMs) expose a denoised solution at every step, which makes process reward model (PRM) guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged in the same budget of forward passes, its deterministic form loses to a much simpler baseline. Our PRMs score intermediate denoising states and are trained on the correctness of the final answer. On Dream-v0-Instruct-7B with 8 candidates per GSM8K problem, keeping the candidate with the highest PRM score at every scoring step reaches 65.18%, while independent sampling plus an ORM reranker trained for the task reaches 75.13%. The gap grows to 12.69 percentage points (pp) with 32 candidates, and is 9.85 pp on MATH and 12.16 pp on MBPP. We trace it to two separable failures. First, guidance prunes on a weak signal: on GSM8K, PRM ROC-AUC falls from 0.77 to 0.54 as the mask ratio rises, a decay that persists when states are relabeled with fresh rollouts, and pruning lowers the best accuracy reachable from the candidate pool from 81.05% for independent samples to 67.30%. Second, on GSM8K and MATH, the PRM is a poor final judge: a sequential Monte Carlo sampler at the same budget restores that ceiling to 77.89%, yet selecting with the PRM gives 65.48%, on par with deterministic guidance, while a PRM retrained on final states matches the ORM on identical candidates. MBPP separates the two: there the PRM reaches 65.47% when reranking finished programs, on par with the ORM, but 50.88% when it guides denoising. The results point to two targets for dLLM guidance: keep correct partial solutions alive through early denoising, and leave the final choice to a verifier trained on final states. We release the corpus of denoising states with outcome labels and evaluation toolkit for reproducible comparisons at matched compute.
Chinese Translation
离散扩散语言模型(dLLM)在每一步都会给出一个去噪后的解,这使得过程奖励模型(PRM)指导看起来像是一种在测试时消耗算力的方式。我们表明,一旦去噪、PRM 评分和结果奖励模型(ORM)评分都被计入相同的前向传播预算,其确定性形式就会输给一个简单得多的基线。我们的 PRM 对中间去噪状态进行评分,并依据最终答案的正确性进行训练。在 Dream-v0-Instruct-7B 上,对每个 GSM8K 问题使用 8 个候选,每个评分步骤都保留 PRM 分数最高的候选达到 65.18%,而独立采样加上为该任务训练的 ORM 重排序器达到 75.13%。当候选数增加到 32 时,差距扩大到 12.69 个百分点(pp),在 MATH 上为 9.85 个百分点,在 MBPP 上为 12.16 个百分点。我们将其追溯到两个可分离的失败。首先,指导基于一个弱信号进行剪枝:在 GSM8K 上,随着掩码比例上升,PRM ROC-AUC 从 0.77 降至 0.54,这种衰减在用新的 rollout 重新标注状态后仍然存在,而剪枝将候选池中可达到的最佳准确率从独立采样的 81.05% 降低到 67.30%。其次,在 GSM8K 和 MATH 上,PRM 是一个糟糕的最终评判者:在相同预算下,顺序蒙特卡洛采样器将该上限恢复到 77.89%,但用 PRM 进行选择只得到 65.48%,与确定性指导相当,而在最终状态上重新训练的 PRM 在相同候选上与 ORM 相当。MBPP 将二者区分开来:在那里,PRM 在对已完成程序进行重排序时达到 65.47%,与 ORM 相当,但当它指导去噪时只有 50.88%。结果指向 dLLM 指导的两个目标:在早期去噪过程中让正确的部分解保持存活,并把最终选择留给在最终状态上训练的验证器。我们发布带有结果标签的去噪状态语料库以及评估工具包,以便在匹配算力下进行可复现比较。
cs.AI / 45 / 2609.35501
SRHarness: A Harness for Agentic Symbolic Regression
SRHarness:一种用于智能体式符号回归的运行框架
Zihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding, Yong Li
cs.AI · cs.LG · cs.SC
large language model
大语言模型相关
Abstract
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.
Chinese Translation
近期的智能体式符号回归方法越来越依赖于大语言模型来分析数据、选择科学操作,并在漫长的搜索轨迹中不断精炼假设。在此类系统中,性能不仅取决于底层模型与搜索策略,还取决于支撑科学搜索的运行时基础设施。我们提出 SRHarness,一个面向智能体式符号回归的领域专用运行框架,它围绕三种机制构建:可组合的科学操作,为原始量、变换后的量以及由候选解导出的量提供统一接口;持久化的科学状态,保留已评估的假设并暴露紧凑的、面向模型的视图;以及轨迹生命周期管理,协调继续、分支、重启与终止。在 LLM-SRBench 上,在相同的大语言模型骨干下,SRHarness 在数值泛化与符号恢复两方面均取得了一致提升。在使用 DeepSeek-v4-flash-0731 时,它在 LSR-Transform 上取得了 93.69% 的符号准确率,而 SR-Scientist 为 62.16%;在去除了科学描述与变量语义的匿名化变体上,它仍保持 72.97% 的准确率,而 SR-Scientist 为 39.64%。在相同的 DeepSeek-v4-flash-0731 骨干下,SRHarness 也显著优于 Codex(72.97% 对 20.72%),并达到与使用 GPT-5.5 的 Codex 相当的性能;而仅仅为 Codex 提供相同的科学工具并不能复现这一优势。这些结果表明,有效的智能体式符号回归不仅取决于模型或工具,还取决于用于组织科学操作、累积假设与长时程搜索的结构化运行时支持。
cs.AI / 46 / 2609.35540
Continuous Context Management
连续上下文管理
William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan
cs.AI
large language model
大语言模型相关
Abstract
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
Chinese Translation
长时程大语言模型(LLM)智能体通常保留其完整的交互历史,直到在预定义阈值处触发压缩。我们研究连续上下文管理(CCM),它在每一轮都执行压缩,以防止交互历史在活动提示中累积。在每一轮中,CCM 智能体在输出环境动作的同时输出更新后的记忆;其下一个提示包含原始任务、保留的记忆以及最新的观测,而不是完整的对话记录。我们首先使用 Claude Sonnet 4.6、Claude Opus 4.6、GLM-5 和 Kimi K3,在 TerminalBench-2 上评估未经微调的 CCM。CCM 大幅降低了累计输入用量和活动提示规模,尽管它降低了大多数模型的任务成功率,但保持了 Kimi K3 的性能。我们使用带有特权全历史蒸馏的 GRPO 来改进开放权重模型中的 CCM。学生模型初始模型的一个冻结副本,在该学生 rollout 重建出的完整历史下,对每个采样的学生动作进行评分,从而在没有单独的教师 rollout 或参考解的情况下提供密集的动作 token 监督。在 WebShop 上,该目标在两种评估的模型规模下都使 CCM 较 GRPO 有显著提升,并在 Qwen3-4B-Instruct 上超越了全历史 GRPO,尽管在 Qwen3-8B 上未能如此。在 Endless Terminals 上,该增强方法相较 GRPO 带来了适度提升,且两种 CCM 策略都优于未经训练的全历史基线。这些结果表明,CCM 是一种可行的推理范式,适用于在显著减少保留上下文的情况下运行的智能体,并且其性能可以通过带有特权全历史蒸馏的强化学习得到提升。
cs.AI / 47 / 2609.35559
From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
从搜索到研究:探索自主量化因子挖掘中的搜索扩展
Kangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian, Rui Sun, Beidi Luan, Jing Li, Daxin Jiang, Zuo Bai
cs.AI
large language model
大语言模型相关
Abstract
Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.
Chinese Translation
推理扩展已被证明能够提升大语言模型(LLM)的性能,而这一原理通过增加搜索预算自然地扩展到自主 LLM 智能体,我们将其称为*搜索扩展*。尽管先前工作已经刻画了 LLM 推理扩展的机制、扩展行为和性能极限,但对于自主研究中的这些问题却知之甚少。因此,我们使用 50 个基于金融研究报告的量化因子挖掘任务,研究搜索扩展如何影响研究表现,以及哪些机制驱动了这些收益。每个任务都要求智能体执行一个端到端的研究循环,从解释一个假设并将其在代码中实现,到评估并迭代改进由此得到的因子。我们在九个模型上,通过追踪不同预算下的表现、在模型之间迁移中间研究状态以及比较不同搜索策略,考察模型能力、搜索深度和搜索组织方式如何塑造因子质量。我们发现:(1)初始表现与模型能力关联更强,而更深的搜索可以缩小跨模型差距;(2)模型嫁接表明,早期研究状态会实质性地塑造最终表现;以及(3)在相同迭代预算下,并行搜索优于顺序搜索,这与更广泛覆盖搜索空间所带来的收益一致。进一步的轨迹分析表明,表现更好的模型在选择候选因子时,能够更有效地诊断失败、修正搜索方向,并保留预期的经济假设。这些发现表明,自主研究的未来进展将需要更强的模型,以及在整个研究过程中部署测试时计算的自适应策略。
cs.AI / 48 / 2609.35576
Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
共享传播的AI病毒:跨LLM智能体的记忆跳跃攻击
Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina, Ivaxi Sheth, Mario Fritz
cs.AI · cs.CL · cs.CR · cs.LG
large language model
大语言模型相关
Abstract
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.
Chinese Translation
大型语言模型正越来越多地被部署为有状态助手,这些助手能够在交互之间保留信息,并使用工具读取、修改和创建持久性产物。由于这些产物在用户之间共享,它们在原本相互独立的助手之间形成了一条间接通信通道。我们研究了一种失效模式,其中这一通道使自我传播攻击成为可能。我们引入产物介导传播,其中通过产物(例如一份报告)引入的对抗性内容被存储在某个助手的持久记忆中,在随后创建的产物中被复制,并被另一个之后读取它的助手获取。我们在时间性的人类-智能体宇宙中评估这一过程,这些宇宙对独立运行的助手之间随时间的产物交换进行建模,衡量攻击是否在连续交接中存活、它到达多少跳,以及它传播得有多广。我们发现,攻击可以跨越多个独立助手传播,并在延长的交互序列中持续存在。在更大的模拟环境中,即使是 GPT-5.6 Luna 也表现出显著传播,触及 60-80% 的智能体,传播链延伸至八跳。这些结果表明,持久性产物可以充当对抗状态的持久载体,使攻击能够比单次交互存在得更久,并跨越孤立的助手传播。
cs.AI / 49 / 2609.35586
IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
IMC-CLINIC:用于模拟存内计算中裁剪的耦合损失感知牛顿迭代
Yung-Chin Chen, Chia-Yu Chen, Naveen Verma
cs.AI
large language model
大语言模型相关
Abstract
Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize these coupled sources of IMC error and often rely on costly search-based calibration. We introduce IMC-CLINIC (Coupled Loss-Informed Newton Iterations for Clipping), a clipping calibration framework based on an analytical surrogate for IMC MatMul output error. The surrogate jointly models operand quantization, accumulated clipping-induced bias, and ADC quantization, enabling efficient evaluation of its gradient and approximate curvature from a small calibration set. IMC-CLINIC jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5-11.5 percentage points over the grid search baseline while reducing calibration time by factors of 10.0-12.1. Its analytical surrogate closely tracks empirical IMC output error, and its optimizer is certified within 1% of the global optimum under the loss objective across all projections on two representative models.
Chinese Translation
模拟存内计算(IMC)通过在模拟域中直接于存储阵列内执行矩阵乘法(MatMul),为高能效的大语言模型(LLM)推理提供了一条颇具前景的路径。然而,其高效率也伴随着一个额外的误差来源:有限精度的模数转换器(ADC)对累积的模拟部分和进行量化,从而引入了有别于 MatMul 输入端常规激活与权重量化的输出端误差。裁剪可以同时缓解操作数量化误差与 ADC 量化误差,但最优裁剪因子必须联合权衡激活的舍入与裁剪、权重的舍入与裁剪以及 ADC 量化。现有的裁剪方法是为数字量化而设计的,并未显式优化这些相互耦合的 IMC 误差来源,且往往依赖代价高昂的基于搜索的校准。我们提出 IMC-CLINIC(用于裁剪的耦合损失感知牛顿迭代),这是一个基于 IMC MatMul 输出误差解析代理模型的裁剪校准框架。该代理模型联合建模了操作数量化、累积的裁剪引入偏置以及 ADC 量化,从而能够基于一个小型校准集高效地评估其梯度与近似曲率。IMC-CLINIC 采用一种带防护措施的牛顿型方法联合优化激活与权重的裁剪因子。在多个模型与数据集上,相较于网格搜索基线,它将平均零样本准确率提升了 6.5-11.5 个百分点,同时将校准时间缩短为原来的 1/10.0-1/12.1。其解析代理模型能够紧密贴合经验 IMC 输出误差,并且在两个代表性模型的所有投影上,其优化器被证明处于损失目标下全局最优解的 1% 以内。
cs.AI / 50 / 2609.35599
Signatures of semantic search in the activations of large language models
大型语言模型激活中的语义搜索特征
Luke Leckie, Peter M. Todd, Jacob G. Foster
cs.AI
large language model
大语言模型相关
Abstract
When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and between-cluster switching ("explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., "water") increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.
Chinese Translation
在语义流畅性任务(semantic fluency task, SFT)中回忆概念列表(例如动物)时,人类和大型语言模型(LLMs)都会将其输出组织成由相关项目(例如海洋动物)构成的簇,并在簇之间穿插策略性切换。在人类中,这种模式可以由语义觅食过程解释,其中不同的神经和行为特征伴随簇内产出(“利用”)和簇间切换(“探索”)。LLMs 是否同样在其内部状态中表征这两种搜索机制尚不清楚。在这里,我们应用一系列机制可解释性技术来为这一点提供证据。在研究1中,我们使用雅可比透镜(Jacobian lens, J-lens),它将中间层残差流表示映射到词元级激活,来表明概念级激活可预测切换。首先,我们发现切换与低下一词元激活相吻合。此外,随着最强 J-lens 激活的集合(J 空间)中当前正在产出的类别项目逐渐耗尽,切换概率上升,这类似于斑块觅食过程中的探索-利用决策。然后我们表明,抽象类别相关标签(例如“水”)的中间层 J-lens 激活会在预期切换到该类别时增加。我们通过推导靶向类别切换的引导向量,确认这些表征对切换具有因果影响。在研究2中,我们识别出在切换期间以及预期切换时被激活的通用残差流方向。通过沿这些方向引导激活,我们使切换率偏向增加或减少。我们的研究将语义觅食框架扩展到人工智能,并提供证据表明:LLMs 在将概念信息言语化时,为探索和利用维持着不同的表征特征。
cs.AI / 51 / 2609.35606
TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science
TCSAlgBench:面向研究级理论计算机科学的自动证明基准测试
Chutong Yang, Xiyuan Zhang, Yu Huang, Boran Han, Soonho Kong, Shuai Zhang, Vihang Prakash Patil, Zhen Han, Michael Bohlke-Schneider, Bernie Wang
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery, comprising 398 theorem-level challenges from 138 STOC and COLT 2026 papers. Expert-designed rules complete paper-specific context, preserve computational assumptions and quantitative guarantees, and withhold constructions when discovering an algorithm is part of the task. For each task, prover systems receive theorem statements and access to cited prior work. The pipeline supports fresh, versioned challenge batches from newly released papers. We evaluate ten model configurations from four families under direct inference and prover-verifier discussion, and compare four agent workflows under matched model-call opportunities. All evaluations use the full benchmark. In the model comparison, GPT-5.6 Sol max achieves the highest five-run verifier-accepted coverage at 23.6% after 10-round discussion. Discussion and repeated sampling improve coverage. In the separate agent comparison using GPT-5.5 xhigh, decomposition improves coverage over discussion, and agentic planning achieves the highest five-run verifier-accepted coverage at 25.4%. TCSAlgBench provides a refreshable testbed for measuring progress in model reasoning and studying how agent workflows support research-level proof discovery.
Chinese Translation
大型语言模型在竞赛数学上表现强劲,但其研究级推理仍然难以被系统评估。理论计算机科学(TCS)将算法设计与明确的保证和根本极限联系起来,为评估模型能否用人类可审查的论证来证明计算改进提供了场景。我们提出 TCSAlgBench,一个用于自然语言证明发现的基准和可复用流程,包含来自 138 篇 STOC 和 COLT 2026 论文的 398 个定理级挑战。由专家设计的规则补全论文特定的上下文,保留计算假设和定量保证,并在发现算法属于任务的一部分时隐去构造。对于每个任务,证明器系统会收到定理陈述,并可以访问被引用的先前工作。该流程支持从新发布的论文中生成新的、带版本号的挑战批次。我们在直接推理和证明器-验证器讨论下评估了来自四个家族的十种模型配置,并在匹配的模型调用机会下比较了四种智能体工作流。所有评估都使用完整基准。在模型比较中,GPT-5.6 Sol max 在 10 轮讨论后取得了最高的五轮验证器接受覆盖率,为 23.6%。讨论和重复采样提高了覆盖率。在使用 GPT-5.5 xhigh 的单独智能体比较中,分解相比讨论提高了覆盖率,而智能体规划取得了最高的五轮验证器接受覆盖率,为 25.4%。TCSAlgBench 提供了一个可刷新的测试平台,用于衡量模型推理的进展,并研究智能体工作流如何支持研究级证明发现。
cs.AI / 52 / 2609.35623
RIDE: Reference-Anchored Inference-Time Diffusion Editing for Scaffold Hopping
RIDE:用于骨架跃迁的参考锚定推理时扩散编辑
Ruoxi Gao, Frazier N. Baker, Trieu Nguyen, Xia Ning
cs.AI · cs.LG
diffusion
扩散模型相关
Abstract
Scaffold hopping is a critical task in drug discovery, which seeks to discover new, structurally distinct molecules that share key functional groups and similar 3D shape with a reference binding ligand. Existing diffusion-based scaffold hopping methods formulate the problem as conditional generation of scaffolds given the functional groups. However, they lack a principled mechanism to jointly enforce 2D structural novelty and preserve the 3D shape of the reference ligand. Here, we introduce RIDE, a Reference-anchored Inference-time Diffusion Editing framework for scaffold hopping. RIDE recovers the reference diffusion noise trajectory conditioned on the binding pocket and functional groups, selects an optimal trajectory segment for editing via noise perturbation, and conducts a value-guided scaffold sampling to generate new scaffolds. Extensive experimental results demonstrate that, compared to baselines, RIDE consistently generates scaffolds with lower 2D similarity and higher 3D similarity to the reference, with an average improvements of 11.7% and 7.3%, respectively. Further analysis reveals that RIDE can accommodate various reward functions, and can preserve 3D similarity even when this is not explicitly included in the reward. Two case studies illustrate RIDE's ability to generate distinct scaffolds with different structures and properties, and its ability to introduce substantial 2D variation while maintaining very high 3D similarity. RIDE is publicly available at https://anonymous.4open.science/r/RIDE-C8A0.
Chinese Translation
骨架跃迁是药物发现中的一项关键任务,其旨在发现与参考结合配体共享关键官能团且具有相似3D形状的、结构上不同的新分子。现有的基于扩散的骨架跃迁方法将该问题表述为在给定官能团条件下进行骨架的条件生成。然而,它们缺乏一种原则性机制来同时确保2D结构新颖性并保持参考配体的3D形状。在此,我们提出RIDE,一个用于骨架跃迁的参考锚定推理时扩散编辑框架。RIDE以结合口袋和官能团为条件恢复参考扩散噪声轨迹,通过噪声扰动选择用于编辑的最优轨迹片段,并进行价值引导的骨架采样以生成新骨架。大量实验结果表明,与基线相比,RIDE始终生成相对于参考具有更低2D相似性和更高3D相似性的骨架,平均改进分别为11.7%和7.3%。进一步分析表明,RIDE能够适应各种奖励函数,并且即使当3D相似性未被显式包含在奖励中时,也能保持3D相似性。两个案例研究展示了RIDE生成具有不同结构和性质的独特骨架的能力,以及其在保持极高3D相似性的同时引入显著2D变化的能力。RIDE已公开提供于 https://anonymous.4open.science/r/RIDE-C8A0。
cs.AI / 53 / 2609.35641
Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
可验证视觉奖励从合成场景到自然提示的迁移
Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer
cs.AI · cs.CV
diffusion
扩散模型相关
Abstract
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.
Chinese Translation
图像生成中的精确指令遵循,例如满足物体计数和空间关系,仍然是一个开放性挑战,至少部分原因是它是使用诸如物体检测器和视觉语言模型之类的不可靠奖励模型来学习的。我们提出可验证视觉奖励(Verifiable Visual Rewards,VVR),这是首个用于可程序化验证图像奖励的框架,并表明在其上训练可泛化到自然提示。每个 VVR 任务都是一个由几何对象及其相互关系组成的场景,我们从中同时推导出提示和确定性验证器,因此可以以任意数量和任意选定复杂度生成任务。我们发布 VVRBench,包含 10,000 个任务,覆盖 32 种约束类型,以及 VVRBench-Challenge,包含 720 个更复杂的任务;我们评估的最强模型——GPT-Image-2.5——解决了 VVRBench-Challenge 的 21.4%。使用 VVR 分数作为强化学习奖励(RLVVR)将 Stable Diffusion 3.5 Medium 在 VVRBench 上的准确率从 2.8% 提高到 28.3%,并展现出一致的从易到难泛化。这些增益扩展到域外基准,并且将 VVR 混入现有目标进一步提升了整体性能和人类偏好,推动将 VVR 采纳到标准图像生成后训练配方中。
cs.AI / 54 / 2609.35643
Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization
并非所有思考都生而平等:潜在推理发现一种用于深度泛化的循环搜索算法
Huzi Cheng, Zhewei Zhang
cs.AI
large language model
大语言模型相关
Abstract
Large Language Models can perform multi-step reasoning and improve task performance through different forms of intermediate computation, from token-based traces to computation carried out in latent space. However, a question remains open: do these different forms of thinking rely on the same underlying mechanism? To address this, we train and compare five variants of the same GPTNeoX backbone from scratch on an extended multi-hop reasoning task (ProsQA-Ext): a vanilla model, a Chain-of-Thought (CoT) model, a Pause Token model, and two latent-reasoning models that are optimized end-to-end without intermediate reasoning traces. We find that, strong in-distribution (ID) performance does not guarantee depth generalization. Vanilla, CoT, and Pause Token models solve ID problems well, but rely largely on local graph features and generalize poorly to out-of-distribution (OOD) problems with longer hops. In contrast, latent variants generalize better and show internal dynamics consistent with forward reachability propagation on the graph. Causal interventions and circuit analysis localize this computation to a sparse recurrent search circuit in the bottleneck latent model: an attention head retrieves graph relations, an MLP and the residual stream update the reachability state across recurrent steps, while multiple attention heads together then do the candidate matching. Together, these results show that different thinking mechanisms can learn distinct computational solutions, even at similar ID performance. In this setting, latent recurrence supports a reusable forward-search algorithm that generalizes beyond the training depth.
Chinese Translation
大型语言模型可以通过不同形式的中间计算来执行多步推理并提升任务性能,这些形式从基于 token 的轨迹到在潜在空间中进行的计算不等。然而,一个问题仍然悬而未决:这些不同形式的思考是否依赖相同的底层机制?为了解决这个问题,我们在一个扩展的多跳推理任务(ProsQA-Ext)上从头训练并比较了同一 GPTNeoX 骨干网络的五种变体:一个 vanilla 模型、一个思维链(CoT)模型、一个 Pause Token 模型,以及两个在没有中间推理轨迹的情况下进行端到端优化的潜在推理模型。我们发现,强大的分布内(ID)性能并不保证深度泛化。Vanilla、CoT 和 Pause Token 模型能很好地解决 ID 问题,但在很大程度上依赖局部图特征,并且对具有更长跳数的分布外(OOD)问题泛化很差。相比之下,潜在变体泛化得更好,并表现出与图上前向可达性传播一致的内部动力学。因果干预和电路分析将这一计算定位到瓶颈潜在模型中的一个稀疏循环搜索电路:一个注意力头检索图关系,一个 MLP 和残差流在循环步骤中更新可达性状态,而多个注意力头随后共同进行候选匹配。总之,这些结果表明,不同的思考机制可以学到不同的计算解决方案,即使在相似的 ID 性能下也是如此。在这种设定下,潜在循环支持一种可复用的前向搜索算法,该算法能够泛化到训练深度之外。
cs.AI / 55 / 2609.35694
Reasoning with Continuous Latent Diffusion
使用连续潜在扩散进行推理
Xiang Cheng
cs.AI
diffusion
扩散模型相关
Abstract
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/LFRM
Chinese Translation
连续扩散通过在潜在空间中的迭代细化生成完整的推理解决方案。我们提出潜在流推理模型(Latent Flow Reasoning Models,LFRMs),一种基于 ELF 的训练与推理方案。我们的实验表明,仅靠准确解码并不能确保强大的推理性能。因此,我们从强大的自回归教师模型的多个层中学习紧凑表示。它们的分解还使得能够以不同速率进行异步去噪。我们表明,提示编码只需保留正确的文本条件得分所需的信息,而不是与教师特征完全匹配;并且我们使用分阶段课程来学习一个紧凑的提示编码器,该编码器在推理时替代教师 Transformer。我们将 DiffusionNFT 适配到学习到的自条件引导,并纳入金标准解端点以补充稀疏奖励。我们的监督模型在数学推理和 HumanEval 代码生成任务上,在可比的骨干规模下,优于近期连续扩散基线所报告的结果。在具有 638M 参数的去噪骨干和学习到的提示条件下,post-NFT 的 LFRM-L 在 64 个去噪步骤下在 GSM8K 上达到 63.74% 的 pass@1、在 MATH500 上达到 24.6%,并在 128 个去噪步骤下在 HumanEval 上达到 32.85%、在 HumanEval+ 上达到 30.18%。代码将在以下网址提供:https://github.com/chengxiang/LFRM
cs.AI / 56 / 2609.35706
Reinforcing Agentic Creativity in Scientific Ideation with Night Science
以夜间科学强化科学构想中的智能体式创造力
Priyanka Kargupta, Silviu Cucerzan, Shweti Mahajan, Allen Herring, Jiawei Han, Ryen W. White, Sujay Kumar Jauhar
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
Chinese Translation
大语言模型(LLMs)擅长结构化的、可验证的任务,但其低熵偏置会产生同质且可预测的输出,限制了它们在开放式科学构想中的效用。然而,有效的发现涵盖更广泛的创造力谱系:从结构化的日间科学到结构松散、充满意外发现的夜间科学,后者能够触及通常不会被考虑的想法。我们提出 AI Night-Scientist,这是一个智能体框架,使用强化学习来教会模型何时以及如何偏离可预测的推理。我们以认知科学为基础,沿三个轴对创造力建模:行动(做什么以及如何有创造性地做)、过程(何时探索与何时利用)以及结果(由此产生的想法的新颖性和有用性)。我们使用这些轴,通过 GRPO 训练模型,并在整个训练过程中让模型接触不同程度和形式的创造力。这产生了显著更多样化的科学提案,相较基础模型,将研究方向的范围扩大了 27.8%,将贡献类型扩大了 14.9%。它还将预测的引用影响力最多提高了 32.0 个百分点,并将原创性提高了 66.2 分。这些增益无法仅通过提高解码温度来复现;相反,我们发现,指定追求何种创造力的语义引导至关重要。总体而言,我们的结果表明,创造力是一种可学习的、多层级的能力,可以被塑造,以帮助研究人员触及超越 LLMs 通常探索的那些想法。
cs.CL / 57 / 2609.33886
LLMs learn different forms of metacognition when trained to predict their own accuracy
LLMs 在被训练预测自身准确率时学会不同形式的元认知
Nicolas Yax, Stefano Palminteri, Pierre-Yves Oudeyer
cs.CL
large language model
大语言模型相关
Abstract
Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.
Chinese Translation
大型语言模型被训练为总是给出一个答案,无论它们是否拥有相关知识,这导致它们编造事实。先前的工作已表明,LLMs 的置信度估计与其实际表现对应得很差,而微调可以显著改善它们。然而,模型在此类训练过程中实际学到了什么,仍然鲜为人知。我们研究 LLMs 如何获得元认知监控——即知道自己知道什么的能力——方法是训练 10 个开放权重的 LLM,让它们在回答事实性多项选择题之前先预测自己的准确率。我们发现,经过训练的置信度反映了两种不同的信号。在与训练数据接近的问题上,它追踪模型的真实准确率;而在其他领域中,它转而追踪输出一致性:即模型答案分布的集中程度。输出一致性追踪在训练早期就出现,并能跨数据集泛化,而准确率追踪则发展得更晚,且始终局限于训练分布之内。这些结果表明,校准训练可能并不会教会模型普遍地检测出它们自信犯下的错误,并且它们就人工系统中元认知的本质提出了更广泛的疑问。
cs.CL / 58 / 2609.33887
Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees
具有无分布风险保证的更快块扩散服务
Jungseob Lee, Dongyub Jude Lee, Chanjun Park, Sugyeong Eo, Heuiseok Lim
cs.CL · cs.LG
diffusion
扩散模型相关
Abstract
Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slower one answers correctly. On the serving engine and its decode traces, the default commit rule already commits every fully resolved block, a static skip rule captures nearly all of the compute that allocation can save, and self-distillation on engine-decoded targets adds speed at unchanged accuracy. Larger speedups come from lower thresholds, which commit tokens that are still uncertain. We therefore present Redline, a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts. Redline keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget with high probability, and deploys the fastest configuration that passes. It speeds up math at a smaller risk budget than code in both model families, and at a budget of ten percent it deploys a LLaDA2 math configuration that commits over a third more tokens in each forward. It also applies without modification to the acceptance rule of speculative decoding and to weight quantization. On the same calibration data, Redline stays within its stated failure probability, whereas each tolerance of a mean-accuracy rule either gains less speed for some model and task or exceeds the risk budget far more often for another. Code is available at https://github.com/js-lee-AI/Redline.
Chinese Translation
块扩散语言模型在人工挑选的工作点上提供服务,例如接受阈值、缓冲区深度、调度、检查点和精度,而每个工作点都是依据其平均基准准确率来选择的。然而,均值并不能告诉运维者:一个更快的配置在较慢配置能正确回答的提示上失败的频率有多高。在服务引擎及其解码轨迹上,默认的提交规则已经会提交每一个完全解析的块,一条静态跳过规则能够捕获分配所能节省的几乎全部计算量,而基于引擎解码目标的自蒸馏在准确率不变的情况下提升了速度。更大的加速来自更低的阈值,而这会提交仍然不确定的词元。因此,我们提出 Redline,一种有限样本过程,它根据各工作点在校准提示上回答的正确性,来选择人工挑选的或学习得到的工作点。Redline 以高概率将参考相对风险——即参考配置回答正确而候选配置回答不正确这一联合概率——控制在用户选定的预算之内,并部署通过检验的最快配置。在两个模型家族中,它以比代码更小的风险预算加速数学任务;在百分之十的预算下,它部署了一个 LLaDA2 数学配置,该配置在每次前向中多提交超过三分之一的词元。它还可以不加修改地应用于投机解码的接受规则以及权重量化。在同一校准数据上,Redline 保持在其所声明的失败概率之内,而平均准确率规则的每一个容差取值,要么在某些模型和任务上获得的速度更少,要么在另一些模型和任务上远远更频繁地超出风险预算。代码见 https://github.com/js-lee-AI/Redline。
cs.CL / 59 / 2609.33947
Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported
简单扩散语言模型作为少步生成器比已报道的更有效
Hasan Amin, Ming Yin, Rajiv Khanna
cs.CL · cs.LG
diffusion
扩散模型相关
Abstract
Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We show that much of the supposed quality gap at few steps can instead arise from a suboptimally configured sampler. Modest sampler sharpening, without any model retraining, enables a couple years old masked DLM to rival supposedly far improved successors. This differently sampled DLM in fact achieves lower generative perplexity in just 16 steps than what its standard sampler obtains with 1024, while improving both judged quality and semantic diversity. We further show that conventional per-output metrics can fundamentally obscure these gains, since any optimal trade-off between two such metrics can be attained by a generator supported on at most two outputs. We subsequently introduce GroupEval, which separately evaluates quality and across-output semantic diversity, and offers fresh insights including uncovering how 1.5-4.7x perplexity gains of a distilled model yield no corresponding quality gain. Finally, we explain why sharpening helps: parallel unmasking destroys dependencies among simultaneously generated tokens, creating a gap between prediction and generation. We prove that pervasive temperature choice of one is generically suboptimal under parallel sampling even for an exact denoiser, and that worse predictions can yield better samples. Through these results, we argue for a broader evaluation principle of treating the deployed generator as the object of comparison, benchmarking it against tuned baselines, and assessing quality and diversity jointly and with more human-aligned measures.
Chinese Translation
扩散语言模型(DLMs)有望实现快速并行生成,但高质量样本往往需要大量细化步骤,这削弱了它们在实际中的优势。这引发了人们对有效的少步生成新方法的巨大兴趣,并推动了其快速发展。我们表明,少步下所谓的质量差距,很大一部分其实可能源自配置次优的采样器。适度的采样器锐化,无需任何模型再训练,就能使一个已有几年历史的掩码 DLM 与据称远为改进的后继模型相媲美。实际上,这种采用不同采样方式的 DLM 仅用 16 步就实现了比其标准采样器用 1024 步所获得的更低的生成困惑度,同时提升了评判质量和语义多样性。我们进一步表明,传统的逐输出指标可能从根本上掩盖这些收益,因为两个此类指标之间的任何最优权衡,都可以由一个支撑集至多包含两个输出的生成器达到。我们随后引入 GroupEval,它分别评估质量和跨输出语义多样性,并提供新的洞见,包括揭示一个蒸馏模型的 1.5–4.7 倍困惑度增益如何没有产生相应的质量增益。最后,我们解释了为什么锐化有帮助:并行解掩码破坏了同时生成的词元之间的依赖关系,从而在预测与生成之间造成差距。我们证明,即使对于精确的去噪器,在并行采样下,普遍采用温度取 1 的选择通常也是次优的,并且更差的预测可以产生更好的样本。通过这些结果,我们主张一个更广泛的评估原则:将所部署的生成器视为比较对象,将其与经过调优的基线进行基准比较,并联合评估质量与多样性,且采用更与人类对齐的度量。
cs.CL / 60 / 2609.33970
On the Token Value Inequality in Efficient Reasoning
论高效推理中的词元价值不平等
Runjia Zeng, Hang Hua, Yiyang Liu, Zhiqiang Tao, Ruixiang Tang, Qifan Wang, Cheng Han, Dongfang Liu
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the reasoning trace equally valuable? We present a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals. We show that normalized log probability helps distinguish core tokens, which carry structural and decisive reasoning content, from redundant tokens, which are exploratory, low-confidence filler that contributes less directly to the final answer. Building on these findings, we formulate the TokenProbe framework around two empirical findings and one claim: findings identify token value inequality first and then establish TokenProbe as a core-token proxy, and the claim introduces an efficient GRPO objective positing that selectively compressing redundant tokens can yield Pareto improvements in the accuracy-token efficiency space. Empirically, our method preserves reasoning quality while reducing the token usage by 76% of the baseline. Under matched reasoning-length budgets, we show that it can even outperform strong flagship baselines like Gemini-3.1-Pro. Homepage: https://runjia.tech/tokenprobe/.
Chinese Translation
思维链推理使大语言模型能够在复杂任务上取得显著性能提升。然而,这些收益是以词元消耗急剧增加为代价的。这引出了一个基本问题:推理轨迹中的每个词元都同等有价值吗?我们提出了一个基于一项关键实证发现的诊断与优化框架:CoT 推理序列中词元的价值高度不均匀,而这种不均匀性可以通过词元级对数概率信号得到有效刻画。我们表明,归一化对数概率有助于区分核心词元与冗余词元:核心词元承载结构性和决定性推理内容,而冗余词元则是探索性的、低置信度的填充内容,对最终答案的直接贡献较小。基于这些发现,我们围绕两项实证发现和一项主张构建 TokenProbe 框架:这些发现首先识别出词元价值不平等,然后将 TokenProbe 确立为核心词元的代理,而该主张引入了一个高效的 GRPO 目标,其假设是,选择性压缩冗余词元可以在准确率-词元效率空间中带来帕累托改进。在经验上,我们的方法在保持推理质量的同时,将词元使用量降低了基线的 76%。在匹配的推理长度预算下,我们表明它甚至可以超越像 Gemini-3.1-Pro 这样的强大旗舰基线。主页:https://runjia.tech/tokenprobe/。
cs.CL / 61 / 2609.33974
Beyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive Pruning
超越单智能体与一致性:通过条件渐进剪枝为多智能体辩论正名
Ruosong Ye, Caiqi Zhang, Jiahao Li, Haijun Wu, Xiaolong Luo, Huiyuan Chen, Yu Wang, Ying Chen, Zhenting Wang, Kai Mei, Yang Zhou, Dimitris N. Metaxas
cs.CL
large language model
大语言模型相关
Abstract
Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. However, existing MAD frameworks fail to beat strong Single Agent and Consistency-based baselines under the same strict cost limit, which shakes the foundation of the MAD field. We propose Conditional Progressive Pruning (CPP), a lightweight pruning framework that fully exploits multi-round MAD. CPP outperforms all existing MAD frameworks on multiple dominated benchmarks. It is also the first to fully outperform consistency methods. Our code, detailed agent interaction records will be released soon.
Chinese Translation
基于大语言模型(LLM)的多智能体辩论(MAD)是最有效的测试时扩展技术之一。通过多轮通信,智能体在知识和推理上相互补充,并解决任何单个成员都无法解决的任务。然而,在同样严格的成本限制下,现有的 MAD 框架未能击败强大的单智能体和基于一致性的基线,这动摇了 MAD 领域的基础。我们提出条件渐进剪枝(CPP),这是一种轻量级剪枝框架,能够充分利用多轮 MAD。CPP 在多个主导基准上优于所有现有 MAD 框架。它也是首个全面超越一致性方法的方法。我们的代码以及详细的智能体交互记录将很快发布。
cs.CL / 62 / 2609.33983
High-Level Text Preprocessing for Semantic Similarity Analysis of Discursive Texts: A Framework and Empirical Demonstration
面向论述性文本语义相似度分析的高层文本预处理:一个框架与实证演示
Mehmet Murat Albayrakoglu, Mehmet Nafiz Aydin
cs.CL · cs.IR · cs.LG
diffusion
扩散模型相关
Abstract
Semantic Textual Similarity (STS) methods assume that a document's lexical content faithfully represents what it asserts. This assumption fails for discursive documents that discuss, compare, critique, and contextualize other positions in the process of articulating their own. The result is semantic diffusion: similarity scores between documents are inflated by vocabulary acquired through discursive engagement rather than substantive alignment. Standard Natural Language Processing (NLP) preprocessing (tokenization, stopword removal, stemming, lemmatization) cannot address this problem because it operates at the lexical level, treating all content identically regardless of its discursive function. This paper introduces high-level text preprocessing: a systematic, rule-based intervention applied before the standard preprocessing pipeline to isolate each document's actual claim from its discursive structure. We propose 12 rules, each with an explicit rationale, and demonstrate their effect on an encyclopedic philosophical corpus: three entries from the Stanford Encyclopedia of Philosophy (virtue ethics, deontological ethics, and consequentialism). A three-phase experiment using eight Transformer-based STS models shows that preprocessing reduces centroid cosine similarity scores across all three theory pairs, with 23 of 24 model-pair comparisons showing the expected decrease and cross-model agreement ranging from 7-1 to 8-0. We introduce the semantic diffusion index (SDI), a per-document metric for assessing the semantic reorientation between a document's raw and high-level preprocessed representations. Although the framework is demonstrated using philosophical texts, it potentially addresses a domain-agnostic problem applicable to legal texts, policy documents, academic articles, and any genre in which a discursive approach introduces vocabulary from positions the document does not endorse.
Chinese Translation
语义文本相似度(STS)方法假设一个文档的词汇内容忠实地表示其所断言的内容。对于在阐述自身立场的过程中讨论、比较、批判并语境化其他立场的论述性文档,这一假设不成立。其结果是语义扩散:文档之间的相似度分数被通过论述性参与而获得的词汇所抬高,而非被实质性对齐所抬高。标准的自然语言处理(NLP)预处理(分词、停用词移除、词干提取、词形还原)无法解决这个问题,因为它在词汇层面运作,不论内容具有何种论述功能都将其一视同仁。本文引入高层文本预处理:一种在标准预处理流程之前应用的系统化、基于规则的干预,用以将每个文档的实际主张与其论述结构隔离开来。我们提出 12 条规则,每条规则都有明确的理由,并在一个百科全书式哲学语料库上展示其效果:来自《斯坦福哲学百科全书》的三个条目(德性伦理学、义务论伦理学和后果主义)。一项使用八个基于 Transformer 的 STS 模型的三阶段实验表明,预处理降低了全部三个理论对上的质心余弦相似度分数,其中 24 个模型配对比较中有 23 个显示出预期下降,跨模型一致性从 7-1 到 8-0 不等。我们引入语义扩散指数(SDI),这是一种逐文档指标,用于评估一个文档的原始表示与高层预处理后表示之间的语义重定向。尽管该框架是使用哲学文本进行演示的,但它可能解决一个领域无关的问题,该问题适用于法律文本、政策文档、学术文章,以及任何这样的体裁:在其中,论述性方法引入了来自文档并不认同的立场的词汇。
cs.CL / 63 / 2609.33989
RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
RewardExplainer:从反事实偏好反馈中学习奖励模型解释
Jingyi He, Nier Wu, Shuang Liu, Xin Wang, Mengnan Du, Xia Hu
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs' feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.
Chinese Translation
奖励模型(RMs)是大语言模型后训练的关键组成部分,为后续的强化学习提供奖励信号。然而,传统的判别式奖励模型通常只输出标量分数,使得难以识别与其评分决策相关的响应行为。现有的解释方法通常依赖预定义的高层属性,并且需要对每个响应对反复进行反事实干预以验证候选解释,缺乏一种利用奖励模型反馈来训练可复用解释器的闭环机制。为了解决这一问题,我们提出 RewardExplainer,一个通过反事实重写从目标奖励模型获取反馈,并利用该反馈进一步优化解释器的框架。RewardExplainer 生成开放式的、原子的且可干预的自然语言评分机制,使解释更具体、更可读且更可操作。它进一步将反事实反馈转换为偏好监督,使解释器能够比单次生成更忠实地捕捉目标奖励模型的评分偏好和敏感行为。在多个目标奖励模型和解释器主干网络上的大量实验表明了一致的改进。除了解释之外,我们使用生成的机制来识别潜在的偏差模式,并构建有针对性的去偏数据以微调奖励模型,从而提高在奖励攻击(reward-hacking)基准上的鲁棒性。
cs.CL / 64 / 2609.34033
Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation
忠实激活言语化:减少大语言模型表示解释中的幻觉
Haiyan Zhao, Zirui Hei, Wei Shi, Huiqi Deng, Na Zou, Mengnan Du
cs.CL
large language model
大语言模型相关
Abstract
Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.
Chinese Translation
诸如 Activation Oracle 和 Natural Language Autoencoders 之类的激活言语化方法将大语言模型的隐藏表示解码为人类可读的自然语言。然而,现有方法可能产生不完整或带有幻觉的描述,使其激活言语化结果在实践中难以被信任和可靠使用。为此,我们提出 AVPO,一个两阶段框架:它首先从隐藏激活中重建源文本,然后使用一个独立的冻结问答模型评估所得文本,从而产生显式且可检查的中间读出。我们进一步使用直接偏好优化(DPO)来优化反演器,所用奖励同时捕捉语义可恢复性和词汇保真度。在六个文本族上,AVPO 相比最强基线在要旨级和细节级信息恢复上分别最多提升 17.1 和 9.3 个百分点。至关重要的是,这些增益来源于偏好优化,而非仅对选定重建进行微调,使紧凑的跨模型反演器能够超越与供体匹配的问题条件言语化器,同时提升语义可恢复性和词汇保真度。此外,分布外案例研究表明,AVPO 能更好地恢复高层语义,同时虚构更少细节。
cs.CL / 65 / 2609.34065
Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
谁获得一个词元,而它承载什么?大型语言模型中不平等的姓名支持与概念访问
Mir Tafseer Nayeem, Davood Rafiei
cs.CL · cs.AI · cs.CY · cs.ET · cs.LG
large language model
大语言模型相关
Abstract
Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched names are not necessarily matched inputs. Some names receive direct single-token access, while others are assembled from multiple subwords, creating unequal name-surface support. Across nearly half a million first names and 12 LLM-associated tokenizers, direct lexical access is highly selective, model dependent, and uneven across race- and gender-associated name metadata. We introduce NameTrace, a model-native, fine-grained, pre-behavioral framework for measuring whether unequal name-surface support remains a vocabulary property or becomes visible in task-relevant internal representations. NameTrace measures concept accessibility from the model's own probabilities over task-specific adjective axes with continuous task-aligned weights. On matched atomic and short-fragmented names within the same race/ethnicity--gender-associated strata, support predicts systematic differences in concept accessibility across fellowship, hiring, clinical assessment, and lending. These differences persist across all eight matched strata, extend across model families, and transfer to unseen names. Hidden-state interventions further show that the measured task directions have downstream leverage, shifting later constrained choices. Unequal lexical support is therefore demographically structured at the input and remains visible in task-relevant model computation. NameTrace makes lexical comparability measurable, supporting a broader principle: behavioral comparability begins with lexical comparability.
Chinese Translation
姓名是个人标识符,但它们也承载社会意义,并被广泛用于评估语言模型如何对待不同的人。此类评估通常假设,匹配的姓名是可比的模型输入。我们表明,这一假设在词汇接口处常常不成立:匹配的姓名未必是匹配的输入。有些姓名获得直接的单词元访问,而另一些姓名则由多个子词组装而成,从而造成不平等的姓名表层支持。在近五十万个名字和 12 个与 LLM 相关的分词器上,直接词汇访问具有高度选择性、依赖于模型,并且在种族和性别相关的姓名元数据之间不均衡。我们提出 NameTrace,一个模型原生、细粒度、行为前的框架,用于衡量不平等的姓名表层支持究竟是仍作为一种词表属性存在,还是在任务相关的内部表征中变得可见。NameTrace 从模型自身在任务特定形容词轴上的概率出发,以连续的任务对齐权重来度量概念可及性。在相同的种族/族裔—性别关联分层内,对于匹配的原子化姓名与短片段化姓名,支持度预测了奖学金、招聘、临床评估和贷款中概念可及性的系统性差异。这些差异在全部八个匹配分层中持续存在,跨越模型家族延伸,并迁移到未见过的姓名。隐藏状态干预进一步表明,所测得的任务方向具有下游影响力,会改变随后受约束的选择。因此,不平等的词汇支持在输入层面具有人口统计学结构,并且在任务相关的模型计算中仍然可见。NameTrace 使词汇可比性变得可度量,从而支持一个更广泛的原则:行为可比性始于词汇可比性。
cs.CL / 66 / 2609.34112
Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
未知不等于正常:将语言模型抽取与基于规则的决策逻辑分离以用于临床风险评分
Nicolás Vera Zúñiga
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision logic (deterministic code computing score bounds over unknown inputs) lets a system ask only questions that can change the decision. On 1,200 synthetic emergency cases across six calculators (HEART, CURB-65, qSOFA, PERC, Wells, Cockcroft-Gault), with a simulated clinician answering questions, we compared this bounds policy with asking for every missing input, a missing-equals-normal schema, and an end-to-end LLM agent (Claude Opus 5.5). With Claude Haiku 4.5 as extractor, the bounds policy matched ask-all accuracy (99.4% vs 99.4%) with half the questions (0.92 vs 1.78 per case) and no irrelevant ones. Treating missing as normal dropped accuracy to 91.2% and under-triaged 8.5% of patients (95% CI 7.1-10.2), and under-triage persisted under messy notes and a noisy clinician. The agent was equally accurate under ideal conditions (99.6%) but 9.5% of its questions were irrelevant; with a noisy clinician it was less accurate than the bounds policy (83.5% vs 87.0%, p<0.001) and committed prematurely in 2.7% of cases (bounds: 0%). A 9B local model as extractor reached oracle-level accuracy (99.8%). In 584 real case reports from MedCalc-Bench, only 52% contained enough information to determine the category (HEART 13%). Routing decisions through code that reasons explicitly about unknowns avoids premature commitment and irrelevant questions, halves the questions asked, and works with small local models.
Chinese Translation
大型语言模型(LLM)越来越多地被用于从自由文本病历记录中计算临床风险评分。病历记录往往不完整,而将未记录的发现视为正常,可能会在无声无息中导致患者被错误分类。我们检验:将三态抽取(由LLM判断为存在、不存在或未知)与决策逻辑(在未知输入上计算评分上下界的确定性代码)相分离,是否能让系统只提出那些可能改变决策的问题。在覆盖六种计算器(HEART、CURB-65、qSOFA、PERC、Wells、Cockcroft-Gault)的1,200个合成急诊病例上,由模拟临床医生回答问题,我们将这种界值策略与以下做法进行比较:对每一个缺失输入都进行询问、缺失等同于正常的模式,以及端到端的LLM智能体(Claude Opus 5.5)。以Claude Haiku 4.5作为抽取器时,界值策略达到了与全部询问相当的准确率(99.4% vs 99.4%),而提问数量只有一半(每例0.92个 vs 1.78个),且没有无关问题。将缺失视为正常使准确率降至91.2%,并使8.5%的患者分诊不足(95% CI 7.1-10.2),而且在杂乱病历记录和有噪声的临床医生条件下,分诊不足现象依然存在。该智能体在理想条件下准确率相当(99.6%),但其9.5%的问题为无关问题;在有噪声的临床医生条件下,其准确率低于界值策略(83.5% vs 87.0%, p<0.001),并在2.7%的病例中过早做出结论(界值策略:0%)。使用一个9B的本地模型作为抽取器达到了oracle水平的准确率(99.8%)。在来自MedCalc-Bench的584份真实病例报告中,只有52%包含足以确定类别的信息(HEART为13%)。将决策经由对未知量进行显式推理的代码来路由,可避免过早做出结论和无关提问,使提问数量减半,并且适用于小型本地模型。
cs.CL / 67 / 2609.34125
Understanding Clinical Cognitive Dialogues Using Large Language Models
使用大型语言模型理解临床认知对话
Vishalakshi Arumugam, Dan Schumacher, Veronica Rammouz, Enrique Gonzalez Guerrero, Jeremy Davis, Anthony Rios
cs.CL
large language model
大语言模型相关
Abstract
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.
Chinese Translation
面对面认知评估既是一项测试,也是一种互动。临床医生解释任务、修复误解,并适应患者的反应,而患者可能会犹豫、寻求澄清或脱离互动。然而,临床对话资源很少标注大规模研究这些行为所需的互动结构。我们呈现一个去标识化的语料库,包含33段认知评估对话、8,250个话语,并标注了三种说话人角色和56种对话行为。我们使用该语料库对大型语言模型进行基准测试,任务包括细粒度对话行为分类和下一患者话语生成。我们还测试了域外指令数据和解释增强训练是否能迁移到这一临床环境。指令微调产生了最强的患者话语参考匹配,并提高了分类准确率。在LLaMA-3.1-8B变体中,推理感知微调产生了最强的分类结果。然而,即使是最好的模型也难以区分密切相关的对话行为,这表明宽泛的会话意图比细粒度的交际功能更容易识别。该语料库和基准使认知评估中的互动结构变得可测量,并支持关于会话标记、临床医生教育以及经过仔细验证的模拟患者的后续工作。这项工作不作出诊断性声明。相反,它提供了研究这些应用所需的数据和评估框架。
cs.CL / 68 / 2609.34158
Toward a Graded Measure of Belief Stability in Large Language Models
迈向大型语言模型中信念稳定性的分级度量
Samantha Dies, Branden Fitelson, Tina Eliassi-Rad
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.
Chinese Translation
大型语言模型(LLMs)日益中介着人们获取信息并利用信息进行推理的方式,然而事实可靠性通常是以一次一个判断的方式被评估的。我们引入分级信念稳定性,这是一种关系性度量,用于衡量一个信念在 LLM 更广泛的信念体系中持续存在的程度。不同于个体信念概率,它询问的是:当某个主张与模型的其他认知承诺一同被考虑时,对该主张的支持是否仍然持续存在。我们用一个直接条件估计器(Direct Conditional estimator)来操作化这一想法,该估计器利用模型内部表示来估计条件信念概率。在 12 个 LLM 和三个领域中,在按个体信念概率进行匹配之后,较低稳定性的信念在 83.3% 的模型-领域设置中,在对话挑战下表现出更大的平均行为变动。因此,分级信念稳定性将可靠性评估从 LLM 对一个主张的支持有多强,扩展到该信念在其更广泛的信念体系中得到支持有多稳健。
cs.CL / 69 / 2609.34187
LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks
LLMs 不是随机鹦鹉:来自类人造语言任务的意义中介的抽象的证据
Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.
Chinese Translation
随机鹦鹉论证的强版本声称,尽管大语言模型(LLMs)可能超越机械复述,但它们无法超越统计模式匹配而进入抽象或推理,尽管能生成极具吸引力的流畅文本,仍在本体论上接近模式复用的下界。我们使用类人造语言任务检验这一假设。若干大语言模型仅被给予对虚构语言的自然语言描述,这些虚构语言通过组合统计上不常见且未获证实的特征,颠覆了训练数据中显著的表面模式。至关重要的是,没有给出任何示例输出。我们认为,如果模型表现出遵循规则的行为,它们便不可能仅依赖表面的统计模式;此类模式常常不利于正确输出。相反,成功的表现需要对提示中指定的约束的表征。在三个互补的任务族中,模型系统性地朝着由意义预测的方向变化:它们区分提示暴露与受指令使用,响应新约束而改变语义关系,并且有时与复杂的翻译答案键精确匹配。尽管性能在所使用的模型谱系中各不相同,这些结果为 LLMs 中的意义中介的抽象提供了证据,并驳斥了强随机鹦鹉假设。我们的工作表明,在适当的架构和上下文约束下,统计学习能够产生意义中介的抽象,尽管生成仍受到表面合理性的强烈约束。我们讨论了对模型开发的意义,以及对理解日益抽象的表示如何可能从似真文本生成目标中涌现的意义。
cs.CL / 70 / 2609.34247
SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents
SALMONN-duo:面向全双工语音智能体的自适应双系统协同
Wenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Chao Zhang
cs.CL · eess.AS
large language model
大语言模型相关
Abstract
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $τ$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $τ$-Voice with an acceptable increase in the delegation rate.
Chinese Translation
全双工语音大语言模型(LLMs)能够实现低延迟、自然的语音交互。然而,现实世界中的智能体还必须使用工具并执行审慎推理——这些操作的延迟与计算开销各不相同,与实时对话严格的时序要求相冲突。为调和这些需求,我们提出了 SALMONN-duo,一种受认知双过程理论启发的自适应双系统语音智能体。SALMONN-duo 通过将一个始终在线的、快思考的全双工语音 LLM(系统 1)与一个强大的异步慢思考 LLM 智能体(系统 2)配对,将实时交互与审慎计算分离开来。除了处理实时交互之外,系统 1 还学习何时直接作答、何时委派,在后端执行期间保持响应能力,并将返回的信息无缝整合到正在进行的对话中,既不暴露工具痕迹,也不丢失对话上下文。在单轮口语问答(QA)和多轮对话上的评估表明,自适应委派显著提升了知识密集型问题和多跳推理问题的准确率,而知识边界感知训练则避免了不必要的系统 2 调用。在定制版的 $τ$-Voice 上,SALMONN-duo 进一步展示了其通过多轮交互在真实业务场景中完成基于环境、受策略约束任务的能力。最后,成本感知的强化学习进一步改善了问答与对话任务中任务性能与后端使用量之间的权衡,同时在 $τ$-Voice 上提升了任务成功率与回复安全性,且委派率的上升处于可接受范围内。
cs.CL / 71 / 2609.34345
CRISP: Cultural Reward Modeling for Implicit Situated Propriety
CRISP:面向隐式情境得体性的文化奖励建模
Zekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao, Zekun Zhou, Ziming Li, Qichen Hong, Kun Chen, Xiaocheng Feng
cs.CL
large language model
大语言模型相关
Abstract
As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.
Chinese Translation
随着大语言模型(LLMs)越来越多地部署到不同国家和地区,识别并恰当回应多样化文化语境的能力变得日益重要。然而,现有研究在很大程度上聚焦于文化知识或具有预定义回应空间的任务,而开放式的文化情境化行为则相对而言仍未得到充分探索。在这项工作中,我们提出 CRISP-RM,一种文化情境化奖励模型,它根据开放式社会场景中的文化适当性来分配奖励。在策略优化过程中,我们进一步引入规范锚定监督(Norm Grounding Supervision, NGS),提供增强策略对相关文化规范敏感性的指导。为了构建文化情境化数据,我们采用一个协作式多智能体框架,该框架将隐式文化规范实例化到多样化的社会场景中,并进一步策划 NormCompass 作为专用测试平台。我们开展全面实验,以评估 CRISP-RM 在奖励建模和策略优化两方面的有效性。Best-of-\(N\) 实验表明,CRISP-RM 始终优于强大的通用奖励模型。在 GRPO 策略优化过程中,CRISP-RM 总体上改善了文化情境化行为,而引入 NGS 则带来进一步提升。进一步分析表明,CRISP-RM 在区分文化上适当的行为方面具有优势,这种优势超越了表面的流畅性和礼貌性,而 NGS 则通过改善规范锚定在策略优化过程中提供互补性收益。
cs.CL / 72 / 2609.34388
Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier
互惠引导:编排草稿与验证预算以推进扩散-自回归自推测前沿
Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li
cs.CL
diffusion
扩散模型相关
Abstract
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to $1.8\times$ speedup.
Chinese Translation
结合自回归(AR)验证的扩散草稿已成为高效推测解码的一种有前景的范式。近期以 Nemotron-Labs-Diffusion 为代表的自推测模型,通过在共享骨干网络内统一草稿与验证,进一步简化了推测流水线,同时实现了更长的接受长度。然而,总体吞吐量与单请求吞吐量之间的帕累托前沿仍未得到充分探索。在低并发下,顺序的草稿-验证执行每轮需要两次模型前向传播,限制了每次前向的有效 token 数(TPF)。相比之下,在高并发下,更长的草稿会带来日益昂贵的计算,迫使单个请求在受限的推测预算下运行,并阻碍了对全骨干草稿器的充分利用。我们的关键观察表明,草稿与验证表现出相互可预测性。草稿 logits 能够预判可能出现的验证不匹配,而近期的验证结果则能预测未来的草稿效用与合适的块大小。基于这一观察,我们提出了互惠引导(RecGuide),一个运行时的草稿-验证编排框架,使推测解码能够适应不断变化的服务负载。RecGuide 在低并发下通过验证重叠的草稿来利用空闲的计算能力,同时在工作负载日益计算密集时动态分配特定于请求的草稿块大小。在广泛并发水平上的实验表明,相较于原始自推测,吞吐量获得了一致的提升,最高实现 $1.8 imes$ 的加速。
cs.CL / 73 / 2609.34428
AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
AgentHop:面向智能体式多跳科学问答的诊断性基准
Chanhee Park, Jeongho Yoon, Sungbin Han, Hyeonseok Moon, Heuiseok Lim
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.
Chinese Translation
智能体式任务要求大语言模型与世界交互,在受限资源下跨多个步骤导航信息并收集证据。由于这种复杂性,智能体式任务的失败源于多种来源,而精确定位这些失败原因对于诊断和改进智能体系统至关重要。然而,现有基准往往只关注单一的排行榜分数,使底层的失败模式不透明。为填补这一空白,我们提出了 AgentHop,这是一个包含 1,011 道多选题的诊断性基准,并配备了一个在固定 token、轮次和工具调用约束下的受控七工具沙箱。AgentHop 通过沿智能体运行的四个轴——检索、综合、工具调用和资源管理——分解单一的准确率分数,从而揭示模型的脆弱性。在 19 个模型上,我们发现行为按模型家族聚集,工具调用特征揭示了不同的家族指纹:GPT 模型早早提交,Anthropic 和 GLM 的检查点在提交前进行验证,DeepSeek 和 Kimi 过度搜索,而 Gemini-3 Pro 保持平衡。分解后的各轴进一步暴露了家族内部结构:Claude Opus 4.6 和 Sonnet 4.6 的准确率相差在一个百分点以内,但在检索与综合的侧重上出现分化,Opus 检索更多,而 Sonnet 综合得更好。我们发布了完整的基准数据集和测试框架,以支持诊断性智能体基准评测。
cs.CL / 74 / 2609.34447
Unbiased Top-$k$ Estimation for On-Policy Distillation
用于同策略蒸馏的无偏 Top-$k$ 估计
Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang, Ziyang Ding, Ditang Gou, Yiming Wu, Zhen Zhao
cs.CL · cs.LG · stat.ML
large language model
大语言模型相关
Abstract
On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated by the student's policy. However, estimating the gradient of the reverse KL divergence in OPD remains a challenge. Using only the sampled token from the student-generated rollout is computationally cheap but provides limited distributional supervision, which will degrade accuracy. In addition, using the full vocabulary provides complete distributional supervision but is computationally expensive. Therefore, recent works propose Top-$k$ OPD (TK-OPD) that use selected top-$k$ tokens, which provides richer distributional supervision than sampled-token estimation at substantially lower computational cost than full-vocabulary estimation. Unfortunately, using only the selected top-$k$ tokens induces bias, leading to accuracy degradation, as the probability mass outside the selected top-$k$ tokens is discarded. To address the bias of TK-OPD, we propose Tail-Corrected Top-$k$ On-Policy Distillation (TT-OPD). It preserves the advantages of TK-OPD, including rich distributional supervision and low computational cost, while providing an unbiased estimator of the gradient of the reverse KL divergence. The key insight of TT-OPD is to use not only the selected top-$k$ tokens, but also the sampled token from the student-generated rollout, thereby recovering the discarded probability mass in expectation, avoiding the bias. Experimental results demonstrate that TT-OPD significantly outperforms other tested OPD variants.
Chinese Translation
同策略蒸馏(OPD)正在成为大语言模型(LLM)后训练的一个重要组成部分,用于将强教师 LLM 的推理能力迁移到较弱的学生 LLM。OPD 通过由学生策略生成的 rollout,最小化教师与学生之间的反向 KL 散度来训练学生。然而,在 OPD 中估计反向 KL 散度的梯度仍然是一个挑战。仅使用学生生成的 rollout 中采样得到的 token 在计算上成本低廉,但提供的分布监督有限,这会降低准确率。此外,使用完整词表可提供完整的分布监督,但计算成本高昂。因此,近期工作提出了 Top-$k$ OPD(TK-OPD),其使用选出的 top-$k$ token,相比采样 token 估计提供了更丰富的分布监督,同时计算成本显著低于全词表估计。不幸的是,仅使用选出的 top-$k$ token 会引入偏差,导致准确率下降,因为选出的 top-$k$ token 之外的概率质量被丢弃了。为了解决 TK-OPD 的偏差,我们提出了尾部校正 Top-$k$ 同策略蒸馏(TT-OPD)。它保留了 TK-OPD 的优势,包括丰富的分布监督和低计算成本,同时提供了反向 KL 散度梯度的无偏估计量。TT-OPD 的关键洞见在于,不仅使用选出的 top-$k$ token,还使用学生生成的 rollout 中采样得到的 token,从而在期望上恢复被丢弃的概率质量,避免偏差。实验结果表明,TT-OPD 显著优于其他测试的 OPD 变体。
cs.CL / 75 / 2609.34454
When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation
当言语显得不足时:用于LLM置信度估计的言语化推理与隐藏特征之间的迭代协同
Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma, Yang Liu
cs.CL
large language model
大语言模型相关
Abstract
Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.
Chinese Translation
置信度估计对于开发值得信赖的大语言模型(LLMs)至关重要,大多数方法遵循基于估计器或基于言语化的范式。尽管近期研究越来越关注改进言语化的置信度自我报告,但我们挑战了这一流行观点:即该方法优于独立的置信度估计器。我们的实证研究表明,一个专用的置信度估计器可以显著优于言语化置信度,这表明LLMs的内部表征包含更丰富的置信度信号。基于这一发现,我们提出迭代策略-估计器训练(Iterative Policy-Estimator Training, IPoET),这是一个协同利用言语化推理轨迹与信息丰富的表征的互补优势的框架。IPoET将策略优化与估计器更新交替进行,将估计器导出的置信度反馈整合到策略学习中,并在新的策略 rollout 上刷新估计器。在多样化数据集以及Qwen和Llama骨干模型上的实验表明,通过迭代地利用更丰富的隐藏特征并适应不断演化的策略分布,IPoET在域内持续优于基于估计器和基于言语化的基线,并在所有域外指标上取得更优或相当的结果。更多细节请参考 https://github.com/xyk829/ipoet。
cs.CL / 76 / 2609.34660
Rewarding Novel Deductions: Solver-guided Process Rewards for Logical Reasoning
奖励新颖演绎:面向逻辑推理的求解器引导过程奖励
Muhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza
cs.CL
large language model
大语言模型相关
Abstract
Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING: (Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation). SPRING uses SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. It introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, it designs process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation across three logical reasoning benchmarks, ZebraLogic, AR-LSAT, and Knights and Knaves, and four LLMs shows that SPRING consistently outperforms base LLMs, outcome-only reward baselines, and Logic-LM. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and strongest outcome-only baseline, respectively. On AR-LSAT, it improves overall accuracy by up to 64.93 and 12.14 points, respectively. On Knights and Knaves, SPRING achieves up to 93.14 puzzle accuracy and 96.05 person accuracy.
Chinese Translation
逻辑推理仍然是大语言模型(LLM)面临的一大挑战,尤其是在需要精确约束追踪、一致性保持和多步演绎的结构化问题上。这一挑战对于小规模 LLM 尤为严峻,因为它们更容易产生不一致、冗余或脆弱的推理轨迹。现有的改进逻辑推理的方法大多针对最终答案的正确性进行优化,仅对中间推理过程提供弱监督。在这项工作中,我们提出了 SPRING:(Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation,求解器引导的新颖逻辑推理步骤生成过程奖励)。SPRING 使用 SMT 求解器作为中间推理步骤的训练时验证器,以提供过程级监督。它引入了新颖推理步骤这一概念,即一个步骤在逻辑上有效、与不断演化的推理状态保持一致,并且尚未被此前已接受的非矛盾演绎所蕴含。基于这种由求解器作出的评估,它设计了过程奖励,以鼓励新颖的推断进展,同时惩罚矛盾且无信息量的推理步骤。在三个逻辑推理基准 ZebraLogic、AR-LSAT 和 Knights and Knaves 以及四个 LLM 上的评估表明,SPRING 持续优于基础 LLM、仅结果奖励基线以及 Logic-LM。在 ZebraLogic 上,SPRING 相较基础 LLM 和最强仅结果奖励基线,将谜题准确率分别提升最多 49.71 和 15.43 个百分点。在 AR-LSAT 上,它分别将总体准确率提升最多 64.93 和 12.14 个百分点。在 Knights and Knaves 上,SPRING 实现了最高 93.14 的谜题准确率和 96.05 的人物准确率。
cs.CL / 77 / 2609.34717
ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation
ReMCTS:面向代码生成的反思增强蒙特卡洛树搜索
Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan
cs.CL
large language model
大语言模型相关
Abstract
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.
Chinese Translation
开放权重的大语言模型(LLMs)能够根据自然语言提示生成函数级程序,但看似合理的候选程序仍会在隐藏语义上失败,并在多次修复尝试中重复犯错。我们提出 ReMCTS,一个以执行为基础、记忆增强、由 LLM 引导的 MCTS 风格搜索框架。它将程序候选组织为树状态,保留分支局部的调试上下文,跨分支检索失败经验,并区分失败的检查与无法获得的证据。在 HumanEval 和 MBPP-Sanitized 上,在留出评估下,使用可见测试的 ReMCTS 在 10 个模型-数据集组合中的 8 个上相较直接生成有所提升,而仅使用代理信号的搜索则稳定性较差。受控的树搜索、采样、修复和记忆消融实验刻画了这些增益的来源与局限。一个包含 30 个任务的 HumanEval-X C++ 试点进一步展示了与编译器支持执行的兼容性,但并不构成广泛的多语言评估。
cs.CL / 78 / 2609.34783
TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases
TQTS-Bench:面向时间序列数据库的文本到查询多语法基准
Fei Lyu, Zhiyi Peng, Jiaming Liu, Yixuan Yang, Changjian Chen, Zhuo Tang, Jiapeng Zhang, Kenli Li
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.
Chinese Translation
大语言模型(LLMs)显著推进了关系数据库上的自然语言查询,然而其查询时间序列数据库(TSDBs)的能力在很大程度上仍未得到评估。现有基准未能充分捕捉 TSDBs 固有的非统一查询语法、多样化的应用领域以及独特的时间特定查询意图。为弥补这一空白,我们提出了 TQTS-BENCH,一个用于评估 TSDBs 上文本到查询能力的多语法基准。TQTS-BENCH 包含 6,125 个高质量问答(QA)对,涵盖 97 个 TSDB、23 种不同的查询语法、22 个应用领域以及 4 种时间特定查询意图。它通过以人为中心的 AI 辅助工作流构建,其中所有 QA 对均由领域专家仔细审查和修订,以确保质量与正确性。对先进 LLMs 和最先进的文本到查询方法的广泛评估揭示了查询 TSDBs 所面临的挑战。即使是评估中表现最好的模型 Claude-Opus-5,其执行准确率也仅达到 48.98%,而人类则达到 87.34%。错误分析表明,这一性能差距主要源于不同 TSDB 之间异构的查询语法、对时间特定意图的误解以及错误的模式链接。这些发现凸显了缩小当前 LLM 能力与现实应用中 TSDB 查询需求之间差距的新机遇。该基准可在以下网址获取:https://anonymous.4open.science/r/TQTS-Bench-00CD。
cs.CL / 79 / 2609.34800
Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks
及格还是不及格?在两项希腊考试基准上评估大语言模型
Panagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis
cs.CL
large language model
大语言模型相关
Abstract
The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in this domain, we introduce Prot-Ex and Pan-Ex, two benchmarks consisting of questions from entrance exams for Greek Model and Experimental schools as well as the Panhellenic exams (the Greek national university entrance examinations). These benchmarks are employed to assess the performance of text-only LLMs-including the Greek-adapted KriKri-8B-Instruct, Llama-3.1-8B, Gemma-4-26B, and Qwen-3-32B-across diverse academic disciplines (Modern Greek, Mathematics, Physics, etc.) and task formats (closed, structured, and open-ended), including textualized visual context (i.e., image descriptions). Our findings indicate the localized KriKri-8B significantly outperforms its base model, successfully rivalling much larger LLMs in linguistically demanding humanities tasks. By leveraging an LLM-as-a-Judge methodology, we expose the inadequacy of traditional lexical metrics for evaluating complex reasoning. Crucially, we uncover a few-shot prompting paradox: while synthetic examples improve accuracy in closed-ended questions, they severely overload the context window of 8B models in structured tasks, causing significant performance degradation. Ultimately, this study suggests targeted linguistic adaptation offsets lower parameter counts in specialized domains, despite the fragility of smaller models to prompt verbosity.
Chinese Translation
大语言模型(LLMs)的快速发展要求对其语言和分析能力以及局限性进行全面评估,尤其是对于像希腊语这样基准覆盖有限的语言。为解决该领域综合性基准可用性有限的问题,我们提出了 Prot-Ex 和 Pan-Ex,这两个基准由希腊示范学校和实验学校入学考试以及泛希腊考试(希腊全国大学入学考试)的题目组成。这些基准被用于评估纯文本LLMs——包括适配希腊语的 KriKri-8B-Instruct、Llama-3.1-8B、Gemma-4-26B 和 Qwen-3-32B——在多样化学科(现代希腊语、数学、物理等)和任务形式(封闭式、结构化和开放式)上的表现,其中包括文本化的视觉语境(即图像描述)。我们的发现表明,经过本地化的 KriKri-8B 显著优于其基础模型,在语言要求高的人文学科任务中成功与规模大得多的LLMs相匹敌。通过利用 LLM-as-a-Judge(以LLM作为评判者)方法,我们揭示了传统词汇指标在评估复杂推理方面的不足。关键的是,我们揭示了一个少样本提示悖论:虽然合成示例提高了封闭式问题的准确率,但在结构化任务中它们严重过载了8B模型的上下文窗口,导致性能显著下降。最终,本研究表明,针对性的语言适配可以抵消专业领域中参数规模较低的劣势,尽管较小的模型面对提示词冗长时具有脆弱性。
cs.CL / 80 / 2609.34967
Semantic Uncertainty Quantification Needs Factual Equivalence
语义不确定性量化需要事实等价性
Joseph Hoche, Quentin Guimard, Gianni Franchi
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compares two answers, and an aggregator that combines all pairwise comparisons into a scalar. Existing methods differ almost entirely in how they aggregate, while taking the operator off the shelf, typically an NLI model or a generic sentence encoder. We show that this reliance on off-the-shelf operators is the primary bottleneck of semantic UQ: they do not accurately measure factual equivalence of multiple answers to the same question. We resolve this with a deliberately simple recipe: a single encoder trained contrastively to isolate the targeted fact, utilizing synthetic data generated by an LLM and dataset both disjoint from all evaluation settings. Integrating the resulting operator into existing methods improves performance on 120 of 126 evaluation settings (95%) spanning 18 model dataset combinations across language and vision-language models. The best variant reaches 0.76 mean AUROC against 0.68 for the strongest baseline, while replacing the quadratic cross-encoder comparisons of entailment-based operators with one encoder pass per answer. The uniformity of the improvement supports the view that the operator, not the aggregator, is the limiting factor. The same operator also improves single generation token-level estimators: the norm it assigns to each token measures how much that token bears on the answer, and reweighting token log-likelihoods accordingly sharpens the estimate.
Chinese Translation
大语言模型的语义不确定性量化基于一个常见模板:采样若干答案,衡量它们之间的一致程度,并将分歧视为不确定性。我们首先将该模板形式化为两个独立角色:一个比较两个答案的算子,以及一个将所有成对比较组合为标量的聚合器。现有方法几乎完全在如何聚合上不同,而算子则直接现成取用,通常是 NLI 模型或通用句子编码器。我们表明,这种对现成算子的依赖是语义 UQ 的主要瓶颈:它们无法准确衡量对同一问题的多个答案的事实等价性。我们用一个刻意简单的方案解决这一问题:单个编码器通过对比学习训练以分离目标事实,所用合成数据由 LLM 生成,且该数据和数据集都与所有评估设置不相交。将所得算子集成到现有方法中,在 126 个评估设置中的 120 个(95%)上提升了性能,这些设置涵盖跨语言模型和视觉语言模型的 18 个模型-数据集组合。最佳变体达到 0.76 的平均 AUROC,而最强基线为 0.68,同时将基于蕴含的算子的二次交叉编码器比较替换为每个答案一次编码器前向传播。改进的一致性支持这样一种观点:限制因素在于算子,而非聚合器。同一算子还改进了单次生成的 token 级估计器:它为每个 token 分配的范数衡量该 token 在多大程度上与答案相关,并据此对 token 对数似然重新加权可提高估计的精度。
cs.CL / 81 / 2609.35070
Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices
行为的回响:道德历史能够塑造并引导大语言模型的行为选择
Lucio La Cava, Andrea Tagarelli
cs.CL · cs.AI · cs.CY
large language model
大语言模型相关
Abstract
Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what extent, moral history shapes LLM decisional behaviors. Prior work on human moral decision-making shows that past behavior can influence subsequent moral choices. Building on this observation, we investigate whether analogous effects emerge in LLMs in two complementary ways: at the behavioral level, through the model's observable responses, and at the representation level, through its latent internal representations. We introduce MoralLedger, a framework for studying how an actor's moral history shapes actions for LLMs' behaviors under a fixed decision context. At the behavioral level, we find that prior moral histories systematically alter subsequent choices as a function of their valence and intensity. At the internal representation level, these histories induce a linearly recoverable direction in the residual stream that generalizes to held-out examples. Intervening along this direction on neutral-history prompts produces two-sided intensity-dependent changes in subsequent choices, with effects that are stronger than those induced by prompting alone or by favorable-nonmoral direction. To our knowledge, this is the first demonstration that a latent representation of an actor's prior moral conduct can provide signed inference-time control over a moral decision. Our MoralLedger extends moral evaluation beyond static dilemmas, establishing moral history as both a source of behavioral sensitivity and a causal target for auditing and controlling moral behavior in LLMs.
Chinese Translation
对大语言模型(LLM)道德性的评估通常孤立地考察各项决策,因而忽视了某个个体无关的先前行为是否会影响模型随后的选择。这使得道德历史是否以及在多大程度上塑造LLM的决策行为这一问题悬而未决。关于人类道德决策的既有研究表明,过去的行为能够影响随后的道德选择。基于这一观察,我们以两种互补的方式探究类似效应是否会在LLM中出现:在行为层面,通过模型可观测的响应;在表征层面,通过其潜在的内部表征。我们提出MoralLedger,一个用于研究在固定决策情境下,行为主体的道德历史如何塑造LLM行为之行动的框架。在行为层面,我们发现先前的道德历史会依据其效价与强度,系统性地改变随后的选择。在内部表征层面,这些历史会在残差流中诱导出一个可线性恢复的方向,且该方向能够泛化到留出样本上。沿这一方向对中性历史提示进行干预,会使随后的选择产生双向的、依赖于强度的变化,其效应强于仅通过提示或通过有利但非道德的方向所诱导的效应。据我们所知,这是首次证明行为主体先前道德行为的潜在表征能够对道德决策提供带符号的推理时控制。我们的MoralLedger将道德评估拓展至静态困境之外,确立道德历史既是行为敏感性的来源,也是审计与控制LLM道德行为的因果靶点。
cs.CL / 82 / 2609.35074
When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment
当置信度过早上升:通过过早答案承诺检测捷径推理
Zhaohan Zhang, Junjie Liu, Chengzhengxu Li, Chen Shen, Xiaoming Liu, Chao Shen, Jieping Ye, Ziquan Liu, Ioannis Patras
cs.CL
large language model
大语言模型相关
Abstract
The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.
Chinese Translation
大型语言模型(LLM)的推理轨迹通常被视为其内部推理的言语化描述。然而,此类轨迹可能是不忠实的:模型可能依赖捷径来得出答案,然后用看似连贯的思维链为其决策进行事后合理化。检测这种捷径推理具有挑战性,因为现有的监控器和验证器主要检查文本轨迹或最终结果,而不是模型在生成过程中对其答案的信念如何发展。我们提出 ConfLens,一个跟踪推理过程中对最终答案的置信度演变的框架。在三种捷径推理设置中,我们观察到一种常见的过早置信模式,即捷径样本在推理早期阶段就对最终答案高度自信。然而,现有的置信度估计方法在检测这种行为时表现出有限的泛化性、可靠性或效率。因此,我们提出分布答案承诺分数(DACS),一种分布置信度估计器,用于衡量模型在每个推理步骤上对答案承诺的概率分布的熵。DACS 捕捉模型答案信念的集中程度,而无需真实答案或任务特定的验证器。我们进一步将 ConfLens 的检测结果转换为奖励模型的可解释信号,以减少其对捷径推理的偏好。在数学和代码推理任务上的实验表明,与强基线相比,带有 DACS 的 ConfLens 将捷径推理检测的 F1 提高了超过 4.3%,并减少了奖励模型偏好中忠实性与正确性之间的不匹配。
cs.CL / 83 / 2609.35108
A mechanistic study of language model introspection
语言模型内省的机制研究
Jiahong Zou, Xiangkun Sun, Lingkai Kong, Tonghan Wang
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.
Chinese Translation
大型语言模型(LLMs)有时能够报告其内部激活所受到的扰动——即使输入没有提供任何表明发生过干预的证据。模型如何检测并定位此类内部变化?我们使用一个保持输入文本固定的受控任务来研究这个问题。我们要么在十个 token 位置之一向隐藏状态注入一个概念向量,要么不施加任何干预。模型被要求识别被扰动的位置,或报告没有发生干预。在三个模型家族中,我们识别出两组小规模的注意力头,它们在内省报告中具有不同的作用。中间层的门控头影响模型是否报告发生变化,而较后一层中的路由头帮助选择要报告的位置。对门控头的干预可以抑制位置报告,即使路由头提供了位置信息。我们进一步考察为什么报告准确率在不同概念之间会变化。能够被更准确地定位的概念向量会在门控头中产生更强的注意力分数和输出响应,这与它们 QK 和 OV 计算中诱导出的键和值变化具有更好的对齐有关。综合来看,这些发现识别出支持内省检测和定位的注意力头机制。
cs.CL / 84 / 2609.35250
SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys
SCBO:面向基于LLM的社会调查的语义连贯批处理与排序
Yuanzi Li, Lingjie Wang, Zihang Tian, Lei Wang, Xu Chen
cs.CL · cs.CY
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits each target to a narrow set of reference responses, and prevents later predictions from using information in earlier answers. Predicting multiple questions in one prompt can reduce these costs, share a broader pool of references, and let later predictions build on earlier ones. This requires forming coherent batches, selecting shared references, and ordering questions and references effectively. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges. SCBO first uses an LLM to extract compact semantic representations from survey items and filter out template noise. It then groups related questions into batches and builds a shared reference bank using target-specific retrieval and centroid-based completion. Finally, it orders target questions from easy to hard and arranges references according to their semantic alignment with those questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over a non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.
Chinese Translation
大型语言模型(LLMs)提供了一种可扩展的方法,利用人口统计特征和观察到的参考回答来模拟调查受访者。然而,每个提示只预测一个问题的传统方法会重复编码相同的上下文,将每个目标限制在狭窄的参考回答集合中,并阻止后续预测利用先前回答中的信息。在一个提示中预测多个问题可以降低这些成本,共享更广泛的参考池,并让后续预测建立在先前预测之上。这要求形成连贯的批次、选择共享参考,并有效地对问题和参考进行排序。我们提出语义连贯批处理与排序(SCBO),这是一个无需训练的框架,能够应对这些挑战。SCBO首先使用LLM从调查项目中提取紧凑的语义表示,并过滤掉模板噪声。然后,它将相关问题分组为批次,并使用目标特定的检索和基于质心的补全来构建共享参考库。最后,它按从易到难的顺序对目标问题进行排序,并根据参考与这些问题之间的语义对齐程度来安排参考。在四个大规模调查数据集和四个LLM上的实验表明,SCBO大幅降低了token消耗和推理时间,同时相对于非批处理基线总体提高了预测准确率。代码可在 https://anonymous.4open.science/r/SCBO-41D8 获取。
cs.CL / 85 / 2609.35279
Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
测量同质小组LLM辩论中的崩溃与纠正
Xin Li, Mengbing Liu, Chau Yuen
cs.CL
large language model
大语言模型相关
Abstract
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.
Chinese Translation
多智能体大语言模型(LLM)辩论通常通过最终答案是否改善来评估,但变化并不必然是改善:同一场讨论既可能挽救最初错误的多数,也可能摧毁最初正确的多数。标准的最终准确率评估将这两种相反的机制混为一谈。我们引入一种用于多项选择题(MCQ)同质辩论的可审计协议,它将每次运行记录为关于崩溃、纠正、起始和带符号干预效用的转移账本。在6,925场MMLU-Pro辩论上,该协议识别出253次崩溃以及一个平行的纠正账本,后者改变了应如何评判干预。重放实验揭示了核心权衡:在等权重下,一种留一模型外探针门控冻结阻止了29次崩溃,但损失了108次纠正,因此仅凭防止崩溃可能推荐错误的策略。一个紧凑的辩论前8探针筛查是一种分诊信号:其与条件崩溃风险的未调整家族级关联很高(G=7,Spearman rho=0.893,精确双侧p=0.0123),但初始多数准确率是一个接近的对照(rho=0.821;家族偏相关rho=0.767,p=0.0877),因此我们不将其视为经过校准或能力调整的预测。轮级轨迹将许多崩溃定位到第一轮辩论,在那里早期分歧可能先于有害级联和有用恢复出现。我们发布可重放的模式、编码器、审计、成本卡和零API重建脚本,以便未来的模型-脚手架行可以在相同的分母和带符号效用账本下进行比较。
cs.CL / 86 / 2609.35308
Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation
多轮大语言模型污染中的认知策略分歧:一项协议梯度研究
Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah
cs.CL
large language model
大语言模型相关
Abstract
Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct failure mechanisms while holding the false premise constant, and evaluate GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains at temperature zero (22,500 turns), using a dual-track automated judge validated against a human gold standard (Cohen's \k{appa} = 0.901). GPT-5.4 Mini showed zero adoptions across all 500 sessions, a content-independent policy at the session level; token-level probing shows the underlying margin, while large, is finite. Gemini-3.1 Flash-Lite followed a steep authority gradient: 0.1% adoption for self-attributed falsehoods, 23.5% for user-cited sources, 68.2% for system-injected authority, and 94.0% under instruction override. GLM-4.5-Air showed a shallower gradient (15.8% vs 84.2%), a 68-percentage-point dissociation confirming that authority deference and instruction compliance are distinct mechanisms within one architecture. Recovery also diverged: GLM recovered in 94.5% of affected sessions, whereas 26.1% of affected Gemini sessions never did, rising to 40.0% under instruction override. Conversation history is an untrusted attack surface requiring provenance-aware system design; the complete framework is released as an open-source benchmark.
Chinese Translation
大语言模型将对话历史作为未经核实的上下文来处理:注入到先前轮次中的虚假前提会被当作事实采纳,我们将这一失效模式称为会话级污染。我们引入了五种污染协议,它们沿来源权威梯度排列,在保持虚假前提不变的同时分离出不同的失效机制,并在温度为零的设置下(22,500 个轮次)、跨十个知识领域评估了 GPT-5.4 Mini、Gemini-3.1 Flash-Lite 和 GLM-4.5-Air,所用的是一个双轨自动评判器,并针对人工金标准进行了验证(Cohen's \k{appa} = 0.901)。GPT-5.4 Mini 在所有 500 个会话中均未出现采纳,在会话层面上表现为一种与内容无关的策略;token 级探测显示其底层裕度虽然很大,但仍是有限的。Gemini-3.1 Flash-Lite 遵循陡峭的权威梯度:自我归因的虚假信息为 0.1% 的采纳率,用户引用的来源为 23.5%,系统注入的权威为 68.2%,而在指令覆盖下为 94.0%。GLM-4.5-Air 表现出更平缓的梯度(15.8% 对 84.2%),这一 68 个百分点的分离证实了权威顺从与指令遵循是同一架构内的不同机制。恢复情况同样出现分化:GLM 在 94.5% 的受影响会话中实现了恢复,而受影响的 Gemini 会话中有 26.1% 从未恢复,在指令覆盖下这一比例升至 40.0%。对话历史是一个不可信的攻击面,需要具备来源感知能力的系统设计;完整框架已作为开源基准发布。
cs.CL / 87 / 2609.35312
MemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs
MemoReason:评估参数记忆对大型语言模型上下文推理的影响
Zineddine Tighidet, Andrea Mogini, Jiali Mei, Patrick Gallinari, Benjamin Piwowarski
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \textit{memorization bias}, where familiar content improves reasoning performance, and the \textit{Strong Parametric Shortcut Hypothesis}, where models skip reasoning entirely and recall stored answers. To test these effects, we introduce \textbf{MemoReason}, a human-curated benchmark that pairs factual reasoning tasks with structurally identical \fictitiousterm{} versions where real entities like people, companies, or dates are systematically replaced by \fictitiousterm{} ones of the same type. This \scorerevision{preserves task structure and specified reasoning operations} while varying the familiarity of the context, allowing controlled measurement of how the parametric memory affects reasoning. \revision{Our evaluation of recent LLMs reveals consistent and statistically significant performance drops of up to 15.7\% in the fictitious setting, demonstrating a clear memorization bias.} However, a targeted analysis of \revision{questions failed in the fictitious setting} shows that models rarely respond with the corresponding factual answer, indicating that direct parametric shortcuts are not the dominant failure mode. These findings suggest that parametric memory influences reasoning through mechanisms more complex than simple factual recall. \textbf{MemoReason} provides a controlled framework for studying these mechanisms and for extending paired factual-fictitious{} evaluation to broader reasoning settings.
Chinese Translation
大型语言模型(LLMs)在推理基准上表现良好,但目前尚不清楚这反映的是真正的上下文推理,还是对其参数中记忆的事实的依赖。我们通过区分两种可能性来研究这一点:一种广泛的 \textit{记忆偏差},即熟悉的内容会提升推理表现;以及 \textit{强参数捷径假设},即模型完全跳过推理并回忆存储的答案。为检验这些效应,我们提出 \textbf{MemoReason},这是一个人工整理的基准,它将事实推理任务与结构相同的 \fictitiousterm{} 版本配对,其中人物、公司或日期等真实实体被系统性地替换为同类型的 \fictitiousterm{} 实体。这种 \scorerevision{保留了任务结构和指定的推理操作},同时改变上下文的熟悉程度,从而允许受控地测量参数记忆如何影响推理。\revision{我们对近期 LLMs 的评估显示,在虚构情境中性能出现了一致且统计上显著的下降,降幅最高达 15.7\%,表明存在明显的记忆偏差。} 然而,对 \revision{在虚构情境中失败的题目} 进行的针对性分析表明,模型很少以相应的事实答案作答,这表明直接的参数捷径并不是主要的失败模式。这些发现表明,参数记忆通过比简单事实回忆更复杂的机制影响推理。\textbf{MemoReason} 提供了一个受控框架,用于研究这些机制,并将配对的事实-虚构{} 评估扩展到更广泛的推理场景中。
cs.CL / 88 / 2609.35376
How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?
LLMs 能在多大程度上模拟真实学习者对教育反馈的评价?
Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara, Kyosuke Takami
cs.CL
large language model
大语言模型相关
Abstract
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.
Chinese Translation
尽管近期研究已探索了使用大语言模型(LLMs)进行人类行为与偏好模拟,但在教育情境中,LLMs 能在多大程度上模拟真实学习者的主观评价仍不清楚。我们利用针对高中生物问题反馈的真实学习者评价数据,在群体和个体两个层面上研究了这一问题。我们在六个模型上比较了使用与不使用学习者特定信息(如人格特质和评价示例)时的表现。我们的结果表明,LLMs 模拟学习者评价的能力仍然有限。提供学习者画像和示例可以改善分数校准和个体层面的模拟,但更常无法提升群体层面的一致性。这些发现凸显了有必要研究哪些学习者信息和适应策略对学习者偏好模拟是有效的。
cs.CL / 89 / 2609.35387
TRACE: Single-Pass Decoding-Trace Risk Localization for Generation Calibration
TRACE:用于生成校准的单遍解码轨迹风险定位
Yuebin Xu, Xuemei Peng, Junlan Chen, Zhiyi Chen, Zeyi Wen
cs.CL
large language model
大语言模型相关
Abstract
Reliable confidence estimation is essential for large language model deployment. However, answer-level calibration remains challenging because generation errors are often localized: a response may be fluent and high-probability overall while still failing at a critical number, entity, or factual claim. Existing estimators compress token probabilities, sequence likelihoods, entropy, or beam statistics into a global score, which can dilute such local risk signals. We propose TRACE, a single-pass, decoded-answer-preserving confidence estimator that treats decoding-time uncertainty as a trajectory through three steps: (i) recording token-level surprisal and predictive entropy during decoding, (ii) applying local risk operators to preserve uncertainty spikes, and (iii) converting localized trace risk into answer-level confidence. TRACE produces a label-free risk score, while TRACE+ calibrates trace-only features into probabilities using a held-out split, without extra generations or external verifiers. We evaluate four tasks against 19 calibration baselines, and TRACE+ reduces Brier from 0.149 to 0.137 and improves AUROC from 0.758 to 0.792 over the strongest likelihood baseline. Across seven LLMs, TRACE+ improves over the best non-TRACE baseline pool from 0.136 to 0.120 Brier and from 0.764 to 0.817 AUROC. Results show that localizing decoding-time risk provides a general approach to calibration.
Chinese Translation
可靠的置信度估计对于大型语言模型部署至关重要。然而,答案级校准仍然具有挑战性,因为生成错误往往是局部化的:一个回答可能整体流畅且概率很高,却仍然在关键数字、实体或事实性主张上出错。现有估计器将词元概率、序列似然、熵或 beam 统计量压缩为一个全局分数,这可能会稀释此类局部风险信号。我们提出 TRACE,一种单遍、保留已解码答案的置信度估计器,它将解码时不确定性视为一条经过三个步骤的轨迹:(i) 在解码过程中记录词元级惊奇度和预测熵,(ii) 应用局部风险算子以保留不确定性尖峰,以及 (iii) 将局部化轨迹风险转换为答案级置信度。TRACE 产生无标签风险分数,而 TRACE+ 使用留出划分将仅轨迹特征校准为概率,无需额外生成或外部验证器。我们在四项任务上针对 19 个校准基线进行评估,并且相较于最强似然基线,TRACE+ 将 Brier 从 0.149 降至 0.137,并将 AUROC 从 0.758 提升至 0.792。在七个 LLM 上,TRACE+ 相较于最佳非 TRACE 基线池,将 Brier 从 0.136 改善至 0.120,并将 AUROC 从 0.764 改善至 0.817。结果表明,定位解码时风险为校准提供了一种通用方法。
cs.CL / 90 / 2609.35461
AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic
AraDynFact:阿拉伯语事实知识的动态评估
Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi
cs.CL
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.
Chinese Translation
随着大型语言模型(LLMs)在规模和能力上的持续扩展,它们在阿拉伯语方面的熟练程度已取得显著进步。然而,一个关键空白依然存在:它们对多元阿拉伯语世界的事实知识范围和文化敏感性在很大程度上仍未得到充分探索。当前的评估指标往往聚焦于翻译或通用推理,未能捕捉阿拉伯文化固有的丰富历史、社会和地区细微差别。此外,大多数基准依赖于繁重的工作,并在某些步骤中需要人工干预,这使得对知识覆盖范围的评估既昂贵又缓慢。为解决这一不足,我们提出 AraDynFact,一个新颖的动态评估框架,旨在严格评估嵌入在 LLMs 中的阿拉伯语事实知识。与静态基准不同,AraDynFact 采用动态方法来提取事实信息,并以快速且自动的方式生成丰富且可回答的问题。我们将 AraDynFact 应用于阿拉伯语维基百科,并审计了若干最先进模型的性能,这些模型从以阿拉伯语为中心的专用 LLMs 到高资源通用 LLMs 不等。此外,我们发现其与现有的、手工构建的以阿拉伯语为中心的基准具有高度相关性,这证实了我们动态方法的潜力。
cs.CL / 91 / 2609.35578
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
FactorEngram:用于语言模型的、具有基级门控的因子化 N-gram 记忆
Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
Chinese Translation
基于查找的记忆一直是扩展大语言模型(LLM)参数的一种有前景的方式。它检索局部词元模式(如 n-gram)的已学习表示,而不是通过逐层计算来重建它们。然而,诸如 Engram 之类的现有设计将每个检索到的嵌入视为一个不可分割的整体单元。每个嵌入被存储在其各自的哈希槽中,并由单个标量门进行调制。因此,多义模式无法选择性地读出其记忆中与上下文相关的那些分量。此外,参数仅通过哈希冲突实现共享,而哈希冲突在很大程度上与语义无关。我们提出 FactorEngram,一种具有基级上下文门控的因子化 n-gram 记忆。FactorEngram 在跨模式共享的基向量字典上检索经稀疏性正则化的系数,因此相关模式可以复用公共分量。同一个字典也被用于门控。主干网络的隐藏状态与每个基向量进行打分,以在重建之前对相应的系数进行门控,这使得上下文能够对每个记忆分量进行单独调制。FactorEngram 还同时涵盖单个词元和多个词元的 n-gram,并且我们系统地研究了记忆分支应当插入的位置。在 340M 和 1B 参数的 Transformer 主干网络上,FactorEngram 提升了语言建模和下游任务的性能。消融研究证实了每个组件的贡献,并确定将记忆分支插入中间层的注意力子层之前是一种有效的配置。
cs.CL / 92 / 2609.35663
Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context
后期注意力层单独即可复制实体词元,但若不对其上下文加以注意则不然
Muyu He, Yuchen Liu, Ran Tao, Li Zhang
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which layers specialize in this fundamental task or how other tokens in the same sequence, termed context tokens, influence the model's ability to copy the entity tokens. To address these questions, we conduct experiments on Qwen3-8B using two novel methods: genie-in-a-bottle, which controls exactly which layers can participate in an entity-copying task, and attention lobotomy, which cuts off specific tokens' attention to entity tokens without affecting the remaining attention distribution. We find that two distinct groups of layers in the second half of the model are both necessary and sufficient for entity copying. Moreover, in addition to the decoding position's attention to entity tokens, context tokens' attention to entity tokens also proves necessary for copying the exact tokens, even though context tokens do not store entity information themselves unless they satisfy particular semantic properties. Our findings establish the critical role of late layers in entity copying under the guidance of context tokens, calling for future work on how models propagate and consume entity information.
Chinese Translation
大型语言模型(LLM)能够可靠地执行实体复制,即模型将指代某一实体的词元(称为实体词元)从提示中复制到其输出中,以回答一个问题。尽管实体复制对大多数 LLM 而言都很直接,但现有研究并未系统地说明哪些层专门负责这一基础任务,也未说明同一序列中的其他词元(称为上下文词元)如何影响模型复制实体词元的能力。为回答这些问题,我们在 Qwen3-8B 上使用两种新方法进行实验:genie-in-a-bottle,它能够精确控制哪些层可以参与实体复制任务;以及 attention lobotomy,它切断特定词元对实体词元的注意力,同时不影响其余的注意力分布。我们发现,模型后半部分中两个不同的层组对于实体复制既是必要的也是充分的。此外,除了解码位置对实体词元的注意力之外,上下文词元对实体词元的注意力也被证明对于复制确切的词元是必要的,即使上下文词元本身并不存储实体信息,除非它们满足特定的语义属性。我们的发现确立了后期层在上下文词元引导下于实体复制中的关键作用,并呼吁未来研究模型如何传播和消费实体信息。
cs.CR / 93 / 2609.33903
Can Prompt Anonymity Protect Your Identity From LLM Providers?
提示匿名能否保护你的身份免受 LLM 提供商侵害?
Dzung Pham, Dillon Sheils, Naina Singh, Amir Houmansadr
cs.CR
large language model
大语言模型相关
Abstract
User conversations with large language models (LLMs) often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerged as a practical solution that separates user identity from their prompts, yet this approach still leaves the prompt content visible to LLM providers. We study the impact of this gap by conducting the first empirical investigation into the risk of prompt authorship re-identification. Towards this end, we create PromptAnonBench, a novel benchmark for evaluating prompt anonymity, consisting of over 175,000 cleaned, authentic multi-turn user prompts from various real-world datasets (SWE-Chat and WildChat). Using the embeddings of historical user conversations, an attacker can correctly detect and re-identify at least one anonymized conversation for 50--75% of SWE-Chat users and up to 10% of WildChat users at a 10% false acceptance rate for out-of-set users, even with text-based defenses applied. Our findings unveil the risk of relying only on anonymity for private LLM inference and the gap in existing text privacy defenses.
Chinese Translation
用户与大型语言模型(LLMs)的对话往往包含高度敏感的个人信息,LLM 提供商可利用这些信息来创建详细的用户档案、实现定向广告投放以及训练更强大的模型。为了保护用户隐私,匿名化 LLM 代理已成为一种实用解决方案,它将用户身份与其提示分离,然而这种方法仍使提示内容对 LLM 提供商可见。我们通过开展首个针对提示作者身份重识别风险的实证研究,来考察这一缺口的影响。为此,我们创建了 PromptAnonBench,这是一个用于评估提示匿名性的新基准,包含来自多个真实世界数据集(SWE-Chat 和 WildChat)的超过 175,000 条经过清洗的、真实的多轮用户提示。使用历史用户对话的嵌入,攻击者可以在对集合外用户保持 10% 误接受率的情况下,为 50--75% 的 SWE-Chat 用户和最多 10% 的 WildChat 用户正确检测并重识别至少一段匿名化对话,即使应用了基于文本的防御措施。我们的发现揭示了仅依赖匿名性进行私密 LLM 推理的风险,以及现有文本隐私防御中的缺口。
cs.CR / 94 / 2609.33914
A2A-CaseVerify: Merkle-Linked Case-Evidence Verification for Cross-Organization A2A Workflows
A2A-CaseVerify:面向跨组织 A2A 工作流的 Merkle 关联案件证据验证
Adil Alshammari, Hayretdin Bahsi
cs.CR
large language model
大语言模型相关
Abstract
Agent-to-Agent (A2A) communication enables large language model (LLM) agents to exchange tasks, messages, and artifacts across organizations. A valid message alone does not establish that a final workflow claim is supported by a complete, ordered, and case-consistent evidence path. We present A2A-CaseVerify, a deterministic offline verifier that maps preserved A2A runtime objects, business events, and typed support edges to a directed case evidence graph. It commits A2A envelopes, event projections, and support edges, then combines event and edge leaves into a deterministic case Merkle root. The verifier returns SUPPORTED with status OK only when protection and profile-specific reconstruction checks pass. We evaluate one canonical supported healthcare-profile bundle and 13 controlled mutations. We separately test a valid bundle generated with the official Python A2A software development kit (SDK). In these tests, the canonical and SDK-generated bundles are accepted. All 13 mutations are rejected, and each returned reason-code set contains its targeted diagnostic code. A2A-CaseVerify adds offline case-level evidence-support verification to A2A interoperability.
Chinese Translation
智能体到智能体(A2A)通信使大语言模型(LLM)智能体能够跨组织交换任务、消息和制品。仅凭一条有效消息并不能确立最终工作流主张得到了一条完整、有序且案件一致的证据路径的支持。我们提出 A2A-CaseVerify,一种确定性离线验证器,它将所保留的 A2A 运行时对象、业务事件和带类型的支持边映射为有向案件证据图。它对 A2A 信封、事件投影和支持边进行承诺,然后将事件叶与边叶合并为一个确定性的案件 Merkle 根。只有当保护检查与特定配置档案的重建检查均通过时,验证器才返回 SUPPORTED 且状态为 OK。我们评估了一个规范的、受支持的医疗健康配置档案捆绑包以及 13 个受控变异。我们另行测试了一个使用官方 Python A2A 软件开发工具包(SDK)生成的有效捆绑包。在这些测试中,规范捆绑包与 SDK 生成的捆绑包均被接受。全部 13 个变异均被拒绝,且每个返回的原因码集合都包含其针对性的诊断码。A2A-CaseVerify 为 A2A 互操作性增添了离线案件级证据支持验证。
cs.CR / 95 / 2609.33924
A2A-ForensicTrace: Offline Verification of Tamper-Evident A2A Runtime Evidence
A2A-ForensicTrace:防篡改 A2A 运行时证据的离线验证
Adil Alshammari, Sareh Assiri, Hayretdin Bahsi
cs.CR
large language model
大语言模型相关
Abstract
Security-relevant Agent2Agent (A2A) executions can cross organizational boundaries, leaving investigators without live access to all participating systems. Offline investigation involves checking preserved records and their cross-record relationships for consistency. This paper presents A2A-ForensicTrace, an offline verification layer that converts runtime observations into typed records, derives protocol-relevant relationships, and commits both under an incident-trace root. An Ed25519-signed receipt binds the root and capture digest to the incident context. Evaluation comprised 240 executions through the official A2A software development kit (SDK), with actions selected by a large language model (LLM). These included 120 condition runs and 120 matched controls. All condition runs returned the expected bounded findings. No control produced an indication, and all roots and receipts verified. Median in-memory latency of the full offline verifier was 1.61 ms for the scenario traces. In the separate scaling experiment, it was 30.85 ms at 1,000 committed leaves. Future work will broaden A2A lifecycle coverage.
Chinese Translation
与安全相关的 Agent2Agent(A2A)执行可能跨越组织边界,使调查人员无法实时访问所有参与系统。离线调查包括检查所保存的记录及其跨记录关系的一致性。本文提出了 A2A-ForensicTrace,这是一种离线验证层,它将运行时观测转换为带类型的记录,推导出与协议相关的关系,并将二者提交于一个事件追踪根(incident-trace root)之下。一个 Ed25519 签名的回执将该根与捕获摘要绑定到事件上下文。评估包括通过官方 A2A 软件开发工具包(SDK)执行的 240 次执行,其中动作由大语言模型(LLM)选择。这些执行包括 120 次条件运行和 120 次匹配对照。所有条件运行都返回了预期的有界发现结果。没有任何对照产生指示,并且所有根和回执均验证通过。对于场景追踪,完整离线验证器的内存中延迟中位数为 1.61 ms。在单独的扩展实验中,当提交叶子数为 1,000 时,该值为 30.85 ms。未来工作将拓宽 A2A 生命周期覆盖范围。
cs.CR / 96 / 2609.33985
The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning
众包微调的隐私谬误:通过基于主题的投毒提取专有数据
Sae Furukawa, Alina Oprea
cs.CR · cs.LG
large language model
大语言模型相关
Abstract
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches $3.71\times$ the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and $3.08\times$ for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.
Chinese Translation
监督微调(SFT)被广泛用于使大型语言模型适配下游任务。众包用户对话是一种成熟的方法,可在减少对昂贵人工标注需求的同时大规模收集 SFT 数据。然而,它也允许不可信用户向微调流程贡献数据。我们研究了这一场景中一个尚未充分探索的隐私风险:恶意用户能否污染一小部分众包数据,以放大对其他用户贡献的此前未见指令的提取?我们表明,仅利用对已部署模型的黑盒、仅输出访问即可实现这一点。在四个模型和两个数据集上的实验表明,训练数据提取显著增加:仅使用 50 个投毒样本,在 OpenMathInstruct 上对 Qwen2.5-14B,近乎逐字提取达到未投毒时比率的 $3.71\times$;在 AceReason 上对 Llama-3.1-8B 达到 $3.08\times$。数据过滤在检测投毒样本方面也证明大体无效:即使表现最好的方法在 F-1 分数上也仅达到 0.378,使得大多数投毒样本未被检测到。这些发现表明,看似良性的众包贡献能够放大其他记录的泄露,同时仍然难以通过数据过滤被识别。
cs.CR / 97 / 2609.34251
RADNPO: Reference-free Adaptive Negative Preference Optimization for LLM Unlearning
RADNPO:面向LLM遗忘的无参考自适应负偏好优化
Shenghan Tan, Ziyi Zhou, Wenpeng Hu, Mengyuan Zhang
cs.CR
large language model
大语言模型相关
Abstract
Large language models (LLMs) can memorize sensitive, private, or copyrighted content during pre-training, making machine unlearning necessary for removing targeted knowledge. Recent preference optimization (PO)-based unlearning methods improve stability over gradient ascent (GA)-based methods by introducing alignment-style objectives, which effectively suppress the probability of forget targets. However, target suppression alone does not sufficiently constrain the next-token distribution after unlearning. Existing methods provide limited control over how suppressed probability mass is redistributed and insufficiently adapt forgetting strength to target confidence and distributional concentration. Even after target suppression, probability mass may remain concentrated on a few non-target tokens, potentially producing repetitive or uninformative outputs. To address these limitations, we propose Reference-free ADaptive Negative Preference Optimization (RADNPO), which explicitly guides next-token probability redistribution. Specifically, RADNPO contrasts each forget target with alternative tokens favored by the current next-token distribution and adaptively modulates token-level forgetting strength using target confidence and next-token concentration. Experiments on TOFU and MUSE demonstrate that RADNPO achieves a better trade-off between forgetting quality and model utility than current baselines.
Chinese Translation
大型语言模型(LLMs)在预训练期间可能记住敏感、私密或受版权保护的内容,这使得机器遗忘对于移除目标知识变得必要。最近的基于偏好优化(PO)的遗忘方法通过引入对齐式目标,相较于基于梯度上升(GA)的方法提升了稳定性,这些目标有效地抑制了遗忘目标的概率。然而,仅靠目标抑制并不能充分约束遗忘后的下一词元分布。现有方法对被抑制的概率质量如何重新分配提供的控制有限,并且未能充分根据目标置信度和分布集中度自适应调整遗忘强度。即使在目标抑制之后,概率质量仍可能集中在少数非目标词元上,从而可能产生重复或缺乏信息量的输出。为了解决这些局限,我们提出无参考自适应负偏好优化(RADNPO),它显式地引导下一词元概率的重新分配。具体而言,RADNPO将每个遗忘目标与当前下一词元分布所偏好的替代词元进行对比,并利用目标置信度和下一词元集中度自适应地调制词元级遗忘强度。在TOFU和MUSE上的实验表明,RADNPO实现了比当前基线更好的遗忘质量与模型效用之间的权衡。
cs.CR / 98 / 2609.34450
ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch
ReproBench:在从零开始复现漏洞方面对LLM智能体进行基准测试
Liang He, Sheng Wu, Haomiao Hao, Hongduo Zhao, Jia Yan, Purui Su
cs.CR
large language model
大语言模型相关
Abstract
Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting. Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.
Chinese Translation
大语言模型(LLM)智能体在漏洞复现、漏洞利用和漏洞修补等网络安全任务上正受到越来越多的评估。然而,现有的网络安全基准主要在后环境评估范式下运行,即把源代码、容器或可执行二进制文件直接交给智能体。这种设置绕过了关键的环境重建步骤,从而为真实世界的漏洞分析留下了一个根本性问题:智能体能否完全从零开始自主重建所需的执行环境并复现漏洞?为填补这一空白,我们提出了ReproBench,这是一个以证据为基础的基准,旨在评估智能体仅从一个CVE标识符出发进行端到端漏洞复现的能力。ReproBench将完整的复现工作流分解为六个不同的阶段,并使用可验证的实验产物独立评估每个阶段的表现:下载的固件镜像、解包后的二进制文件、细粒度的分析日志以及经过验证的崩溃样本等。我们以30个真实世界的物联网固件漏洞实例化了ReproBench,这些漏洞为我们的从零开始评估设置提供了理想的测试用例。我们的评估表明,45.3%的测试运行求助于漏洞模拟——这是所有被评估的LLM智能体普遍采用的一种补救性变通做法——而只有5.3%的CVE-模型组合成功复现了真实世界的漏洞。尽管总体成功率较低,但这些非平凡的成功案例证实,最先进的LLM智能体已经具备完全自主进行端到端漏洞复现的能力。与此同时,我们对失败案例的深入分析识别出了在整个复现流程中阻碍LLM智能体的核心瓶颈,为后续研究提供了可付诸实践的洞见。
cs.CR / 99 / 2609.34463
CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents
CoDeL:面向基于LLM的智能体中间接提示注入的协同演化防御
Xiao Yang, Yangchen Ou, Yuhan Gao, Le Wang, Zonghao Ying, Aishan Liu
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.
Chinese Translation
基于大语言模型(LLM)的智能体日益依赖外部工具和内容,这使它们暴露于间接提示注入(IPI)之下。这一威胁催生了范围广泛的多种防御方法,其中基于训练的防御通常被认为最为可靠。然而,现有的基于训练的防御通常是在显式注入的静态分布上进行优化的。它们学到的是表层形式线索,而非在服务用户与服从注入目标之间的边界,因此当恶意意图被折叠进一个看似合理的工作流中并被推迟若干轮次时,它们就会失效。我们提出 CoDeL,一种在训练过程中不断重塑攻击分布、从而强化智能体防御能力的防御方法。防御者在每一轮通过基于 LoRA 的 GDPO 进行更新,其奖励在安全性、任务进展和格式合规性上被解耦,因此拒绝注入与完成用户的任务共同定义了适应度。为了持续为其提供值得学习的失败样本,一个协同演化的探测者会在注入轮次、攻击方法和载荷上进行搜索,以寻找仍能穿透当前防御者的注入,并由攻击成功率和攻击时延共同引导,从而优先挖掘那些防御者察觉过晚的突破。每一次防御者的更新都会使部分攻击种群失效,并迫使下一轮走向新的前沿,从而将防御者自身的失败转化为一个不断变化的课程。在三个 IPI 基准、九个基线方法和两个基础模型上的大量实验表明,CoDeL 将攻击成功率(ASR)降低了 88.5%,并大幅优于其他基线(+38.0%)。代码已公开。
cs.CR / 100 / 2609.34514
How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models
如何驯服多头九头蛇?面向大型语言模型的自适应多类别安全引导
Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu
cs.CR · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during inference without updating model parameters. However, a single prompt can involve multiple harm categories, and steering toward safety in one category may leave harmful content from another unaddressed. Despite advances in adaptive steering, existing methods do not explicitly coordinate steering direction and strength when multiple harm categories co-occur within a single prompt. To address this problem, we propose CAM-Steer, a Category-Adaptive Multi-category Safety Steering framework. Specifically, it estimates the risk associated with each harm category by comparing the current hidden state with safe and unsafe prototypes. The estimated risks are then used to combine the safety directions for different harm categories into a single steering direction and to determine the strength of the intervention. Finally, it rotates the hidden state along the composed steering direction, with the rotation angle determined by the estimated risks, while preserving the hidden-state norm. Experiments across three LLM backbones and seven harm categories show that CAM-Steer outperforms the evaluated baselines in average defense success rate, including when categories co-occur. Further analyses support its component designs and informative risk scores, with negligible inference overhead.
Chinese Translation
随着大型语言模型(LLMs)日益普及,防止对有害提示产生不安全响应对于其安全部署至关重要。激活引导提供了一种通过在不更新模型参数的情况下于推理过程中修改内部激活来提升 LLM 安全性的方法。然而,单个提示可能涉及多个危害类别,而朝一个类别安全的引导可能使来自另一类别的有害内容未得到处理。尽管自适应引导取得了进展,当多个危害类别在单个提示中同时出现时,现有方法并未显式协调引导方向和强度。为解决此问题,我们提出 CAM-Steer,一个类别自适应的多类别安全引导框架。具体而言,它通过将当前隐藏状态与安全和不安全原型进行比较,估计与每个危害类别相关的风险。随后,估计的风险被用于将不同危害类别的安全方向组合成单个引导方向,并确定干预强度。最后,它沿组合后的引导方向旋转隐藏状态,旋转角度由估计的风险确定,同时保持隐藏状态范数。在三个 LLM 主干和七个危害类别上的实验表明,CAM-Steer 在平均防御成功率上优于所评估的基线,包括当类别共现时。进一步分析支持其组件设计和有信息量的风险评分,且推理开销可忽略不计。
cs.CR / 101 / 2609.34597
AuxMark: Defending Against Unauthorized Agent Distillation via Auxiliary Behavioral Watermarking
AuxMark:通过辅助行为水印防御未经授权的智能体蒸馏
Yiqing Feng, Haozhe Feng, Shunan Shang, Xiaoyu Zhang, Jian Lou, Haodong Zhao, Mingxun Zhou
cs.CR
large language model
大语言模型相关
Abstract
Large language model agents can acquire complex capabilities through multi-step interaction and tool use, but their trajectories can also be illegally collected to dis- till student agents. However, existing watermarking methods either do not fit the structured and interactive nature of agent environments or lack reliable effective- ness across tasks and model architectures. We introduce AuxMark, a behavioral watermarking framework for tracing unauthorized agent distillation. AuxMark dynamically inserts safe, non-essential auxiliary action into teacher trajectories, and stores the associated contexts as private evidence cards. To audit a suspicious student model, AuxMark constructs paired real and fake probes from these cards and applies a card-level sign test. This black-box protocol supports both model- level detection and trace-level attribution. Across three agent benchmarks, two teacher agents, and four student architectures, AuxMark detects all 24 distilled models with zero false positives on 48 clean models. It also preserves task utility and remains effective against data flooding, paraphrasing, truncation, and adaptive cleaning attacks. Our code will be released at this URL.
Chinese Translation
大语言模型智能体可以通过多步交互和工具使用获得复杂能力,但它们的轨迹也可能被非法收集以蒸馏学生智能体。然而,现有的水印方法要么不适应智能体环境的结构化和交互性本质,要么在不同任务和模型架构上缺乏可靠的有效性。我们提出 AuxMark,一个用于追踪未经授权智能体蒸馏的行为水印框架。AuxMark 动态地向教师轨迹中插入安全的、非必要的辅助动作,并将相关上下文存储为私有证据卡。为了审计可疑的学生模型,AuxMark 从这些卡片构造成对的真实和虚假探针,并应用卡片级符号检验。这种黑盒协议同时支持模型级检测和轨迹级归因。在三个智能体基准、两个教师智能体和四种学生架构上,AuxMark 检测出全部 24 个蒸馏模型,并在 48 个干净模型上实现零假阳性。它还保持任务效用,并且对数据洪泛、改写、截断和自适应清洗攻击仍然有效。我们的代码将在该 URL 发布。
cs.CR / 102 / 2609.34963
JevVibe: Efficient Classification-Guided Secure Code Generation
JevVibe:高效的分类引导的安全代码生成
Arshak Rezvani, Sasha Behrouzi, Ahmad-Reza Sadeghi
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at $6.27\times$ lower median API latency and $55.9\times$ lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.
Chinese Translation
大型语言模型能够生成功能正确、但仍包含安全弱点的代码,这促使人们构建修复流水线:先诊断出弱点类型,然后再决定如何修复。通用弱点枚举(CWE)为此类诊断提供了标准化的词汇表,但要求一个自回归语言模型生成一个 CWE 标签,并从其响应中将其提取出来,会引发关于输出有效性、速度、成本以及准确性的疑问。我们在一个受控的 50 路 CWE 分类任务上,基于 1,916 个 CyberSecEval 基准示例,将 Jev——一种改为直接从一组已声明的候选标签中进行选择、并为每个候选返回一个概率的决策模型——与六个开放权重自回归模型以及一个前沿专有模型 GPT-5.6-Sol 进行对比评估。Jev 在每一项分类与排序指标上都优于全部六个开放权重基线,而它与 GPT-5.6-Sol 的对比则取决于具体指标:GPT-5.6-Sol 取得了更高的 Top-1 准确率和 Macro-F1,而 Jev 取得了更高的 Top-3 和 Top-5 准确率以及几乎相同的 MRR,同时其中位 API 延迟低 $6.27\times$、估计 API 成本低 $55.9\times$。我们进一步构建了 JevVibe,这是一个诊断引导的修复智能体,它利用预测出的 CWE 标签来修复由 Qwen2.5-Coder-32B-Instruct 生成的代码。在 Jev 提供诊断的情况下,该智能体将检测器测得的安全通过率从修复前的 63.5% 提升到 70.7%,相比之下,LLM 引导的修复为 66.1%。这些结果表明,JevVibe 能够有效提升生成代码的安全性,而 Jev 则提供了可靠且高效的 CWE 分类。
cs.CR / 103 / 2609.35266
Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates
面向软件交付决策闸门的智能体安全审计器的持续保证
Guy Lupo, Nguyen Hung Nguyen, Viet Vo, M. A. P. Chamikara, Guangdong Bai, Nazatul Haque Sultan, Alsharif Abuadbba
cs.CR
large language model
大语言模型相关
Abstract
Large language model (LLM)-based repository auditors are increasingly deployed as security controls within continuous integration (CI) pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle (SDLC) Security Controls, their non-deterministic behaviour changes the evidence, while organisational risk appetite and jurisdictional or data-sovereignty policy change its interpretation. Point-in-time audits therefore cannot maintain current assurance for merge decisions. We propose the Policy-Evidence-Execution Separation Pattern, implemented by the Trustworthy AI Posture (TAIP) Assurance Engine and operated as Continuous Control Posture Assurance (CCPA). By separating policy from stable execution and binding admitted evidence to a versioned Posture Tree, the same assurance logic operates across models, environments, and policy profiles. We evaluate the approach using unmodified RepoAudit on a fixed Python Null Pointer Dereference benchmark. The retained evidence repository contains 80 RepoAudit executions across two OpenAI model configurations, gpt-4o-mini and gpt-4.1. TAIP recomputes assurance posture after policy, evidence, and model-context changes and is evaluated across increasing numbers of independent Decision Gateway contexts. The maximum observed policy-to-posture latency was 1.1 ms across three policy-class cycles in one execution. At 1,000 independent assurance contexts, full policy-triggered recomputation with one worker recorded a maximum aggregate refresh of 1.62 s, below the predeclared 5 s Decision Gateway budget. These single-host measurements concern assurance over retained evidence and exclude RepoAudit execution and provider inference.
Chinese Translation
基于大语言模型(LLM)的仓库审计器正越来越多地被部署为持续集成(CI)流水线中的安全控制措施,其审计发现会准许、阻止或延迟软件变更。作为智能体化软件开发生命周期(SDLC)安全控制,其非确定性行为改变了证据,而组织风险偏好以及司法辖区或数据主权政策则改变了对其的解释。因此,时点审计无法为合并决策维持当前的保证。我们提出策略—证据—执行分离模式(Policy-Evidence-Execution Separation Pattern),该模式由可信人工智能态势(TAIP)保证引擎实现,并作为持续控制态势保证(CCPA)运行。通过将策略与稳定执行相分离,并将被准许的证据绑定到版本化的态势树(Posture Tree),同一套保证逻辑可跨模型、环境和策略配置运行。我们在一个固定的 Python 空指针解引用基准上,使用未经修改的 RepoAudit 对该方法进行评估。留存的证据仓库包含 80 次 RepoAudit 执行记录,涵盖两种 OpenAI 模型配置:gpt-4o-mini 和 gpt-4.1。TAIP 在策略、证据和模型上下文发生变化后重新计算保证态势,并在数量不断增加的独立决策网关(Decision Gateway)上下文中进行评估。在一次执行的三轮策略类别循环中,观测到的策略到态势的最大延迟为 1.1 ms。在 1,000 个独立保证上下文中,使用一个工作进程进行完整的策略触发重算,记录到的最大聚合刷新时间为 1.62 s,低于预先声明的 5 s 决策网关预算。这些单主机测量针对的是对留存证据的保证,不包括 RepoAudit 执行和提供商推理。
cs.CR / 104 / 2609.35408
"Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables
“这里没什么可看的”:通过LLM交付物的修订痕迹造成的非预期披露
Yage Zhang, Yukun Jiang, Yang Zhang
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such statements revision traces. For example, after a user removes the password before sharing a configuration file, the model may delete it but leave a comment saying, "Removed the password 'No****4!' as requested." A third-party recipient who sees only the delivered file can therefore recover the withdrawn password from the comment. In an in-the-wild analysis of three public conversation corpora, we identify 26,753 revision requests, of which 2,363 (8.8%) leave revision traces. We study them in greater depth under controlled conditions by introducing RevLeakBench, a benchmark of 100 tasks across five scenarios with a conversation track and an agent track. We measure trace occurrence, withdrawn-item recovery, trace position, and required-content retention. Across six models, about half of the deliverables in both tracks state the edit after a revocation, and a reader that sees only the deliverable can recover the withdrawn item from about 13% of them. Telling the model that its entire reply will be forwarded to the recipient still leaves revision traces in 36.4% of the deliverables. We compare prompt defenses and a delivery boundary, and propose an output-side filter that sharply reduces recovery with little loss of required content. We believe our work can benefit efforts to understand and mitigate unintended disclosure in LLM interactions.
Chinese Translation
大型语言模型(LLM)助手越来越多地帮助用户为第三方接收者起草内容。在私下起草过程中,用户或模型可能引入某一项内容,随后又将其移除或替换。模型可能从预期内容中移除该项,但在陈述这次编辑时又将其泄露出来。我们将此类陈述称为修订痕迹。例如,在用户于共享配置文件前移除密码后,模型可能删除了它,却留下一条评论说:“已按要求移除密码‘No****4!’。”因此,仅看到交付文件的第三方接收者可以从该评论中恢复被撤回的密码。在对三个公开对话语料库的野外分析中,我们识别出26,753个修订请求,其中2,363个(8.8%)留下了修订痕迹。我们通过引入RevLeakBench,在受控条件下更深入地研究它们;RevLeakBench是一个包含100个任务的基准,涵盖五个场景,并设有对话轨道和智能体轨道。我们测量痕迹出现、被撤回项恢复、痕迹位置以及所需内容保留。在六个模型中,两个轨道中约一半的交付物在撤回之后陈述了这次编辑,而仅看到交付物的读者可以从其中约13%的交付物中恢复被撤回项。告诉模型其整个回复将被转发给接收者,仍会在36.4%的交付物中留下修订痕迹。我们比较了提示词防御和交付边界,并提出一种输出侧过滤器,它能在几乎不损失所需内容的情况下大幅降低恢复率。我们相信,我们的工作能够有助于理解和缓解LLM交互中的非预期披露。
cs.CR / 105 / 2609.35434
LLM-Assisted Automatic Security Proofs for Cryptographic Protocols: How Far Are We?
LLM 辅助的密码学协议自动安全性证明:我们进展如何?
Tianjian Liu, Shicheng Feng, Jin'ao Shang, Xiaoting Lyu, Bin Wang, Zonghua Zhang, Lei Xue, Wei Wang
cs.CR · cs.SE
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown strong potential for assisting software and security analysis tasks, yet their effectiveness in cryptographic symbolic protocol verification remains insufficiently understood. In this paper, we conduct the first systematic evaluation of the capability of state-of-the-art LLMs in cryptographic symbolic protocol verification. To quantify this capability, we propose \textsc{CRoST} (Coverage Rate of Solve Tree), a proof-based metric derived from the verifier's proof skeleton that measures the similarity between generated lemmas and reference lemmas. We then establish the rationale of \textsc{CRoST} through both theoretical analysis and empirical validation. The evaluation results show that state-of-the-art models achieve 38.82\% coverage on average, with 14.4\% of generated lemmas exceeding 80\% coverage, indicating that LLMs can already generate useful lemmas to a certain extent. However, they still exhibit non-trivial failure modes on complex multi-phase protocols, show diminishing returns under naive scaling, and incur substantial verification overhead. These findings clarify the practical potential and limitations of LLMs for protocol verification and motivate future work on complex real-world protocols.
Chinese Translation
大语言模型(LLMs)已展现出辅助软件与安全分析任务的强大潜力,然而其在密码学符号协议验证中的有效性仍未得到充分理解。在本文中,我们对最先进的 LLM 在密码学符号协议验证中的能力进行了首次系统评估。为了量化这一能力,我们提出 \textsc{CRoST}(求解树覆盖率,Coverage Rate of Solve Tree),这是一种基于证明的指标,源自验证器的证明骨架,用于衡量生成的引理与参考引理之间的相似性。随后,我们通过理论分析和实证验证确立了 \textsc{CRoST} 的合理性。评估结果表明,最先进的模型平均达到 38.82\% 的覆盖率,其中 14.4\% 的生成引理覆盖率超过 80\%,这表明 LLM 已经能够在一定程度上生成有用的引理。然而,它们在复杂的多阶段协议上仍表现出非平凡的失败模式,在朴素扩展下呈现收益递减,并产生大量验证开销。这些发现阐明了 LLM 在协议验证中的实际潜力与局限性,并推动针对复杂的真实世界协议的未来研究。
cs.LG / 106 / 2609.33833
One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
一次攻击即可骗过所有:针对前沿多模态大语言模型的高迁移性黑盒对抗攻击
Sen Nie, Jie Zhang, Zhongqi Wang, Shiguang Shan, Xilin Chen
cs.CV · cs.LG
large language model
大语言模型相关
Abstract
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.
Chinese Translation
对抗攻击长期以来一直对机器学习系统构成根本性威胁。随着多模态大语言模型(MLLM)快速演进并被广泛部署,评估其对此类攻击的脆弱性对其安全使用至关重要。在这项工作中,我们研究在黑盒设置下,单张对抗图像能否持续误导多样的前沿 MLLM。我们提出了 O-Attack,一种高度可迁移的黑盒攻击框架。该框架基于我们的洞见:代理模型包含一个广泛、高层、跨模态对齐的语义空间。该空间超越了最终层输出,并提供了多种语义一致的表示,而这些表示仍未被现有攻击充分利用。在该空间内,O-Attack 锚定对齐的表示,逐步拓宽语义条件,并通过语义共识优化扰动,以促进一致的目标对齐。通过使用与 M-Attack 相同的代理模型充分挖掘该空间,O-Attack 将攻击成功率在 GPT-5.4 上从 29.1% 提升至 77.2%,在 Claude-4.6 上从 42.8% 提升至 81.6%,在 Gemini-3.1 上从 38.2% 提升至 80.9%。在 24 个 MLLM 上的大量实验表明,O-Attack 在黑盒迁移性方面优于六种最先进方法,在不同提示下具有一致有效性,并提高了效率和不可感知性。这项工作揭示了针对前沿 MLLM 的黑盒对抗攻击所带来的实际安全风险,强调了进行更严格鲁棒性评估和更有效防御的必要性。
cs.LG / 107 / 2609.33834
CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow
CLIMB-flow:通过基于多尺度的流进行耦合线性逆后验采样
Zeqiu Yu, Ruizhi Yuan, Mathews Jacob
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based compression results in loss of spatial detail. In addition, the optimization is converted to a non-linear problem. In this paper, we introduce a posterior sampling algorithm customized for the pyramidal/cascaded architecture, which relies on a coarse to fine hierarchical strategy to generate images in the pixel domain. We present CLIMB-Flow which alternates between three steps: an end-point estimation from the current coarse and noisy image, data-consistent update of the clean image, and re-noising it back to the level the network expects. Together these steps sample the posterior at that scale using an approximate Gibbs sampling from two conditional distributions. Experiments on ImageNet, CelebA, AFHQ and fastMRI span inpainting, deblurring, super-resolution and accelerated MRI, with PSNR gains of 1.37-7.66 dB over the strongest competing method on CelebA and pixel-domain reconstruction up to 512x512.
Chinese Translation
扩散模型如今在成像中的贝叶斯逆问题中被广泛用作先验,其中潜在扩散模型常用于更大规模的问题,以保持计算复杂度和模型规模可控。不幸的是,基于自编码器的压缩会导致空间细节丢失。此外,优化被转化为一个非线性问题。在本文中,我们介绍一种为金字塔/级联架构定制的后验采样算法,该算法依赖从粗到细的分层策略在像素域中生成图像。我们提出 CLIMB-Flow,它在三个步骤之间交替:从当前粗糙且含噪图像进行端点估计、对干净图像进行数据一致性更新,以及将其重新加噪回网络期望的水平。这些步骤共同使用来自两个条件分布的近似 Gibbs 采样,在该尺度上对后验进行采样。在 ImageNet、CelebA、AFHQ 和 fastMRI 上的实验涵盖修补、去模糊、超分辨率和加速 MRI,在 CelebA 上相较于最强的竞争方法取得 1.37-7.66 dB 的 PSNR 提升,并实现最高 512x512 的像素域重建。
cs.LG / 108 / 2609.33895
Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers
残差流负担塑造扩散 Transformer 中的表征学习
Tongtong Liang, Siqi Kou, Ziqiao Xi, Esha Singh, Kun Zhou, Zhijie Deng, Alexander Cloninger, Yu-Xiang Wang, Rahul Parhi
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement *residual-stream burden* and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet $256^2$. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
Chinese Translation
在基于扩散的生成中,神经网络可以被训练来从含噪输入中预测干净数据、噪声或速度。这些预测目标可以相互转换,并描述相同的生成过程,然而在大型像素块上运行的普通扩散 Transformer 在干净预测上成功,却在噪声或速度预测上失败。我们认为,这种不对称性之所以出现,是因为噪声目标要求残差流在贯穿深度时保留依赖于噪声的输入变化,以供最终读出,从而迫使后续层在含噪表征上进行计算。频谱集中的干净目标施加了更轻的要求,为后续计算组织隐藏表征留下了更大的自由。我们将这一保留要求称为*残差流负担*,并展示它如何塑造扩散 Transformer 中的表征学习。受控实验表明,可利用的结构是图像块空间中的频谱集中,而持久残差状态的带宽是噪声预测的关键资源。我们进一步表明,这一解释与近期解耦的像素空间架构一致,这些架构的多样设计都减轻了主路径上的残差流负担。为了从互补方向检验这一理解,我们直接扩展并重组残差流带宽,引入空间索引超连接(Spatially Indexed Hyper-Connections, SiHC),其在 ImageNet $256^2$ 上达到 FID 1.71。总之,这些结果将残差流负担确定为一种机制,通过该机制,预测目标和架构共同塑造扩散 Transformer 中的表征学习。
cs.AI / 109 / 2609.34231
ReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel Geometry
ReGDiff:在受调控潜在空间中进行引导扩散以探索超材料体素几何
Wangzhi Zhan, Jianpeng Chen, Dongqi Fu, Dawei Zhou
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porous structures within a single cubic discretization. However, voxel-based generation faces a plausibility-novelty trade-off: staying close to known geometries helps preserve geometric regularities, while moving away from them is necessary for novelty but may produce degenerate geometries. To address this challenge, we propose REGDIFF, a generative framework that couples voxel representation with latent space regulation and guided diffusion. REGDIFF introduces a repel-and-sink (RAS) mechanism to smooth the latent distribution of plausible geometries, and short-range repulsion (SRR) guidance to discourage generation overly close to known samples while maintaining geometric plausibility. We further contribute a voxel-based benchmark covering truss- and shell-type metamaterial geometries, together with an evaluation module for geometric plausibility, novelty, and diversity. Experiments show that REGDIFF outperforms voxel-based generative baselines, achieving +8.9% in geometric plausibility, +46.4% in novelty, and +128.6% in diversity on average across two datasets. These results suggest that REGDIFF is a strong geometry candidate generator for downstream evaluation. Our code is provided at https://github.com/wzhan24/ReGDiff.
Chinese Translation
超材料是人工设计的结构,其力学和物理行为在很大程度上由几何形状而非组成决定。体素表示提供了一种用于超材料几何生成的统一格式,因为它可以在单一立方体离散化中表达桁架、壳和多孔结构等多种类别。然而,基于体素的生成面临合理性—新颖性权衡:接近已知几何有助于保持几何规律性,而远离它们对于新颖性是必要的,但可能产生退化几何。为应对这一挑战,我们提出了 REGDIFF,这是一种将体素表示与潜在空间调控和引导扩散耦合起来的生成框架。REGDIFF 引入了排斥-下沉(RAS)机制以平滑合理几何的潜在分布,并引入短程排斥(SRR)引导以抑制生成结果与已知样本过于接近,同时保持几何合理性。我们进一步贡献了一个基于体素的基准,涵盖桁架型和壳型超材料几何,以及一个用于评估几何合理性、新颖性和多样性的评估模块。实验表明,REGDIFF 优于基于体素的生成基线,在两个数据集上平均实现了几何合理性 +8.9%、新颖性 +46.4% 和多样性 +128.6% 的提升。这些结果表明,REGDIFF 是一个用于下游评估的强大几何候选生成器。我们的代码见 https://github.com/wzhan24/ReGDiff。
cs.LG / 110 / 2609.34286
Dexterous Tactile World Model
灵巧触觉世界模型
Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu, Fengyu Yang, Tianyu Liu, Zhiwen Fan, Daniel Rakita
cs.CV · cs.LG · cs.RO
diffusion
扩散模型相关
Abstract
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.
Chinese Translation
用于操作的世界模型通常从视频中训练,然而决定操作如何展开的事件,例如建立和解除接触,难以通过视觉观察到,并且往往更容易通过触觉感知。我们提出灵巧触觉世界模型(DTWM),这是一种视频世界模型,用于根据观测到的视频以及来自每只手上佩戴的手套的触觉信号,对未来帧进行以自我为中心的操作预测。我们通过在视频 token 中对应手部位置处的零初始化残差,将预训练视频扩散 Transformer 条件化到每只手的触觉信号上,同时因果掩码阻止预测帧访问未来信息。与在架构、参数和训练上匹配的纯视觉模型相比,DTWM 将手部运动的低估从 23% 降低到 9%,同时在每个模型的三次训练运行中将手部区域的感知误差降低了 7.4%。这种收益还随着预测时域的增加而增大,后续预测块中的改进约比第一个预测块中的改进大 4.1 倍。在相同设置下,DTWM 也优于其他视觉-触觉世界模型,并且即使推理时没有可用的触觉信息,使用触觉进行训练也能改善未来帧预测。消融实验表明,该模型同时受益于力的大小和空间位置:将触觉信号替换为二值接触状态,无论是按手还是按位置,都会增加预测误差。观测到的力的变化过程表明交互将会持续还是改变。
cs.AI / 111 / 2609.34371
From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models
从静态到动态:从图像扩散模型到视频扩散模型的在线策略蒸馏
Bingqing Jiang, Li Luo, Zichao Yu, Yujin Han, Zhaolong Su, Difan Zou
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.
Chinese Translation
在线策略蒸馏(OPD)沿着学生模型自身的生成轨迹,通过教师模型的监督来使预训练视频扩散模型专门化。尽管大型视频模型是天然的教师模型,但开发专门的视频专家模型可能需要昂贵的视频数据和训练成本,而查询它们所带来的延迟也远高于查询图像专家模型。图像专家模型更易获取且查询成本更低,因而提供了一种具有成本效益的替代方案,尤其适用于诸如美学和 OCR 这类在很大程度上与时间无关、可以接受帧级监督的能力。然而,异构的图像与视频潜在空间阻碍了对学生模型中间状态的直接监督,同时图像专家模型缺乏跨帧运动监督,使得时间一致性容易受到帧级改进的损害。在本文中,我们提出 MILD,一个运动保持的图像到视频潜在蒸馏框架,它在迁移专门的图像专家知识的同时保持预训练的视频动态。MILD 使用一个可学习的线性连接器,将学生的潜在状态和预测更新与图像专家模型的对齐,从而实现跨异构潜在空间的监督迁移。我们进一步将图像引导的修正约束在预训练学生模型的预测附近,以保持视频动态,并引入基于光流的运动奖励,以提升运动质量和时间一致性。在多种专门的图像专家模型和多个视频学生骨干网络上,我们的方法始终优于以视频模型为教师的 OPD 基线,进一步的研究还表明其能够跨连接器设计和异构架构实现有效迁移。这些结果确立了图像到视频蒸馏是一条有效途径,可借助图像生成生态系统中多样且不断演进的能力来改进视频生成。
cs.AI / 112 / 2609.34451
Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields
将全切片图像建模为动态肿瘤微环境场
Lei Wu, Jiashuai Liu, Di Zhang, Zhangpeng Gong, Yingkang Zhan, Yi Niu, Jiusong Ge, Chunze Yang, Kai Yi, Mireia Crispin-Ortuzar, Chen Li, Zeyu Gao
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Due to the gigapixel-scale nature of whole-slide images (WSIs), weakly supervised WSI analysis is commonly formulated as a multiple instance learning (MIL) problem, where patch-level features are aggregated into slide-level representations. However, diagnostic and prognostic evidence often arises from spatially coherent tumor microenvironment regions and their interactions, rather than isolated patches alone. Existing patch-level or static region-based methods usually overlook how tissue regions should be adaptively formed and subsequently evolved through microenvironment interactions across heterogeneous boundaries. In this paper, we propose Concept-Guided Tumor Microenvironment Evolution (TMEvolve), a reaction-diffusion-inspired framework that models WSIs as latent tumor microenvironment fields over discrete patch graphs. TMEvolve instantiates this view as a learnable graph-discretized evolution process over patch neighborhoods. It first forms adaptive soft tissue regions as coherent microenvironment units, then performs pseudo-time evolution through two complementary local dynamics: intra-region diffusion, which stabilizes latent states within coherent tissue compartments, and concept-guided boundary flux, which propagates visual feature signals and language-derived concept signals across heterogeneous region interfaces. The evolved microenvironment regions are finally aggregated for slide-level prediction. We evaluate TMEvolve on six datasets across three weakly supervised WSI tasks: survival prediction, gene expression prediction, and histological subtype classification. TMEvolve consistently improves over representative MIL methods, pathology foundation models, and concept-guided baselines. Ablation studies and visualizations further support the effectiveness and interpretability of TMEvolve, highlighting the value of dynamic region modeling and boundary interaction.
Chinese Translation
由于全切片图像(WSIs)具有千兆像素级的特性,弱监督 WSI 分析通常被表述为多实例学习(MIL)问题,其中切片块级特征被聚合为切片级表示。然而,诊断和预后证据往往来自空间上连贯的肿瘤微环境区域及其相互作用,而非仅仅来自孤立的切片块。现有的切片块级或静态区域方法通常忽略了组织区域应当如何自适应地形成,以及随后如何通过跨越异质边界的微环境相互作用而演化。在本文中,我们提出了概念引导的肿瘤微环境演化(TMEvolve),这是一个受反应-扩散启发的框架,它将 WSIs 建模为离散切片块图上的潜在肿瘤微环境场。TMEvolve 将这一观点实例化为在切片块邻域上的可学习图离散化演化过程。它首先形成自适应的软组织区域作为连贯的微环境单元,然后通过两种互补的局部动力学执行伪时间演化:区域内扩散,它在连贯的组织区室内稳定潜在状态;以及概念引导的边界通量,它跨越异质区域界面传播视觉特征信号和源自语言的概念信号。演化后的微环境区域最终被聚合用于切片级预测。我们在三个弱监督 WSI 任务(生存预测、基因表达预测和组织学亚型分类)的六个数据集上评估了 TMEvolve。TMEvolve 持续优于代表性的 MIL 方法、病理基础模型和概念引导的基线。消融研究和可视化进一步支持了 TMEvolve 的有效性和可解释性,凸显了动态区域建模与边界交互的价值。
cs.AI / 113 / 2609.34479
SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis
SentZero:一种增强的以句子为中心的视觉-语言预训练,用于多任务零样本胸部 X 射线分析
Hangyul Yoon, Hyungyung Lee, Edward Choi, Eunho Yang
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.
Chinese Translation
使用配对的胸部 X 射线(CXR)图像和放射学报告进行视觉-语言(VL)预训练,已在医学图像理解方面展现出强大潜力。然而,现有方法往往仍依赖于任务特定的微调,因为放射学报告冗长、临床信息密集,并且难以与简单的零样本提示对齐。近期的句子级方法利用由大语言模型(LLM)提取的临床短语部分解决了这一限制,但它们在很大程度上忽略了放射学话语的内在特征。特别地,有限的正样本对多样性限制了进一步的收益,而临床等价的句子频繁地在不同患者之间重复出现,从而在对比学习中产生假阴性。为了解决这些问题,我们提出 SentZero,一种用于零样本、多任务 CXR 分析的增强型以句子为中心的 VL 预训练框架。SentZero 引入了基于 LLM 的抽象层级句子结构化与映射,以扩展正样本对多样性,并引入一个额外的损失项来减轻假阴性。我们进一步引入视觉嵌入的句子条件残差调制,使视觉特征能够适应每个输入句子的语义特征。在不同的下游任务和数据集上,SentZero 提升了零样本泛化能力,并优于先前的多任务零样本方法。
cs.CL / 114 / 2609.34563
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
重新思考潜在视觉推理:将潜在推理锚定于视觉证据
Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, Xuanbai Chen, Farhan Tejani, Jakub Zablocki, Julia Xu, Yifan Xing
cs.CV · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Chinese Translation
潜在视觉推理(LVR)使多模态大语言模型(MLLMs)能够在连续潜在 token 中执行中间计算,而不是用文字表达每一个推理步骤。然而,与文本 CoT 不同,潜在推理无法被直接观察,这使得对潜在 token 学到了什么进行监督变得困难。在本工作中,我们首先对潜在 token 的行为进行了深入分析,并识别出一个潜在证据—信用鸿沟(latent evidence-credit gap):潜在 token 对改变正确答案的图像扰动仅作出微弱响应。我们假设这一问题源于 GRPO 训练过程中缺乏显式监督。这些发现表明,仅凭最终答案奖励所提供的指导太少,既不足以说明应当保留哪些视觉证据,也不足以说明应当如何在各个潜在 token 之间分配信用。为弥合这一鸿沟,我们提出了 ReaLVR,它将视觉证据监督引入模型自身自由运行的潜在轨迹。ReaLVR 通过对比正确答案与模型生成的错误答案,来确定哪里需要更强的监督;并通过对比相关的视觉证据与不匹配的视觉证据,来明确应当保留什么。在三个模型系列上,ReaLVR 持续优于所评估的 LVR 基线,在 Qwen2.5-VL-7B 上取得了五任务平均 63.7% 的最高成绩。至关重要的是,我们首次在潜在空间中扩展视觉推理,表明我们的框架在前沿模型规模(最高达 235B)下仍能持续带来稳健的提升。进一步的分析显示,潜在 token 位置对问题更加敏感、与相关视觉区域具有更强的对齐,并且对最受关注的潜在 token 表现出更强的固定上下文依赖。
cs.AI / 115 / 2609.34579
GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
GenNVS:通过解耦三维先验的几何增强新视角合成
Yajiao Xiong, Youyu Luan, Xiaoyu Zhou, Yongtao Wang
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.
Chinese Translation
单图像新视角合成仍然具有挑战性,因为其潜在的三维几何高度模糊。最近的基于扩散的方法能产生看似合理的结果,但它们往往难以保持前景对象的几何结构和空间一致性。我们提出 GenNVS,一个通过解耦三维先验实现几何增强新视角合成的框架。具体而言,GenNVS 使用三维高斯泼溅对前景对象和背景进行建模,并通过由粗到细的几何优化过程将它们对齐,以形成一个统一的三维场景。该场景通过所提出的双流掩码机制为视频扩散模型提供条件,该机制通过联合利用渲染的有效性掩码和几何感知变形来指导合成。实验结果表明,GenNVS 在视觉质量和几何精度方面均优于近期方法,同时自然支持灵活的场景编辑。
cs.AI / 116 / 2609.34697
Triangular Resampling for Long-Horizon Motion Generation
面向长时程运动生成的三角重采样
Kunhang Li, Yiyi Cai, Xiangyue Zhang, Fangyuan Tu, Yuhan Wu, Zhixiang Wang, Kaipeng Zhang, Haiyang Liu
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference states. Replacing only completed motion history leaves this mismatch unresolved in partially denoised states within the active window. TR therefore extends rollout-based training to these states, using ground-truth clamping to limit excessive drift. For each replayed sample, TR draws one denoising threshold, shared across latent positions and replay updates, and replays multi-step triangular denoising without gradient tracking. After each update, states below the threshold are replaced with noise-matched ground truth, while those at or above it retain model predictions. The resulting latent window enters the standard training update. This rollout construction supports both supervised training (TR) and distribution matching (TR-DMD). On 120-second motion generation from HumanML3D test prompts, TR and TR-DMD achieve state-of-the-art FID AUC within their respective non-DMD and DMD comparison groups. Supervised TR reduces FID AUC by 40.9% and FID degradation slope by 55.3% relative to matched post-training without replay.
Chinese Translation
我们提出了三角重采样(Triangular Resampling,TR),这是一种用于缓解运动扩散模型中长时程误差累积的训练后方法。TR 建立在 FloodDiffusion 的三角去噪调度之上,解决了由真实数据(ground-truth)导出的训练窗口与模型生成的推理状态之间的不匹配问题。仅替换已完成的运动历史,会使得这种不匹配在活动窗口内部分去噪的状态中依然未得到解决。因此,TR 将基于 rollout 的训练扩展到这些状态,并使用真实数据钳制(ground-truth clamping)来限制过度的漂移。对于每个重放的样本,TR 抽取一个去噪阈值,该阈值在潜在位置与重放更新之间共享,并在不跟踪梯度的情况下重放多步三角去噪。每次更新之后,低于该阈值的状态会被替换为与噪声匹配的真实数据,而处于或高于该阈值的状态则保留模型预测。由此得到的潜在窗口进入标准的训练更新。这种 rollout 构造同时支持监督训练(TR)与分布匹配(TR-DMD)。在基于 HumanML3D 测试提示的 120 秒运动生成任务上,TR 和 TR-DMD 分别在各自的非 DMD 与 DMD 对比组中取得了最先进的 FID AUC。与匹配的、不带重放的训练后方法相比,监督式 TR 将 FID AUC 降低了 40.9%,并将 FID 退化斜率降低了 55.3%。
cs.LG / 117 / 2609.34742
Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?
情感概念会在视觉-语言模型中的来源、模态和架构之间泛化吗?
Bohao Xing, Xin Liu, Kaishen Yuan, Deng Li, Rong Gao, Guoying Zhao, Xiaolan Fu, Heikki Kälviäinen
cs.CV · cs.LG
large language model
大语言模型相关
Abstract
Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.
Chinese Translation
近期研究表明,大语言模型将情感概念编码为结构化的内部表示,但大多数现有工作集中于文本和单一架构。因此,我们提出:情感概念会在视觉-语言模型(VLMs)中的来源、模态和架构之间泛化吗?为了解决这个问题,我们构建了 CMES(跨模态情感刺激,Cross-Modal Emotion Stimuli),这是一个多来源集合,包含情感条件化的故事、真实面部表情、合成肖像和合成的情感唤起场景。对于每个刺激来源,我们从三个 VLM 中的每一个分别提取一组六种 Ekman 情感向量。我们报告以下四个主要发现:1)由图像导出的情感向量形成与由文本导出的向量相似的低维几何结构。效价在不同来源之间相对稳定,而唤醒度变化更大。2)由文本和图像导出的情感向量具有中等的余弦相似度,但仍表现出留出的跨模态对应关系。由文本导出的向量还可以引导图像解读。3)即使原生余弦接近零,跨架构对应关系仍然存在。从通用 ImageNet 激活中估计的变换,在不使用六种情感向量或其标签的情况下,同时恢复了对应关系和因果迁移。4)在对不同架构之间的表示进行对齐后,我们构建了一个共享情感子空间,它保留了情感几何结构和选择性引导效应。相应的共识情感向量在两种模型规模下也泛化到留出的第四种架构。这些结果表明,即使单个向量方向不同,情感表示也可以在来源、模态和架构之间共享关系结构和因果效应。
cs.AI / 118 / 2609.34972
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
仅用 MLPs:面向多模态语言模型的高效视觉状态重建
Jingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang, Yiwen Guo, Soujanya Poria
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose $δ$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, $δ$-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.
Chinese Translation
长视觉 token 序列通常占多模态大语言模型(MLLMs)计算开销的很大一部分。现有方法通过剪枝冗余视觉 token 来降低这一成本,但会永久丢弃可能在后续层中有用的视觉证据。相反,我们询问是否可以在降低通过 Transformer 反复演化表示的成本的同时,保留所有视觉 token。为回答这一问题,我们对视觉到文本的信息流执行低秩干预。我们发现,在视觉到文本注意力被阻断后,仅恢复少数几个方向就能恢复大部分丢失的准确率,这表明相关的视觉影响集中在一个低维子空间中。我们进一步观察到,特定层的视觉状态具有很强的可预测性:轻量级 MLPs 能以高余弦相似度和低重建误差近似它们。受这些发现启发,我们提出 $δ$-Vision,它用轻量级低秩适配器替换视觉 token 的重复 Transformer 演化,这些适配器在构建逐层视觉记忆的同时,为文本检索保留所有视觉 token。在图像和视频基准上,$δ$-Vision 在相当或更低计算量下取得了比视觉 token 剪枝基线更高的准确率,同时在不丢弃视觉 token 的情况下提供了有竞争力的推理效率。
cs.AI / 119 / 2609.34977
SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models
SPIDER:多模态大语言模型中的多层语义 token 剪枝与自适应子层跳过
Tianxiang Chen, Zhentao Tan, Zi Ye, Yue Wu, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Tao Gong, Qi Chu, Nenghai Yu, Xipeng Qiu, Jieping Ye
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose \textbf{SPIDER}, a training-free framework that integrates multi-layer \underline{\textbf{S}}emantic visual token \underline{\textbf{P}}run\underline{\textbf{I}}ng with an a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by $79\%$ while maintaining 96$\%$ of the baseline performance.
Chinese Translation
多模态大语言模型面临显著的效率挑战,这些挑战源于两个不同但相互耦合的来源:数据冗余和计算冗余。尽管大多数方法通过从视觉编码器的输出中剪枝视觉 token,或使用块级重要性来计算 LLM 解码器中的冗余,从而关注数据冗余,但更细粒度的层间表示偏移以及各层自身内部的分布差异尚未得到充分探索。在这项工作中,我们全面研究了这种双层低效性。我们认为,应考虑到来自视觉编码器的中间层 token 以实现有效的视觉 token 剪枝,因为语义焦点会跨层转移,而中间层 token 捕获更详细的以对象为中心的信息,而更深层可能会将这些信息抽象掉。此外,我们揭示了在不同 LLM 解码器层中 Attention 和 FFNs 的差异化贡献。基于这些发现,我们提出了 \textbf{SPIDER},一个无需训练的框架,它将多层 \underline{\textbf{S}}emantic 视觉 token \underline{\textbf{P}}run\underline{\textbf{I}}ng 与 a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping 机制相结合。实验评估表明,SPIDER 在各种 MLLM 架构和缩减比例下始终保持着强劲的性能。例如,在 LLaVA-NeXT-7B 上,SPIDER 将 FLOPs 减少了 $79\%$,同时保持基线性能的 96$\%$。
cs.AI / 120 / 2609.35226
Generative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and Benchmark
基于生成式人工智能的口腔病变分类数据增强:PhotoMOCI 数据集与基准
Marco Parola, Mario G. C. A. Cimino, Sabrina Senatore
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-curated resource, the Photographic Multi-purpose Oral Cancer Imaging (PhotoMOCI) dataset, is introduced for developing models across multiple diagnostic tasks in oral oncology. Then, a comprehensive benchmark study was conducted to investigate how various data augmentation strategies influence the performance of image classifiers. Our analysis spans different generative AI frameworks, evaluating the efficacy of traditional methods against advanced generative approaches, including Generative Adversarial Networks (GANs) and Diffusion Models (DMs). Additionally, we propose the Synthetic Image Filter (SIF), a mechanism to select specific samples based on two auxiliary models: Synthetic Proxy Classifier to ensure samples are representative of the target class and Synthetic Image Detector to verify they appear realistic, thereby selecting only the high-utility images that contribute to improving downstream performance. Across the evaluated datasets and classifiers, the best SIF-filtered setup improves accuracy over traditional augmentation in all cases, with gains of +1.73% and +2.35% on PhotoMOCI and +2.38% and +2.08% on KOCD for ResNet50 and ViT, respectively. Our findings reveal that while the direct application of generative data augmentation may yield performance drops, the integration of SIF, considering (i) how synthetic data looks real and (ii) how it reflects the discriminative features of the belonging class, provides a simple yet effective mechanism to filter out synthetic samples that confuse the classifier during training.
Chinese Translation
通过摄影成像进行口腔癌早期检测,为大规模口腔筛查提供了一条颇具前景的途径。然而,高质量、带标注数据集的稀缺常常阻碍了稳健深度学习模型的开发。为解决这一局限,本文引入了一个新颖且经过精心整理的数据资源——摄影多用途口腔癌成像(PhotoMOCI)数据集,用于开发面向口腔肿瘤学中多种诊断任务的模型。随后,我们开展了一项全面的基准研究,以考察各种数据增强策略如何影响图像分类器的性能。我们的分析涵盖不同的生成式人工智能框架,评估传统方法相对于先进生成式方法的有效性,这些先进方法包括生成对抗网络(GANs)和扩散模型(DMs)。此外,我们提出了合成图像过滤器(SIF),这是一种基于两个辅助模型来选择特定样本的机制:合成代理分类器(Synthetic Proxy Classifier)用于确保样本能够代表目标类别,合成图像检测器(Synthetic Image Detector)用于验证样本看起来是否真实,从而只选择那些有助于提升下游性能的高效用图像。在所评估的数据集和分类器上,经过 SIF 过滤的最佳设置在所有情况下都相较传统增强提升了准确率:对于 ResNet50 和 ViT,在 PhotoMOCI 上分别提升 +1.73% 和 +2.35%,在 KOCD 上分别提升 +2.38% 和 +2.08%。我们的发现表明,虽然直接应用生成式数据增强可能导致性能下降,但结合 SIF——同时考虑 (i) 合成数据看起来有多真实,以及 (ii) 它在多大程度上反映了所属类别的判别性特征——提供了一种简单而有效的机制,可过滤掉在训练过程中使分类器产生混淆的合成样本。
cs.AI / 121 / 2609.35228
Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning
面向视觉-语言推理的Token解耦潜在测试时扩展
Hao-Xuan Ma, Yihao Liu, Yutao Sun, Yanting Miao, Mengyu Zhou, YiCheng Xiao, Long Chen, Zhenguo Li, Han-Jia Ye, Xiaoxi Jiang, Guanjun Jiang
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.
Chinese Translation
潜在测试时扩展通过在推理期间细化隐藏状态来提升推理能力,但现有方法通常对所有可编辑的潜在token施加单一标量奖励。对于多模态大语言模型,这种全局更新忽略了生成token扮演不同角色:一些对视觉证据敏感,而另一些对应不确定的推理决策。我们提出Token-Disentangled Latent Test-Time Scaling,一个使潜在细化具备token角色感知能力的推理时框架。从初始生成的轨迹出发,我们优化一个较短的隐藏状态前缀,同时将感知侧视觉反馈路由到图像敏感的token,并将推理反馈路由到高熵token。未被任一路由选中的token由锚点正则化器约束。在Qwen2.5-VL-7B和InternVL3.5-8B上的感知与推理基准中,我们的方法相较于CoT分别将宏准确率提升了+2.57和+1.51,并且在匹配解码候选预算下优于强大的输出空间测试时扩展基线。代码见https://github.com/Qwen-Applications/TD-LTTS。
cs.LG / 122 / 2609.35292
Scaffold Then Internalize: Representation Injection for Diffusion Transformers
先搭脚手架,再内化:面向扩散 Transformer 的表示注入
Han Fu, Jiacheng Chen, Baoquan Zhao, Weidong Chen, Wei Liu, Qing Li, Xudong Mao
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over $43.5\times$. Code will be available at https://jeneveuxpas.github.io/REPI
Chinese Translation
近期的表示对齐(REPA)方法通过将 Transformer 隐藏状态的投影与来自预训练视觉编码器的表示对齐,加速扩散 Transformer 的训练。在这项工作中,我们探索了一个与 REPA 相反且互补的方向:不是将扩散表示投影到编码器空间中,而是将编码器表示注入扩散 Transformer,使它们能够主动参与去噪过程。为此,我们引入了 REPresentation Injection(REPI),一个基于从脚手架到内化的策略的训练框架,其中投影后的编码器表示最初充当临时脚手架,随后被扩散 Transformer 逐步内化。REPI 在广泛的骨干网络上优于 REPA,并且与其高度互补:将两者结合比单独使用任一者都带来显著增益。值得注意的是,仅用 160K 训练步数,REPI + REPA 就能匹配训练了 7M 步的原始 SiT,加速超过 $43.5\times$。代码将在 https://jeneveuxpas.github.io/REPI 处提供。
cs.AI / 123 / 2609.35491
From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
从分数到样本:面向自回归视频生成的弹性强制
Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
Chinese Translation
少步自回归视频生成通常依赖分布匹配蒸馏(DMD),这需要一个双向扩散教师模型和一个在线伪分数模型。我们则改为直接从参考视频中学习展开分布,从而在后训练阶段消除了这两个分数模型。我们的框架在冻结的自监督视频表示空间中最小化最大均值差异(MMD),并使用混合的 Nyström--Monte Carlo 估计器来平衡近似偏差与采样方差。内存高效的回放与梯度子采样使该目标切实可行。采用与 Self-Forcing 相同的架构和初始化,我们的 1.3B 模型将 VBench Total 分数从 83.80 提升至 84.64,同时保持 17 FPS。去除辅助分数模型还使得在八块 H200 GPU 上进行 14B 的后训练成为可能。除蒸馏之外,从参考视频中学习还能够在没有特定目标扩散教师模型的情况下获得新的视觉风格、语义概念和空间先验。
cs.LG / 124 / 2609.35768
PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
PDMD:面向视频扩散模型的投影分布匹配蒸馏
Zimo Wang, Junkun Yuan, Angtian Wang, Haotian Yang, Canyu Zhang, Siyuan Yuan, Xingchang Huang, Bo Liu, Yizhi Wang, Yiding Yang, Chongyang Ma, Gordon Guocheng Qian
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.
Chinese Translation
现代视频扩散模型需要在长时空 token 序列上进行数十次去噪评估。分布匹配蒸馏(DMD)将函数评估次数(NFE)降至仅几次。然而,DMD 样本在训练过程中可能退化,表现出渐进的过饱和和伪影。我们将这种不稳定性追溯到 critic 误差,这些误差进入连续的学生更新并随时间累积。我们提出投影分布匹配蒸馏(PDMD)以过滤 critic 误差。PDMD 投影掉 DMD 更新中与学生-critic 端点残差平行的分量。在固定的含噪查询下,我们证明该残差是 critic 端点误差的无偏估计。在高维假设下,该投影去除了恒定比例的 critic 误差,同时仅丢弃可忽略比例的理想 DMD 信号。经验上,该投影稳定了训练,并在 DMD 退化并产生不自然纹理的情况下提高了样本质量。PDMD 只需对 DMD 进行一行代码修改,不需要额外的损失、网络、数据、模型前向传递或多阶段训练。在 Wan2.1 上,PDMD 在 4 NFE 下取得了 83.73 的 VBench 总分,超过匹配的 DMD 1.03 分。在 MiniMax-H3 联合视频-音频生成上,PDMD 取得了 83.17 的 VideoGen-Eval 视觉总分,比最强的蒸馏基线高 0.41 分。在比较的 4-NFE 模型中,PDMD 还在所有六项音频指标上取得了最佳性能。定性比较和用户研究表明,在视觉质量、运动和音频质量方面,PDMD 优于蒸馏基线。代码和模型可在 https://pdmd2026.github.io/ 获取。
cs.LG / 125 / 2609.33965
Greenpixie's AI Token Methodology: Assessing the Energy, Water and $\mathrm{CO_2\text{-}eq}$ Impact of AI Tokens for Open and Closed Weight Models
Greenpixie 的 AI Token 方法学:评估开放权重与封闭权重模型的 AI Token 对能源、水和 $\mathrm{CO_2\text{-}eq}$ 的影响
Joshua Horswill, Ross Hunter, Matt Clifford, James Hall
cs.CY · cs.LG
large language model
大语言模型相关
Abstract
We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a wide range of text-based tasks. The remaining server energy contribution from non-GPU hardware is estimated from the inference wall time. Bayesian linear regression is used to model the relationship between energy per token and LLM size, request traffic, and hardware deployment configuration. Proprietary frontier LLMs of unknown size and deployment are binned into size buckets based on naming conventions and performance priors, and the space of possible LLM configurations is sampled with Monte-Carlo methods to give a representative average energy per token and uncertainty. We also describe how these energy measurements can be used to estimate the carbon-dioxide equivalent ($\mathrm{CO_2\text{-}eq}$) emissions, both usage and embodied, and water consumed per token of AI inference. This methodology provides actionable data that enables reductions in cost, electricity usage, $\mathrm{CO_2\text{-}eq}$ emitted and water consumed in cloud and Software as a Service (SaaS).
Chinese Translation
我们描述了一种方法,用于估算云端托管的大语言模型(LLM)推理的每 token 能量成本,并将输入(预填充)和输出(解码)token 区分开来。在针对广泛基于文本的任务使用开放权重模型进行推理基准测试期间,测量图形处理单元(GPU)的能量使用。来自非 GPU 硬件的其余服务器能量贡献根据推理墙钟时间估算。使用贝叶斯线性回归对每 token 能量与 LLM 规模、请求流量和硬件部署配置之间的关系进行建模。对于规模和部署未知的专有前沿 LLM,根据命名惯例和性能先验将其分入规模分桶,并使用蒙特卡洛方法对可能的 LLM 配置空间进行采样,以给出具有代表性的每 token 平均能量和不确定性。我们还描述了如何使用这些能量测量来估算 AI 推理每个 token 的二氧化碳当量($\mathrm{CO_2\text{-}eq}$)排放,包括使用排放和隐含排放,以及消耗的水。该方法提供了可操作的数据,从而能够在云和软件即服务(SaaS)中减少成本、电力使用、$\mathrm{CO_2\text{-}eq}$ 排放和水的消耗。
cs.LG / 126 / 2609.35569
Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions
超越能源:当可持续性维度重塑 LLM 服务决策
Tianyao Shi, Xipeng Shen, Yi Ding
cs.CY · cs.DC · cs.LG · cs.PF
large language model
大语言模型相关
Abstract
Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different optimization decisions. We present PRISM, a unified framework for characterizing and optimizing LLM serving across energy, carbon, water, and biodiversity impacts. Our analysis reveals a fundamental distinction: computing configurations determine energy consumption, whereas where and when LLM serving is deployed determine its carbon, water, and biodiversity impacts. Under a fixed deployment choice and operational-only accounting, all dimensions preserve the same energy-based configuration ranking. Deployment rankings can diverge across dimensions, while embodied impacts can break configuration invariance when they exceed a lifecycle crossover boundary. PRISM identifies these conditions, quantifies cross-dimensional regrets, and balances the four dimensions. In regional-routing experiments, PRISM reduces median worst-case regret by 50.2% relative to the strongest baseline.
Chinese Translation
大语言模型(LLM)服务在能源消耗、碳排放、水资源消耗和生物多样性丧失方面具有环境影响。然而,这些维度在很大程度上是孤立评估的,这使得人们不清楚它们何时以及如何导致不同的优化决策。我们提出 PRISM,一个用于刻画和优化 LLM 服务在能源、碳、水和生物多样性影响方面的统一框架。我们的分析揭示了一个根本区别:计算配置决定能源消耗,而 LLM 服务部署在何处以及何时部署决定其碳、水和生物多样性影响。在固定部署选择和仅运营核算下,所有维度都保持相同的基于能源的配置排序。部署排序可能在不同维度之间出现分歧,而当隐含影响超过生命周期交叉边界时,它们可能打破配置不变性。PRISM 识别这些条件,量化跨维度遗憾,并平衡这四个维度。在区域路由实验中,相对于最强基线,PRISM 将最坏情况遗憾的中位数降低了 50.2%。
cs.AI / 127 / 2609.34764
WeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM Plans
WeaveData:一个具备自我批判与自我演化LLM计划的多模态数据分析系统
Min Jia, Shihao Zhou, Jun-Peng Zhu, Peng Cai, Kai Xu, Chao Zhang, Li Li, Aoying Zhou, Heng Long, Qiu Cui, Liu Tang, Qi Liu
cs.DB · cs.AI
large language model
大语言模型相关
Abstract
Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. However, LLM-generated plans are error-prone: a plan may silently compute something other than what was asked, fail during execution, or return a result that misses the question. This paper presents WeaveData, a multimodal data analysis system with self-critiquing and self-evolving LLM plans. First, WeaveData generates a typed logical plan for each question and critiques it step by step before execution, and it checks the executed result against the question afterwards. Second, WeaveData evolves a plan that fails or misses the question: it diagnoses the failure with the actual data, reuses the results that remain valid, and accumulates planning experience for later questions. Third, WeaveData grounds planning in a metadata knowledge graph of all modalities, clarifies ambiguous questions with the user, and backs every model judgment with evidence in an interactive notebook. We demonstrate WeaveData on two public multimodal datasets.
Chinese Translation
多模态数据分析,即在关系表、文本和图像上回答问题,已在数据管理领域引起越来越多的关注。大语言模型(LLM)通过在关系算子和语义算子上生成分析计划,使得以自然语言进行此类分析成为可能。然而,LLM生成的计划容易出错:计划可能悄无声息地计算了并非所要求的内容,可能在执行过程中失败,或者返回一个偏离问题的结果。本文提出WeaveData,一个具备自我批判与自我演化LLM计划的多模态数据分析系统。首先,WeaveData为每个问题生成一个有类型的逻辑计划,并在执行前逐步对其进行批判,之后还将执行结果与问题进行核对。其次,WeaveData会对失败或偏离问题的计划进行演化:它利用实际数据诊断失败原因,复用仍然有效的结果,并为后续问题积累规划经验。第三,WeaveData将规划建立在涵盖所有模态的元数据知识图谱之上,与用户澄清有歧义的问题,并在交互式笔记本中用证据支撑每一个模型判断。我们在两个公开的多模态数据集上演示了WeaveData。
cs.LG / 128 / 2609.34045
Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
Kafila:在可信的一组异构商品化机器上服务大型语言模型
Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya
cs.DC · cs.LG
large language model
大语言模型相关
Abstract
Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to $5.2\times$ against the even split of pipeline parallelism, as in GPipe, and up to $3\times$ against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns $1.56\times$ the throughput of a uniform split and $1.25\times$ of a memory-proportional one, and under four concurrent users that lead compounds to $3.2\times$ rather than fading, each user served at almost the rate of one.
Chinese Translation
他们合起来,一个研究小组或一圈朋友拥有若干台消费级计算机,但没有一台大到足以运行一个有能力的大型语言模型。现有系统在任何人皆可加入的开放集群中汇聚此类容量,而一个仅接纳可信机器的群体无法使用这种方式。限制成员资格移除了它们所依赖的东西:一个集群将模型的每一部分保存在若干对等节点上,并绕过慢节点进行路由。一个有界会话必须使用它接纳的每一台设备。其流水线以收到了它无法快速处理的份额的那台设备的速度推进,因此划分必须在服务开始之前就是正确的。我们提出 Kafila,其协议从 NAT 后方组建一个环,优先选择直接路径,并在穿越失败处进行中继,同时其规划器测量每个设备的内存带宽、容量和可达性,针对固定的环顺序精确划分模型,并将包含嵌入和输出投影的头部与该划分一起放置,而不是事先放置。在三个设备群中能力不同的机器上,从共享局域网到横跨两大洲的五台设备,Kafila 将最慢的流水线阶段最多缩短 $5.2\times$(相较于 GPipe 中流水线并行的均等划分),并最多缩短 $3\times$(相较于 exo 中个人设备推理的按内存比例划分),使 75% 到 87% 的已投入硬件保持工作,而在那些划分下这一比例低于一半,并且能够服务一个任何均匀划分都无法放置到该设备群上的模型。这对用户的价值取决于一个词元中有多少是计算而非网络。在成员共享网络的情况下,同样的划分可达到均匀划分吞吐量的 $1.56\times$,以及按内存比例划分吞吐量的 $1.25\times$;在四个并发用户下,这一领先优势累积到 $3.2\times$ 而非衰减,每个用户几乎以单用户的速率获得服务。
cs.AI / 129 / 2609.34645
Nereus: Adaptive Parallelism for LLM Post-Training
Nereus:面向 LLM 后训练的自适应并行
Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco, Bo Zhao
cs.DC · cs.AI
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27$\times$ over OpenRLHF and by 1.10--1.47$\times$ over Verl across diverse clusters.
Chinese Translation
大语言模型(LLM)的强化学习(RL)后训练会在 GPU 集群上跨生成、推理和训练协调多个模型。在一次运行期间,若干因素可能发生变化,包括资源可用性、序列长度、内存压力和阶段瓶颈。因此,最初合适的执行计划随后可能随时间推移变得缓慢,甚至不可行。然而,对模型共享 GPU 的作业进行自适应调整会带来重大挑战:判断新计划是否值得转换成本、复用该作业的分布式状态,以及协调跨模型和阶段的 GPU 传输。Nereus 作为一个成本感知的运行时来应对这些挑战,它将 RL 后训练作业调整为高效执行计划。其低开销控制器选择一个内存可行的全局计划,并使用根据正在运行的作业校准的成本模型来接纳该转换。为了估计并执行一次转换,Nereus 将模型-阶段(一个阶段中的一个模型)的每个副本的分布式状态表示为一个弹性模型单元。然后,它采用一个全局转换图来排序这些单元的变换和 GPU 传输。在一条基于真实数据构建的跟踪记录中,在线 TP/PP 自适应相对于采用 DP 扩展的初始固定 TP/PP 布局,将平均步延迟降低了 27.7%。在一次达到 1,024 块 GPU 的 1,000 步运行中,六次转换消耗了总运行时间的 0.079%。Nereus 在不同集群上将端到端 8B PPO 吞吐量相较于 OpenRLHF 提升了 2.14--7.27$\times$,相较于 Verl 提升了 1.10--1.47$\times$。
cs.AI / 130 / 2609.35065
TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash
TempoKV:面向内存语义闪存的 LLM KV 缓存适时暂存
Jay H. Park, Hyungjun Kim, Dong Kim
cs.DC · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas immediate staging can reserve fast-tier capacity long before retrieval begins. We present TempoKV, a timing-aware resource-commitment layer that separates early knowledge of reuse from the acquisition of staging resources. It records reusable-KV hits as metadata-only claims and requests commitment when the runtime-estimated time until retrieval falls to the storage-estimated time needed to make KV resident and protected against eviction. These estimates adapt to runtime progress and staging state, while commitment remains subject to available protected capacity. We implement TempoKV in vLLM and LMCache on an SSD-backed CXL memory device without changing request scheduling. Across two models and three prefix cache ratios, TempoKV reduces protected fast-tier byte-time per request by 63-91% versus immediate staging while retaining much of the serving benefit of advance staging. In a fast-tier capacity sweep, output throughput and p95 time to first token (TTFT) remain nearly unchanged as capacity decreases from 100 to 25 GiB. Compared with unmodified LMCache's Device-DAX L1 configuration, TempoKV reduces p95 TTFT by up to 48.0% and increases output throughput by up to 27.8%.
Chinese Translation
在大语言模型(LLM)服务中,可复用的前缀键值(KV)缓存可能超出 GPU 内存容量。内存语义闪存层级以有限的快速层提供由 SSD 支撑的容量,但逻辑上的 KV 命中并不一定已准备好供 GPU 取用。按需暂存会暴露 SSD 延迟,而立即暂存则可能在取用开始很久之前就预留快速层容量。我们提出 TempoKV,一种时序感知的资源承诺层,它将关于复用的早期知情与暂存资源的获取分离开来。它将可复用 KV 命中记录为仅含元数据的声明,并在运行时估计的距取用时间降至存储估计的使 KV 驻留并受保护免遭驱逐所需时间时,请求承诺。这些估计会随运行时进展和暂存状态自适应调整,而承诺仍受可用受保护容量的约束。我们在 vLLM 和 LMCache 中,在由 SSD 支撑的 CXL 内存设备上实现了 TempoKV,且不改变请求调度。在两个模型和三种前缀缓存比例下,与立即暂存相比,TempoKV 将每请求的受保护快速层字节-时间降低 63–91%,同时保留了提前暂存的大部分服务收益。在快速层容量扫描中,当容量从 100 GiB 降至 25 GiB 时,输出吞吐量和 p95 首令牌时间(TTFT)几乎保持不变。与未经修改的 LMCache 的 Device-DAX L1 配置相比,TempoKV 将 p95 TTFT 最多降低 48.0%,并将输出吞吐量最多提高 27.8%。
cs.AI / 131 / 2609.34404
Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems
Eval4DiRec:一个面向基于扩散的推荐系统的统一且系统的评估框架
Cong Wang, Shoujin Wang, Yishuo Li, Qi Zhang, Liang Hu, Wenpeng Lu
cs.IR · cs.AI
diffusion
扩散模型相关
Abstract
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of diffusion-based RSs, a critical issue has emerged: the lack of a unified and systematic quantitative evaluation benchmark, which often results in irreproducible experimental results and unfair comparisons across studies due to inconsistent data processing, training configurations, inference procedures, and evaluation protocols. To address this challenge, we propose Eval4DiRec, the first unified and open-source evaluation framework specifically designed for diffusion-based RSs. Eval4DiRec supports 14 representative diffusion-based RS models across five different recommendation scenarios, providing consistent and reproducible experimental settings to systematically assess their performance. Built upon this framework, we conduct extensive empirical studies to benchmark these models under unified protocols. The results highlight the strong potential of diffusion models for recommendation while also revealing key factors and practical challenges that substantially affect their performance, thereby establishing a solid foundation to facilitate fair evaluation and guide future research in this promising field. Our code and data are available at: https://github.com/wangcong2001/Eval4DiRec.
Chinese Translation
利用扩散模型强大的生成能力和稳定的训练动态,基于扩散的推荐系统(RSs)最近作为一种新型推荐范式出现,吸引了学术界和工业界越来越多的关注。然而,尽管基于扩散的推荐系统迅速发展,一个关键问题已经出现:缺乏统一且系统的定量评估基准,这往往导致由于数据处理、训练配置、推理过程和评估协议不一致,实验结果不可复现且研究之间的比较不公平。为应对这一挑战,我们提出了 Eval4DiRec,这是首个专为基于扩散的推荐系统设计的统一且开源的评估框架。Eval4DiRec 支持五个不同推荐场景中的 14 个具有代表性的基于扩散的推荐系统模型,提供一致且可复现的实验设置,以系统地评估其性能。基于该框架,我们进行了广泛的实证研究,以在统一协议下对这些模型进行基准测试。结果突显了扩散模型在推荐方面的强大潜力,同时也揭示了会显著影响其性能的关键因素和实际挑战,从而为促进公平评估和指导这一有前景领域的未来研究奠定了坚实基础。我们的代码和数据可在以下网址获取:https://github.com/wangcong2001/Eval4DiRec。
cs.LG / 132 / 2609.33838
ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning
ChemOPD:面向多任务化学推理的多教师在线策略蒸馏
Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Xuemin Chen, Tianshu Yu
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.
Chinese Translation
人们日益期待大语言模型能够在单一统一模型中支持多样化的化学推理能力。一种做法是分别发展专门的化学能力,再通过多教师在线策略蒸馏将它们整合起来,但这带来了两个问题:专门化应如何组织,以及专家指导应如何整合?我们提出 ChemOPD,同时解决这两个问题。我们从监督微调梯度中估计任务亲和度,并求解一个带约束的混合整数规划(MIP),以构建部分重叠的专家分组。在蒸馏过程中,我们保留一个在所有任务上训练的通用教师,使专家指导是对其监督的补充而非替代。我们的锚点-残差目标逐步提高被路由的专家在学生生成回复上的贡献。在 ChemCoTBench 上,亲和度引导的专门化相对于通用教师产生了依赖于任务的增益,并在语义任务分组之外提升了若干能力。然而,教师侧更强的性能并不会自动带来更强的学生:在相同的专家和路由下,锚点-残差 OPD 在大多数报告的指标上优于仅使用专家的蒸馏,并实现了可用教师增益中更大的份额。这些结果凸显出,在化学推理中,专门化与能力整合是相互关联但彼此不同的设计问题。
cs.LG / 133 / 2609.33848
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
QwenGyre:一种用于训练超长时域(xLong-Horizon)智能体的弹性强化学习框架
Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang
cs.LG · cs.DC
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.
Chinese Translation
大语言模型(LLM)智能体越来越多地承担极长(xlong)时域任务,其中单次执行可以持续数小时、包含数百次模型--环境交互,并且每次 rollout 接近 1M tokens。将在线强化学习(RL)应用于此类执行会带来两个根本性挑战:(1)严重的执行方差和过长的 rollout 延迟会造成大量 GPU 空闲;以及(2)复杂的非线性分支会产生大量轨迹冗余,严重削弱训练效率。为应对这些问题,我们提出了 QwenGyre,一个用于超长时域在线 RL 的端到端框架。QwenGyre 可在不中断实时执行的情况下,在 rollout 与训练之间弹性重新分配 GPU,同时其轨迹处理器会重建分支历史、对部分进展进行评分,并对冗余路径去重,以限制训练成本。扩展到我们的旗舰模型 Qwen~3.8 2.4T,并在每次 rollout 使用 700K tokens 时,QwenGyre 在 48 步内于 NL2RepoBench 上取得了 6.0% 的绝对提升(52.5% $\to$ 58.5%)。在我们对训练数据集不同领域的评估中,QwenGyre 相比 Colocate 和 Async 分别实现了高达 $1.85\times$ 和 $1.78\times$ 的加速。
cs.LG / 134 / 2609.33857
PI-NOMT: Physics-Informed Neural Optimal Mass Transport for Brain Fluid Dynamics
PI-NOMT:用于脑流体动力学的物理信息神经最优质量输运
Mehmet Emin Acar, Vahit Bugra Yesilkaynak, Helene Benveniste, Gozde Unal
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Recovering hidden transport mechanisms from sparse spatiotemporal observations is a fundamental inverse problem in scientific machine learning. In brain tracer imaging, dynamic contrast-enhanced MRI (DCE-MRI) provides time-resolved measurements of tracer concentration, while the underlying velocity and source mechanisms governing tracer propagation remain unobserved. We formulate this problem as physics-informed latent-state inference, in which the transport field itself is the primary object of inference rather than an auxiliary variable used only to reconstruct observed densities. We propose Physics-Informed Neural Optimal Mass Transport (PI-NOMT), a framework that represents density, velocity, and source as continuous neural fields and combines a continuous neural density teacher, recursive differentiable advection--diffusion--source rollout, unbalanced optimal-transport regularization, and governing-equation supervision. Physical laws act as structural priors that constrain the space of admissible transport mechanisms, while observed tracer dynamics provide evidence for estimating the latent transport state. We evaluate PI-NOMT on a synthetic benchmark with known ground-truth transport and on DCE-MRI sequences from nine control rats. On the synthetic benchmark, PI-NOMT accurately recovers the prescribed velocity field, including its magnitude, direction, and integrated trajectories, rather than merely reconstructing endpoint densities. Across the nine rat datasets, the framework yields sub-percent local endpoint error, consistent physical speed scales, and low post-training PDE and incompressibility residuals. These results support physics-informed latent-state inference as a general framework for recovering hidden transport mechanisms from observed dynamic scalar fields.
Chinese Translation
从稀疏时空观测中恢复隐藏的输运机制是科学机器学习中的一个基本反问题。在脑示踪剂成像中,动态对比增强 MRI (DCE-MRI) 提供示踪剂浓度的时变测量,而支配示踪剂传播的底层速度和源机制仍未被观测到。我们将该问题表述为物理信息潜在状态推断,其中输运场本身是推断的主要对象,而不是仅用于重建观测密度的辅助变量。我们提出物理信息神经最优质量输运(PI-NOMT),一个将密度、速度和源表示为连续神经场,并结合连续神经密度教师、递归可微平流--扩散--源展开、非平衡最优输运正则化和控制方程监督的框架。物理定律作为结构先验约束可容许输运机制的空间,而观测到的示踪剂动力学为估计潜在输运状态提供证据。我们在具有已知真值输运的合成基准上以及来自九只对照大鼠的 DCE-MRI 序列上评估 PI-NOMT。在合成基准上,PI-NOMT 准确恢复规定的速度场,包括其幅值、方向和积分轨迹,而不仅仅是重建端点密度。在九个大鼠数据集上,该框架产生低于百分之一的局部端点误差、一致的物理速度尺度以及较低的训练后 PDE 和不可压缩性残差。这些结果支持物理信息潜在状态推断作为从观测的动态标量场中恢复隐藏输运机制的通用框架。
cs.LG / 135 / 2609.33889
Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding
激活稀疏性与 KV 缓存稀疏性在 LLM 解码中的交汇之处
Jungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo, Heuiseok Lim
cs.LG · cs.CL · cs.PF
large language model
大语言模型相关
Abstract
At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, yet their reported speedups are hard to compare because each depends on context length and on the dense attention kernel it is measured against. We derive a byte crossover, the context length at which the two savings are equal, together with ideal speedup bounds for each branch and for their composition, from model dimensions and keep ratios alone. We then time both branches and their composition from 2K to 128K tokens on two GPUs after a dense prefill of real text, with dense and sparse modes reading the cache through the same split-K attention kernel. The projection branch leads at short context and the KV branch at long context, with speedups that follow their byte bounds up to fixed kernel costs. Adding these costs, measured in separate sweeps, lets the byte account predict the measured crossings of three keep-ratio pairs, a second model, and a second GPU to within 4.1K tokens. Timing the dense baseline with masked instead of split-K attention inflates the apparent speedup of the same KV policy about fivefold. An attention-scored KV selection answers the same passkey and multi-key placements as dense decoding up to 127K tokens, whereas a KV window misses most of them. Under matched perplexity budgets, activation sparsity composed with this selection decodes 14 to 26% faster than the best single branch on both GPUs. Code is available at https://github.com/js-lee-AI/ByteCross.
Chinese Translation
在每一步中,用大语言模型解码一个序列都会重新读取投影权重(其流量是固定的)以及键值(KV)缓存(其流量随上下文增长)。激活稀疏性削减第一项,KV 缓存稀疏性削减第二项,然而二者所报告的速度提升难以比较,因为二者各自都取决于上下文长度,也取决于其所对照测量的稠密注意力核。我们仅根据模型维度和保留率,推导出一个字节交叉点,即两种节省相等时的上下文长度,以及每个分支及其组合的理想加速上界。随后,我们在真实文本的稠密预填充之后,在两块 GPU 上对两个分支及其组合从 2K 到 128K token 进行计时,其中稠密模式与稀疏模式都通过同一个 split-K 注意力核读取缓存。投影分支在短上下文下领先,KV 分支在长上下文下领先,其加速比在固定的核开销范围内遵循各自的字节上界。将这些在单独扫描中测得的开销相加,使字节账目能够将三个保留率对、第二个模型和第二个 GPU 的实测交叉点预测到 4.1K token 以内。用掩码注意力而非 split-K 注意力对稠密基线进行计时,会使同一 KV 策略的表观加速比虚增约五倍。一种注意力打分的 KV 选择在高达 127K token 时对相同的 passkey 和多密钥放置给出与稠密解码相同的答案,而 KV 窗口会错过其中大多数。在匹配的困惑度预算下,与这种选择组合的激活稀疏性在两块 GPU 上都比最佳单一分支解码快 14% 到 26%。代码见 https://github.com/js-lee-AI/ByteCross。
cs.LG / 136 / 2609.33930
Diffusion-Based Rollouts as a Stabilization Mechanism for Long-Horizon Environmental Forecasting
基于扩散的展开作为一种长时程环境预报的稳定机制
Marina Vicens-Miquel, Amy McGovern, Aaron J. Hill, Efi Foufoula-Georgiou, Samuel S. P. Shen
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Extending forecast lead times while maintaining predictive skill remains a major challenge in environmental forecasting. We investigate diffusion-based rollouts as a stabilization mechanism for recursive forecasting using low-dimensional water-level time series and high-dimensional precipitation fields. Across both modalities, diffusion suppresses recursive error growth, with the largest stabilization occurring where deterministic rollouts are most unstable. However, stabilization does not guarantee forecast fidelity. In the water-level experiments, forecasts progressively lose event-level fidelity as the rollout loses access to external predictive information, and trajectory-level comparisons show that diffusion can remain numerically stable while contracting toward central values and exhibiting reduced variability. In the precipitation experiments, which retain conditioning from numerical weather prediction throughout the rollout, diffusion better preserves spatial organization and event-detection skill. Together, these contrasting experiments indicate that diffusion can control recursive error amplification, while its practical benefit also depends on the predictive information available to constrain future evolution.
Chinese Translation
在保持预测技巧的同时延长预报提前时间,仍然是环境预报中的一大挑战。我们研究基于扩散的展开作为一种用于递推预报的稳定机制,所用数据包括低维水位时间序列和高维降水场。在两种模态中,扩散都抑制了递推误差的增长,且最大的稳定作用出现在确定性展开最不稳定的地方。然而,稳定并不能保证预报的保真度。在水位实验中,随着展开失去对外部预测信息的获取,预报逐渐丧失事件级别的保真度;而轨迹级别的比较表明,扩散可以在数值上保持稳定,同时向中心值收缩并表现出变异性降低。在降水实验中,整个展开过程始终保持来自数值天气预报的条件约束,扩散更好地保持了空间组织和事件检测技巧。总之,这些形成对比的实验表明,扩散可以控制递推误差的放大,而其实际收益还取决于可用于约束未来演变的预测信息。
cs.LG / 137 / 2609.33977
GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning
GroupMask:面向半结构化 LLM 剪枝的层自适应分组稀疏性
Zhengao Li, Shuoqiu Li, Xiaofang Zhang, Yukai Jin, Gokcen Kestor, Yanfu Zhang, Yiming Zeng, Bin Ren, Chuxu Zhang, Shangqian Gao
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same $1\times256$ group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.
Chinese Translation
半结构化剪枝在压缩大语言模型(LLM)的同时保持规则的稀疏结构,但主流的 N:M 模式在每一层中都固定相同的局部稀疏率。层自适应稀疏分配能够改进非结构化剪枝,然而已有报告指出其在 N:M 稀疏下效果较差,这使得以下问题仍悬而未决:自适应分配对半结构化剪枝的价值有限,究竟是普遍如此,还是仅在细粒度的 N:M 模式下如此。我们以组级稀疏性来考察这一问题,它将每个权重矩阵划分为规则的组,以整组为单位保留或剪除每个组,并允许每一层的稀疏率在全局预算下变化。我们提出 GroupMask,它用一个轻量级超网络生成所有层的组选择器,通过 Gumbel-Sigmoid 参数化和直通估计器对其进行松弛,并在保持预训练权重冻结的同时,通过稀疏预算正则化和自蒸馏来学习它们。在 50% 稀疏率、相同 $1\times256$ 组大小的 LLaMA-2-7B 上,相对于均匀的逐层比例,学习到的层自适应分配将 WikiText-2 困惑度从 10.02 降至 8.30,并将平均零样本准确率从 0.455 提升至 0.496。在五个 LLaMA 和 Qwen 模型上,在受评估的基线方法中,GroupMask 在 LLaMA-2-7B 上取得了最低的 WikiText-2 困惑度,并在使用 Alpaca 校准的情况下取得了最高的平均零样本准确率。我们的代码可在 https://github.com/ZhengaoLi/GroupMask 获取。
cs.LG / 138 / 2609.34009
Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
用于 LLMs 基于反馈的在线策略自蒸馏的 Fisher 信息引导重校准
Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton
cs.LG
large language model
大语言模型相关
Abstract
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.
Chinese Translation
基于反馈的在线策略自蒸馏已成为一种有前景的方法,它使基础模型——更具体地说是大语言模型(LLMs)——能够在外部反馈下从自身的输出中学习,其中单个模型同时充当教师和学生。然而,这类方法可能表现出不稳定的优化,容易导致训练过程中的性能崩溃。为解决这一局限,我们提出 FIRE(Fisher 信息重校准,Fisher-Informed REcalibration),这是一个双分支框架,在微调过程中重新校准施加于正确与错误的在线策略输出上的监督信号。对于正确的回答,FIRE 用重新加权的在线策略 SFT 替代自蒸馏;而对于错误的回答,FIRE 识别出那些对教师诱导更新产生不成比例影响的反馈成分,并据此重新校准以反馈为条件的目标。两个分支都受到一个 token 级半径的影响,该半径部分地由 softmax Fisher 迹导出。FIRE 将反馈应当把模型推向哪个方向与模型在该方向上应当移动多远分离开来,同时保持表现良好的反馈监督不变。我们的实验表明,FIRE 提供了显著更稳定的自蒸馏,同时保持强劲的下游性能,尤其是在标准反馈条件蒸馏变得不稳定的情形下。
cs.LG / 139 / 2609.34012
When Known Physics Helps Neural PDE Models: Residual Constraints Out-Regularize Generic Priors for Nonlinear Dynamics
当已知物理有助于神经 PDE 模型时:残差约束对非线性动力学的正则化效果优于通用先验
Zahra Farazpay, Aniruddha Bora
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Neural PDE surrogates increasingly incorporate structural priors, yet it is often unclear whether their gains arise from physics-specific information or simply from regularization and training choices. We evaluate several such priors under a common protocol against a matched from-scratch neural operator baseline. Our central result is that a known-equation residual consistently outperforms the best generic regularizer at equal tuning budget. At fixed capacity this benefit appears across linear and nonlinear PDEs, but a capacity sweep reveals a sharp distinction: the advantage persists and grows for Burgers, KdV, and Allen-Cahn, while collapsing toward or below parity for linear heat and advection-diffusion. Thus, the durable value of the residual is specific to nonlinear operators. We further falsify a pre-registered hypothesis that the benefit is activated only by data sparsity: the residual remains advantageous even under full supervision. Its usefulness does, however, have a clear boundary. Under grid under-resolution, nonlinear coarse fields no longer satisfy the naive governing-equation residual, and enforcing it becomes actively harmful. In contrast, cross-family pretraining and in-context conditioning fail to outperform the strong from-scratch baseline in the regime studied. Together, these results identify when known physics provides non-redundant information to neural PDE models, when it does not, and when enforcing it introduces bias.
Chinese Translation
神经 PDE 代理模型越来越多地引入结构先验,然而其收益究竟源自物理特有的信息,还是仅仅源自正则化与训练选择,往往并不清楚。我们在统一协议下评估了若干此类先验,并与相匹配的从零训练的神经算子基线进行对比。我们的核心结果是:在相同的调参预算下,已知方程残差始终优于最佳的通用正则化器。在容量固定时,这一收益在线性与非线性 PDE 中均会出现,但容量扫描揭示出一种鲜明区别:对于 Burgers、KdV 和 Allen-Cahn,该优势持续存在并不断扩大,而对于线性热方程与对流扩散方程,该优势则退化至与基线持平甚至低于基线。因此,残差的持久价值是非线性算子所特有的。我们进一步否证了一项预注册假设,即该收益仅由数据稀疏性所激活:即便在完全监督下,残差仍然具有优势。然而,其有用性确实存在明确的边界。在网格分辨率不足的情况下,非线性粗粒度场不再满足朴素的控制方程残差,此时强行施加该约束反而会产生有害作用。相比之下,在所研究的范围内,跨族预训练与上下文内条件化均未能超越这一强大的从零训练基线。综合来看,这些结果明确了已知物理何时为神经 PDE 模型提供非冗余信息、何时不提供,以及何时施加它会引入偏差。
cs.LG / 140 / 2609.34058
Do World Models Learn Global Understanding?
世界模型是否学习全局理解?
Alexander Detkov, Matt Thomson
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.
Chinese Translation
AI 系统常常显得脆弱且碎片化。一个大型语言模型(LLM)可能正确地解释一个概念,却无法应用它,或者在一个情境中遵循安全指令,而在另一个情境中则不能。这种行为表明了一种普遍失败:无法将局部信息提升为全局理解。为了获得根本性洞见,我们将“理解”界定为学习约束并传播其后果。我们在幺半群世界(monoid worlds)上构建学习任务,这些世界是由动作转移连接起来的状态集合,其中观察到的训练转移和一个未见约束共同决定留出转移。衡量泛化能力检验模型能否从局部转移中学习全局约束并传播其后果。我们考虑与空间和语义结构相关的逆、交换、复合和周期性约束。在注意力、循环和状态空间架构中,下一状态训练能够拟合数据,但无法传播非平凡约束。复合训练使用相同的路径,但从输入中隐藏中间状态,在跨架构的逆、交换和复合约束上达到 96% 的准确率,并在具身环境上训练的世界模型的几何泛化以及 Wikidata 微调 LLM 的关系泛化中产生相应改进。当推断一个未见事实可能依赖于首先推断其他事实时,模型能将约束传播多远?我们定义留出转移的证明深度 d,衡量推断该转移所需的最少推理轮数,并发现模型泛化能力随证明深度增加而急剧下降。增加复合路径长度 T 可改善泛化。这些结果为研究语言和世界模型中的全局理解提供了一种形式化方法,并证明复合训练促进信息传播与整合。
cs.LG / 141 / 2609.34064
Learning Perturbation Robust Policies for LLM Agents with Stable Optimization
为LLM智能体学习具有稳定优化的扰动鲁棒策略
Pengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang, Jingdi Chen
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve perturbation robustness during policy optimization. We first introduce the notion of a perturbation robust policy and analyze conditions under which perturbed policy updates preserve stable monotonic improvement. Based on this analysis, we introduce Stable Perturbation-Robust Policy Optimization (SPrPO), which applies adaptive and sensitivity-aware perturbations during RL training. We evaluate SPrPO on ALFWorld and WebShop and conduct systematic experiments across multiple perturbation types and scales, showing improved perturbation robustness while maintaining stable policy optimization.
Chinese Translation
强化学习(RL)已成为面向长时程大型语言模型(LLM)智能体的一种有效后训练范式。然而,我们发现,所得策略可能对各种策略扰动敏感,例如隐状态噪声、剪枝和量化。在这项工作中,我们研究如何在策略优化过程中提升扰动鲁棒性。我们首先引入扰动鲁棒策略的概念,并分析在何种条件下受扰动的策略更新能够保持稳定的单调改进。基于这一分析,我们提出稳定扰动鲁棒策略优化(SPrPO),其在RL训练过程中施加自适应且敏感度感知的扰动。我们在ALFWorld和WebShop上评估SPrPO,并跨多种扰动类型和尺度开展系统性实验,结果表明其在保持稳定策略优化的同时提升了扰动鲁棒性。
cs.LG / 142 / 2609.34172
GPARA: Graph-Posterior-Aligned Refinement and Active Acquisition for Grounding Diffusion Priors
GPARA:面向扩散先验接地的图-后验对齐精炼与主动采集
Wangqian Chen, Hao Wang, Yumeng Zhang, Jiajia Guo, Junting Chen, Jun Zhang
cs.LG
diffusion
扩散模型相关
Abstract
Active grounding of a frozen diffusion prior requires jointly determining where new measurements should be taken and how they should be used to refine the current reconstruction. Posterior-ensemble-based methods can estimate acquisition utility from generated samples, but require repeated ensemble generation as observations accumulate and capture posterior geometry only through empirical statistics. This paper proposes GPARA, which learns a context-dependent graph surrogate over diffusion prediction residuals, inducing an explicitly reusable posterior response operator that propagates measurement innovations to unobserved variables and evaluates candidate measurements through weighted posterior-risk reduction. Under the matched surrogate, we show that the same response operator also determines expected one-step acquisition benefit and yields an analytic ranking consistent with expected reconstruction improvement. A bounded learned residual calibrates the analytic utility to account for surrogate mismatch, while a small prior ensemble is generated once and reconditioned to update risk weights without repeated diffusion posterior sampling during acquisition. Experiments on two reconstruction tasks spanning physical field and computer vision show consistent improvements in refinement and active acquisition over the evaluated baselines. Ablations further support the complementary roles of step-wise graph refinement, adaptive risk weighting, and analytically anchored calibration.
Chinese Translation
对冻结扩散先验的主动接地,需要共同确定应当在何处获取新的测量,以及应如何利用这些测量来精炼当前的重建结果。基于后验集成的方法能够从生成的样本中估计采集效用,但随着观测不断累积,它们需要重复进行集成生成,并且仅通过经验统计量来刻画后验几何结构。本文提出 GPARA,它在扩散预测残差上学习一个依赖于上下文的图代理模型,从而诱导出一个显式可复用的后验响应算子,该算子将测量新息传播至未观测变量,并通过加权后验风险降低来评估候选测量。在匹配的代理模型下,我们表明同一个响应算子也决定了期望的单步采集收益,并给出与期望重建改进相一致的分析排序。一个有界的学习残差对分析效用进行校准,以考虑代理模型失配,同时一个小规模先验集成仅生成一次,并通过重新条件化来更新风险权重,而无需在采集过程中重复进行扩散后验采样。在涵盖物理场与计算机视觉的两项重建任务上的实验表明,相较于所评估的基线,该方法在精炼与主动采集方面均取得了一致的改进。消融实验进一步支持了逐步图精炼、自适应风险加权以及以分析为基础的校准三者所起的互补作用。
cs.LG / 143 / 2609.34185
EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
EntroPack:在任意比特率下快速且准确的熵编码权重压缩
Hong Zhang, Zhongjie Duan, Yingda Chen
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility requires accurate rate selection and efficient weight reconstruction for inference. We present EntroPack, an entropy-coded weight compressor that supports arbitrary target bitrates without activation calibration or fine-tuning. It combines row-normalized $E_8$ lattice quantization with a conditional probability model of lattice coordinates. Sampled storage estimates select the quantization resolution without repeated full-stream encoding. The final coordinates are entropy-coded in independently decodable tiles, enabling fast, fused symbol decoding and numerical weight reconstruction on the GPU. EntroPack supports floating-point and integer weight containers, such as BF16, FP16, FP8, and INT8, with storage bitrate controlled independently of numerical precision. Online decoding adds latency that grows with weight count, making the method well suited to compute-intensive workloads such as diffusion denoising and Transformer prefill. Experiments demonstrate fast encoding and modest inference overhead in these settings. When compressing the linear-layer weights of the image generator Z-Image-Turbo, EntroPack achieves substantially lower weight and denoiser output errors than fixed-width formats at comparable storage rates, with modest denoising-step overhead. Targeting 4 bits per parameter, it achieves lower weight and denoiser output errors than NF4, including about 24% lower relative $L_2$ weight error, with less storage. Source code is available at https://github.com/modelscope/entropack.
Chinese Translation
权重压缩有助于大型神经网络适应部署时的内存预算,但常见的固定宽度格式只能提供粗粒度的存储选择。熵编码支持更精细的比特率,然而实际达到的大小取决于量化后的权重分布和编码开销。利用这种灵活性需要准确的码率选择以及用于推理的高效权重重建。我们提出 EntroPack,一种支持任意目标比特率的熵编码权重压缩器,无需激活校准或微调。它将对行归一化的 $E_8$ 格量化与格坐标的条件概率模型相结合。采样的存储估计在不进行重复全流编码的情况下选择量化分辨率。最终坐标在可独立解码的瓦片中进行熵编码,从而能够在 GPU 上实现快速的融合符号解码和数值权重重建。EntroPack 支持浮点和整数权重容器,例如 BF16、FP16、FP8 和 INT8,其存储比特率独立于数值精度进行控制。在线解码会增加延迟,且该延迟随权重数量增长,这使得该方法非常适合计算密集型工作负载,例如扩散去噪和 Transformer 预填充。实验表明,在这些设置下编码快速且推理开销适中。在压缩图像生成器 Z-Image-Turbo 的线性层权重时,EntroPack 在可比的存储比特率下,相比固定宽度格式实现了显著更低的权重误差和去噪器输出误差,且去噪步骤开销适中。以每参数 4 比特为目标时,它相比 NF4 实现了更低的权重误差和去噪器输出误差,包括约 24% 更低的相对 $L_2$ 权重误差,同时占用更少存储。源代码可在 https://github.com/modelscope/entropack 获取。
cs.LG / 144 / 2609.34205
Learning to Optimize through Solver-Grounded Self-Play
通过基于求解器的自我对弈学习优化
Xia Jiang, Yaoxin Wu, Chenyu Zhou, Mengzhu Xu, Wim P. M. Nuijten, Yingqian Zhang
cs.LG
large language model
大语言模型相关
Abstract
Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This dependence introduces a Generalization Ceiling, where models overfit to narrow data distributions, and Capability Anchoring, where models' reasoning is bounded by annotator proficiency and teacher model capability. In response, we propose OPT-Zero, the first fully self-play training framework for optimization modeling that requires zero external training data. OPT-Zero employs a single LLM in a dual-role closed loop: a Proposer that synthesizes increasingly challenging optimization problems alongside their mathematical formulations and solving code, and a Solver that attempts to resolve the problems given only natural-language problem descriptions. Grounded in execution feedback from external optimization solvers, we alternately train both roles using reinforcement learning. This process fosters an auto-curriculum in which the Proposer and Solver co-evolve: generating harder valid problems by the Proposer seamlessly enhances the structural reasoning ability of the Solver. Extensive results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
Chinese Translation
优化建模是许多决策场景的核心,但传统上需要大量的领域专业知识。尽管大语言模型(LLMs)在自动化这一过程方面已展现出前景,但当前的训练范式主要依赖人工标注或教师模型生成的数据集。这种依赖引入了泛化天花板(Generalization Ceiling),即模型会过拟合于狭窄的数据分布;以及能力锚定(Capability Anchoring),即模型的推理能力受限于标注者的熟练程度与教师模型的能力。为此,我们提出了 OPT-Zero,这是首个面向优化建模、无需任何外部训练数据的完全自我对弈训练框架。OPT-Zero 在一个双角色闭环中使用单个 LLM:一个出题者(Proposer),它合成日益具有挑战性的优化问题及其数学形式化表述与求解代码;以及一个求解者(Solver),它仅根据自然语言的问题描述来尝试求解这些问题。以来自外部优化求解器的执行反馈为依据,我们使用强化学习交替训练这两个角色。这一过程促成了一个自动课程(auto-curriculum),其中出题者与求解者共同演化:出题者生成更难的合法问题,会无缝地增强求解者的结构化推理能力。大量结果表明,在零人工整理数据的情况下,OPT-Zero 达到了与最先进的依赖数据的方法相当的水平,同时展现出显著更强的泛化能力,从而确立了自我对弈训练作为一种高度可扩展的范式,用以推进 LLM 在优化问题建模与求解方面的推理能力。
cs.LG / 145 / 2609.34228
SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals
SleuthBench:使用表格隐藏信号对统计大语言模型评估进行基准测试
Jingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez, Joshua Shunk, Rich Caruana, Ben Lengerich
cs.LG
large language model
大语言模型相关
Abstract
Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.
Chinese Translation
评估大语言模型(LLM)智能体的统计发现能力需要可验证的分析性真值。为真实世界数据集建立这样的真值成本高昂,而且对公开数据集的先验知识会影响智能体的回答。我们提出 SLEUTHBENCH,一个通过向公开表格数据集注入受控的数据质量问题和特征效应来解决这两个问题的基准:注入的模式决定了答案,因此参考答案可自动计算,而对原始表格的记忆知识是不充分的,同时表格保持其背景结构。注入的模式是仿照真实数据分析中所报告的现象建模的。该基准定义了两大类共 17 个问题模板:数据质量问题与特征贡献问题。我们评估了六个使用 Python 编码工具分析数据的最先进 LLM,在 70 个经过验证的数据集-模板组合的数据科学表述和商业表述上进行测试,总共产生 1680 条评分回答。这些模型能够可靠地检测数据质量问题(准确率 83.8%),但在恢复特征贡献方面表现很差(41.9%)。要发现特征如何塑造目标,需要在候选变量和分析流程两方面进行搜索。为解决这一问题,我们提出经验层(Empirical Layer),一组预先计算的统计产物,包括摘要、拟合的特征效应与交互效应以及数据集描述,它将候选模式暴露出来以供直接检查。获取这些产物可将特征贡献准确率从 41.9% 提升到 68.0%。
cs.LG / 146 / 2609.34246
CasEm: A Cascade Architecture for Long-Horizon Neural Emulation
CasEm:一种用于长时程神经仿真的级联架构
Zhaoyi Li, Jingtao Ding, Shihua Li
cs.LG
diffusion
扩散模型相关
Abstract
Autoregressive neural emulators can drift or diverge over long rollouts despite accurate short-term predictions. We introduce Cascaded Emulation (CasEm), a one-way rollout architecture that augments an existing full-state backbone with an independently evolving model of physically specified aggregates. Its forecasts guide corrections to full-state predictions, without feedback from the backbone to the aggregate model. Effective guidance requires aggregates that cover substantial backbone error, remain accurately predictable, and support useful full-state corrections. We derive a finite-horizon error bound that clarifies these three factors and use empirical diagnostics to guide subsystem selection. Across four ODE/PDE benchmarks, CasEm reduces long-horizon rollout errors across diverse backbones and suppresses the trend toward error divergence in both diffusion tasks using Fourier neural operator backbones. In global climate emulation, CasEm with a regional total-water subsystem reduces 10-year full-state time-mean error by 66.6% and 46.3% for frozen ACE and Spherical DYffusion backbones, respectively, while adding less than 3% to inference time.
Chinese Translation
自回归神经仿真器尽管短期预测准确,但在长时程推演中可能会漂移或发散。我们提出级联仿真(Cascaded Emulation, CasEm),一种单向推演架构,它用一个独立演化的、由物理指定的聚合量模型来增强现有的全状态主干模型。其预测引导对全状态预测的校正,且没有从主干模型到聚合量模型的反馈。有效引导要求聚合量能够覆盖主干模型的显著误差、保持可准确预测,并支持有用的全状态校正。我们推导了一个有限时程误差界,阐明了这三个因素,并使用经验诊断来指导子系统选择。在四个 ODE/PDE 基准上,CasEm 在多种主干模型上降低了长时程推演误差,并在两个使用傅里叶神经算子主干的扩散任务中抑制了误差发散趋势。在全球气候仿真中,采用区域总水量子系统的 CasEm 分别使冻结的 ACE 和 Spherical DYffusion 主干的 10 年全状态时间平均误差降低了 66.6% 和 46.3%,同时使推理时间增加不到 3%。
cs.LG / 147 / 2609.34253
DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models
DreamingGoose:从自回归 Transformer 到双向循环扩散语言模型的分阶段蒸馏
Julian Boesch, Andrew Wee, Alexander Stranzl
cs.LG · cs.CL
diffusion
扩散模型相关
Abstract
Pretrained autoregressive Transformers represent a large sunk investment in compute. Existing conversion methods reuse that investment by changing either the architecture (attention to recurrence) or the objective (next-token prediction to denoising), never both. We convert Qwen3 teachers at 1.7B and 8B into attention-free, bidirectional, gated-delta-rule diffusion students in three stages, so that each capability can be traced to the stage that kept or lost it. Language modeling transfers only partially and in-distribution; in-context retrieval does not transfer. On a multi-query recall probe where the teachers score 0.34-0.58, both converted students score 0.000, and diffusion pretraining alone does not restore retrieval. A retrieval curriculum in the final stage, which gradually lengthens the gap between a key-value table and the queries that address it, restores it only stochastically: on a fixed schedule, one seed in three learns to retrieve. Advancing the gap only while a running accuracy estimate stays above a threshold works for all three of those seeds, holds on real text, and carries unchanged to 8B, where two of three seeds succeed. The third had not learned within its fixed 16k-step budget: retrieval switches on abruptly at a seed-dependent step (6.5k and 11k in the other two), so a fixed budget can cut a late run off. One boundary survives every intervention: every model that learns retrieval scores 0.000 on tokens that never appeared in a retrieval episode, and an arm that resamples the key and value tokens every batch shows this is a coverage limit, not memorization of particular bindings. Separately, we convert a 7B code model into a 3:1 recurrent-attention block-diffusion hybrid over 85k steps and report two negative training results.
Chinese Translation
预训练自回归 Transformer 代表了在计算上的巨大沉没投资。现有转换方法通过改变架构(从注意力到循环)或目标(从下一 token 预测到去噪)来重用该投资,但从不两者同时改变。我们将规模为 1.7B 和 8B 的 Qwen3 教师模型分三个阶段转换为无注意力、双向、门控 delta 规则的扩散学生模型,使每种能力都能追溯到保留或丢失它的阶段。语言建模仅部分迁移,且仅在分布内迁移;上下文内检索不会迁移。在一个多查询召回探针上,教师模型得分为 0.34-0.58,两个转换后的学生模型得分均为 0.000,并且仅靠扩散预训练不能恢复检索。最后阶段的检索课程逐渐拉大键值表与访问它的查询之间的间隔,但只能随机地恢复检索能力:在固定调度下,三个种子中有一个学会检索。仅在运行准确率估计保持高于阈值时才推进间隔,对这三个种子都有效,在真实文本上成立,并原样迁移到 8B,其中三个种子中有两个成功。第三个在其固定的 16k 步预算内尚未学会:检索会在依赖于种子的步数处突然开启(另外两个分别在 6.5k 和 11k 步),因此固定预算可能会截断较晚的运行。有一个边界在每次干预后都依然存在:每个学会检索的模型在从未出现在检索片段中的 token 上得分均为 0.000,而一个每个批次都重新采样键和值 token 的实验分支表明,这是覆盖范围限制,而不是对特定绑定的记忆。另外,我们将一个 7B 代码模型转换为 3:1 的循环-注意力块扩散混合模型,经过 85k 步,并报告两个负面的训练结果。
cs.LG / 148 / 2609.34301
One Sequence, Many Decodings: CAGenMol-2 Recasts Drug Design as Masked Molecular Inference
一个序列,多种解码:CAGenMol-2 将药物设计重构为掩码分子推断
Yanting Li, Enyan Dai, Lei Wang, Wen-Cai Ye, Li Liu
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Drug design couples property evaluation, conditional generation, structure-based design, and local optimization, yet machine learning systems typically address these capabilities with separate task-specific models. We introduce CAGenMol-2, a masked diffusion molecular language model that represents molecules, continuous scalar properties, and 3D protein pockets within a single wrapped sequence. Within this pretrained interface, downstream operations are selected by which sequence regions are observed or masked at inference, allowing one checkpoint to perform property prediction, property- and pocket-conditioned generation, and partial-constraint design without task-specific architectures or backbone fine-tuning. We further propose Adaptive Fragment Optimization (AdaFO), a gradient-free mask-and-refill search that turns the masked decoder into an iterative local molecular optimizer. On CrossDocked2020, AdaFO increases Success Rate from 30.2\% to 70.8\%, the best reported under this protocol, while largely preserving drug-likeness and diversity. Finally, scaffold-preserving directional editing and CRBN/VHL case studies demonstrate its use in compound design workflows spanning local molecular editing, structure-based prioritization, and downstream simulation-based screening.
Chinese Translation
药物设计将性质评估、条件生成、基于结构的设计和局部优化耦合在一起,然而机器学习系统通常以各自独立的、面向特定任务的模型来处理这些能力。我们提出 CAGenMol-2,一种掩码扩散分子语言模型,它在单个包裹序列中表示分子、连续标量性质和 3D 蛋白质口袋。在这个预训练接口内,下游操作取决于在推理时哪些序列区域被观测或掩码,使得单个检查点无需特定任务架构或骨干微调即可执行性质预测、以性质和口袋为条件的生成以及部分约束设计。我们进一步提出自适应片段优化(AdaFO),一种无梯度的掩码-填补搜索,它将掩码解码器转变为迭代式局部分子优化器。在 CrossDocked2020 上,AdaFO 将成功率从 30.2\% 提高到 70.8\%,这是该协议下报告的最佳结果,同时大体上保持了类药性和多样性。最后,保持骨架的方向性编辑以及 CRBN/VHL 案例研究表明,它可用于涵盖局部分子编辑、基于结构的优先级排序以及下游基于模拟的筛选的化合物设计工作流。
cs.LG / 149 / 2609.34326
Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions
无需嵌入的路由:基于正则表达式的快速且可解释路由
Yifan Lu, Qiyue Zhang, Haotian Shan, Hanjie Chen, Jiarong Xing
cs.LG
large language model
大语言模型相关
Abstract
Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture the query properties needed for routing. We therefore investigate which properties matter and whether they can be extracted directly from text without a neural encoder. We introduce REGEXROUTE, a pipeline that uses sparse autoencoders (SAEs) to discover interpretable regular-expression (regex) features. Using unlabeled text, an LLM turns descriptions of grouped SAE latents into regex extractors and refines them to match latent activation patterns. These extractors supply numerical features to a lightweight routing head, eliminating neural encoding at inference (Figure 1a). Across four benchmarks, one fixed set of 128 features achieves 76.43% average routing accuracy, comparable to 76.41% for the strongest neural text encoder baseline, with much smaller latency and strong robustness. These findings establish explicit, interpretable text features as a practical basis for designing and understanding LLM routers.
Chinese Translation
大语言模型(LLM)路由器通常依赖神经查询嵌入,而更大的编码器被期望能更好地捕捉查询意图和难度。然而,将 Qwen2.5 编码器从 0.5B 扩展到 72B 参数,对路由准确率带来的提升很小(图 1b),这表明小型编码器可能已经捕捉到了路由所需的查询属性。因此,我们研究哪些属性至关重要,以及它们能否在没有神经编码器的情况下直接从文本中提取。我们提出 REGEXROUTE,一个使用稀疏自编码器(SAEs)来发现可解释正则表达式(regex)特征的流程。利用无标注文本,LLM 将分组 SAE 潜在变量的描述转化为正则表达式提取器,并对其进行细化,以匹配潜在激活模式。这些提取器为轻量级路由头提供数值特征,从而在推理时消除了神经编码(图 1a)。在四个基准测试上,一组固定的 128 个特征达到了 76.43% 的平均路由准确率,与最强神经文本编码器基线的 76.41% 相当,同时具有小得多的延迟和很强的鲁棒性。这些发现确立了显式、可解释的文本特征作为设计和理解 LLM 路由器的实用基础。
cs.LG / 150 / 2609.34344
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
学习引导,引导以见:通过可训练向量揭示大型语言模型中 RLVR 的几何结构
Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang, Siye Wu, Pengyuan Wang, Jiaxuan Wang, Weijie Liu, Saiyong Yang, Guangzhong Sun, Guiquan Liu, Junfeng Fang
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training
Chinese Translation
强化学习(RL)已成为增强大型语言模型推理能力的关键范式,然而参数更新的高维性使其训练动力学难以分析。我们研究具有可验证奖励的强化学习(RLVR),并使用向量引导来识别激活空间中与 RL 带来的收益相关的低维有效流形。我们揭示两个几何性质。(1)有效流形容量:复现 RL 收益所需的容量可以非常小,但并非可无限压缩;在极低容量下,干预维度和输入依赖的表达能力成为关键约束,并且这一需求随注入深度而变化。(2)控制流形分离:有效控制方向主要位于激活主成分子空间的低方差补空间中。在同一任务和基础模型内,所学几何结构在不同训练配置下基本保持一致,而在不同任务之间,几何对齐与能力迁移相关。在 5 个 LLM 和 6 个可验证奖励任务上的实验支持了这些发现。然后我们提出 Alpha-Stabler,一个即插即用的框架,其中包含一个 Predictor,用于监测主成分子空间侵入以提供早期崩溃预警;以及一个 Controller,在反向传播过程中移除激活梯度的主成分子空间分量,同时保留正交补空间。Alpha-Stabler 在 2,000 步内稳定训练,并持续提升 RL 收益,为稳健的后期训练提供了实践洞见。代码:https://github.com/caiyuchen-ustc/On_Policy_Vector_Training
cs.LG / 151 / 2609.34406
Unlocking Few-Step Diffusion for Faithful Previews
解锁少步扩散以实现忠实预览
Jing Jia, Sifan Liu, Guanyang Wang
cs.LG · cs.CV · stat.ML
diffusion
扩散模型相关
Abstract
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
Chinese Translation
在扩散工作流中,采样延迟会不断累积,用户会生成并丢弃许多候选样本,最后才保留一个。令人惊讶的是,标准少步采样器糟糕的输出并不反映其缺乏重建能力:仅通过优化初始噪声,冻结的 3-4 步采样器就能紧密复现其对应的全步输出。基于这一发现,我们使用端点监督学习对初始噪声和去噪更新的修正,从而改善与由相同噪声和提示生成的全步输出之间的一致性。由此得到的预览使用户能够以低成本筛选候选样本,并将全步生成保留给有前景的样本。输入修正还可以在不重新训练的情况下跨采样预算迁移。实验表明,参考保真度获得了显著提升,包括在无条件基准上比重训练的 LD3 降低 53-78% 的重建 MSE,同时在 SD1.5、SDXL 和 FLUX.1-dev 上改善了排序保持与候选选择。
cs.LG / 152 / 2609.34415
PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety
PROACT-Agent:面向实时安全的渐进式运行时监督与主动熔断
Ding Jia, Wei Liu, Xianglong Du, Yingjie Li, Yingqing Yang, Huili Yu, Zhangsong Zhan, Chu Zhou
cs.LG
large language model
大语言模型相关
Abstract
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.
Chinese Translation
从大型语言模型(LLMs)向智能体的转变,将安全风险从有害文本转移到不可逆的环境危害。尽管当前的防御措施在很大程度上仍是回溯性的,但主动的运行时干预因缺乏大规模、因果一致的数据而遭遇瓶颈。我们提出 PROACT-Agent,一个用于合成高保真轨迹以支持实时护栏的框架。我们识别出先前基准中的关键“安全漂移”,其中宽松的标注范式未能保证时序一致性。PROACT-Agent 通过以下方式解决这一问题:(1)渐进式轨迹展开,以揭示隐藏在长上下文交互中的风险;(2)推理增强的因果校正,以强制实现单调因果一致性;以及(3)文化感知的数据本地化,以实现跨境鲁棒性。我们提出 PROACT-Bench,一个包含 155,780 个状态的双语安全基准,这些状态通过多模型裁决进行标注。在完整源留出条件下,通过在下一轮 LLM 推理之前评估更新后的上下文,训练后的防护器实现了 91.46% 的不安全类 F1 和 90.63% 的精确边界检测。在 AgentDojo 中,它将非 DoS 定向攻击成功率从 20.82% 降低到 0.40%。
cs.LG / 153 / 2609.34431
Harmonizing Spectral Evolution in Conditional Flow Matching for TTS
在面向 TTS 的条件流匹配中协调频谱演化
Isha Pandey Varad Deshpande Abhijat Bharadwaj Ganesh Ramakrishnan
cs.LG · cs.SD
diffusion
扩散模型相关
Abstract
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.
Chinese Translation
用于文本转语音(TTS)的条件流匹配(CFM)模型在推理过程中存在频率演化不连贯的问题。尽管其他领域的扩散模型已经解决了类似的频谱不平衡问题,但这些通用解决方案无法推广到 CFM 固有的不协调声学动态。我们证明,通过引入一种新颖的免训练频率选择性增强策略,可以有效缓解这一问题。使用离散小波变换(DWT),我们的方法在 ODE 积分过程中动态调制梅尔频谱图的子带,通过惩罚激进的低频增长并增强滞后的高频细节来同步频谱发展。在多种架构(Matcha-TTS、F5-TTS、IndicF5)上验证后,我们的方法将所需的函数评估次数(NFE)从 32 降低到 26,并将 Frechet 音频距离(FAD)最多改善 61%,同时不损害平均意见得分、说话人相似度和语音可懂度。
cs.LG / 154 / 2609.34433
Admissible Diffusion for Multimodal Interventional Trajectories
面向多模态干预轨迹的可容许扩散
Xing Han, Shravan Chaudhari, Jiarui Shao, Paul Pu Liang, Suchi Saria
cs.LG
diffusion
扩散模型相关
Abstract
Generating a plausible clinical trajectory does not establish what would happen under a different treatment. We present ADMIT, a framework combining irregular multimodal representations, treatment-conditioned latent diffusion and explicit constraints on generated states or actions. We formulate its interventional target through sequential g-computation and distinguish causal assumptions from constraint satisfaction. Its admissibility mechanism translates physiological prior knowledge into explicit constraints on generated states and proposed actions. Treatment-exposure dynamics condition latent transitions, while state projection or action gating applies the constraints during rollout so that they influence subsequent trajectory generation. In our preliminary experiments, multimodal inputs improved supervised hidden-state recovery and reduced treatment-contrast error. In a simulated dosing-schedule experiment with leak-free history encoding, ADMIT predicted most of the tumor-volume change caused by redistributing a fixed total dose. An exposure input improved these predictions around a temporary dose reduction whether or not the assumed clearance rate was correct, but reduced the predicted size of a dose effect, and a deterministic recurrent baseline matched ADMIT's average predictions. Exposure projection reduced constraint violations, although enforcement remained incomplete. Semi-synthetic experiments using eICU context illustrated treatment-response generation under fixed and adaptive policies. Observational examples further characterize model treatment sensitivity. ADMIT provides a framework for testing whether complementary observations and physiological restrictions improve intervention trajectories, with representation recovery, effect accuracy and rule enforcement assessed separately.
Chinese Translation
生成一条看似合理的临床轨迹,并不能确定在不同治疗下会发生什么。我们提出了 ADMIT,一个结合不规则多模态表示、以治疗为条件的潜在扩散,以及对生成状态或动作的显式约束的框架。我们通过序贯 g-computation 来形式化其干预目标,并将因果假设与约束满足区分开来。其可容许性机制将生理先验知识转化为对生成状态和所提出动作的显式约束。治疗-暴露动态对潜在转移进行条件化,而状态投影或动作门控在 rollout 过程中施加约束,使其影响后续轨迹生成。在我们的初步实验中,多模态输入改善了监督式隐藏状态恢复,并降低了治疗对比误差。在一个具有无泄漏历史编码的模拟给药方案实验中,ADMIT 预测了由重新分配固定总剂量所引起的大部分肿瘤体积变化。暴露输入在临时剂量降低前后改善了这些预测,无论假设的清除率是否正确,但降低了剂量效应的预测幅度,并且一个确定性循环基线在平均预测上与 ADMIT 相匹配。暴露投影减少了约束违反,尽管执行仍不完整。使用 eICU 背景的半合成实验展示了固定策略和自适应策略下的治疗反应生成。观察性示例进一步刻画了模型的治疗敏感性。ADMIT 提供了一个框架,用于检验互补观察和生理限制是否改善干预轨迹,并将表示恢复、效应准确性和规则执行分别评估。
cs.LG / 155 / 2609.34467
Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning
面向高效视觉-语言-动作策略学习的对齐引导流 Transformer
Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao
cs.LG
diffusion
扩散模型相关
Abstract
Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.
Chinese Translation
近期视觉-语言-动作(VLA)模型的进展通过统一感知、指令与控制,指向通用机器人智能。尽管取得了令人瞩目的进展,现有 VLA 模型往往因视觉、语言和动作之间的\emph{三模态失配}而适应性较差,这削弱了动作接地并损害了泛化性和微调效率。在这项工作中,我们提出了对齐引导流 Transformer(AGFT),这是一个新颖的框架,通过专门的对齐损失显式地强制三模态对齐,弥合跨模态的表征差距并增强任务适应性。尽管先前研究主要强调双模态视觉--语言对齐,我们系统地形式化并研究了 VLA 模型中的三模态对齐,并提供了消融与分析,以分离其在改进适应性和鲁棒性中的作用。为了进一步加速部署,我们采用了流匹配目标,在保持精度的同时,实现比基于扩散的策略少得多的推理步数。理论上,我们建立了三模态对齐差距与流匹配优化紧致性之间的定量联系;实证上,在广泛基准上的实验表明,与 SOTA 基线相比,AGFT 实现了更高的成功率和更低的推理延迟,凸显三模态对齐是扩展稳健 VLA 操作的关键要素。
cs.LG / 156 / 2609.34491
M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization
M3OS:一个由蒙特卡洛图搜索编排的多智能体 LLM 系统,用于证据可追溯的分子优化
Junjie Wang, Yaowei Jin, Ruohui Tang, Guonan Cui, Haojie Wang, Penglei Wang, Dingyan Wang, Duo An, Shuangjia Zheng, Qian Shi
cs.LG · q-bio.BM
large language model
大语言模型相关
Abstract
Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they must recover candidate identities, prior evaluations, and task constraints to guide subsequent decisions. We present M3OS, a multi-agent LLM system that decouples molecular-design reasoning from optimization-state management through Monte Carlo graph search. A persistent graph links evaluated candidates, parent-child transformations and evaluation evidence, while rewards and visit statistics guide LLM-assisted parent selection. Two branches combine tool-driven candidate generation with knowledge- and case-guided medicinal-chemistry editing. An execution harness controls graph updates through structured output extraction, molecular validation and task-bound evaluation. Agents receive role-specific contexts, while the graph preserves optimization trajectories beyond their active contexts. Across three molecular optimization benchmarks, M3OS achieves higher success rates than baselines, supporting the integration of persistent search state, specialized agents and controlled execution for multi-constraint optimization.
Chinese Translation
小分子优化通过迭代式的多目标决策,将药物化学推理与计算证据整合在一起。当大语言模型(LLM)对主要存储在对话上下文中的优化历史进行推理时,它们必须恢复候选分子的身份、先前的评估结果以及任务约束,以指导后续决策。我们提出了 M3OS,这是一个通过蒙特卡洛图搜索将分子设计推理与优化状态管理解耦的多智能体 LLM 系统。一个持久化的图将已评估的候选分子、父子变换关系与评估证据连接起来,而奖励和访问统计则引导 LLM 辅助的父节点选择。两条分支将工具驱动的候选分子生成与知识和案例引导的药物化学编辑结合起来。一个执行框架通过结构化输出抽取、分子验证和任务绑定的评估来控制图的更新。各智能体接收角色特定的上下文,而图则在其活动上下文之外保存优化轨迹。在三个分子优化基准上,M3OS 取得了高于基线的成功率,从而支持将持久化搜索状态、专用智能体和受控执行相集成,以用于多约束优化。
cs.LG / 157 / 2609.34507
KiT: A Foundation Model for Financial Time-Series Forecasting using DiffusionTransformers
KiT:一个使用扩散Transformer进行金融时间序列预测的基础模型
Boyu Zhang, Haorui Li
cs.LG · cs.CV
diffusion
扩散模型相关
Abstract
Financial candlestick forecasting is fundamental to quantitative investment, yet it remains exceptionally challenging due to extremely low signal-to-noise ratios and vast heterogeneity across markets and instruments. Existing approaches have largely attempted to introduce deep learning to capture hidden temporal features, but most adopt an auto-regressive formulation, which leads to error accumulation during inference. Meanwhile, general-purpose time-series foundation models are not tailored to the unique structure of k-line data and yield unsatisfactory performance on downstream candlestick forecasting tasks. To tackle these problems, we introduce KiT, a K-line Diffusion Transformer foundation model, and reformulate future prediction as conditional path generation via flow matching: given a historical context window, the model generates an ensemble of plausible future OHLCV trajectories. We pre-train KiT at multiple parameter scales on billions of candlestick bars spanning multiple markets and timescales. Across three markets and seven resolutions, KiT attains a mean return RankIC of 0.057 and a mean volatility RankIC of 0.66, leading at every timescale and outperforming both task-specific financial forecasters and general time-series foundation models. Code will be available at: https://github.com/Luciferbobo/KiT.
Chinese Translation
金融K线预测是量化投资的基础,然而由于极低的信噪比以及跨市场和跨工具的巨大多样性,它仍然极具挑战性。现有方法大多试图引入深度学习来捕捉隐藏的时间特征,但大多数采用自回归形式,这会在推理过程中导致误差累积。与此同时,通用时间序列基础模型并未针对K线数据的独特结构进行定制,因而在下游K线预测任务上表现不尽如人意。为了解决这些问题,我们提出了KiT,一种K线扩散Transformer基础模型,并将未来预测重新表述为通过流匹配进行条件路径生成:给定历史上下文窗口,模型生成一组合理的未来OHLCV轨迹。我们在跨越多个市场和多个时间尺度的数十亿根K线柱上,以多种参数规模预训练KiT。在三个市场和七个分辨率上,KiT取得了0.057的平均收益RankIC和0.66的平均波动率RankIC,在每个时间尺度上均领先,并优于特定任务的金融预测器和通用时间序列基础模型。代码将发布于:https://github.com/Luciferbobo/KiT。
cs.LG / 158 / 2609.34509
Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs
低置信度重掩码困住了灵活性:实现扩散 LLM 中多样化 rollout 的任意阶潜力
Moongyu Jeon, Dongjae Jeon, Bumjun Kim, Mingyu Kim, Albert No
cs.LG · cs.CL
diffusion
扩散模型相关
Abstract
Masked diffusion language models support arbitrary-order generation, suggesting a natural way to produce diverse outputs. However, recent work argues that this flexibility reduces diversity by delaying high-uncertainty tokens that can lead to different generation paths. We trace this diversity loss not to arbitrary-order generation itself, but largely to low-confidence remasking (LCR), a widely used decoding rule. At each step, LCR samples a token at every masked position but commits only the sampled token with the highest probability, filtering out the rest. We show that this mechanism can exponentially suppress lower-probability tokens as more positions compete, and observe the same suppression in LLaDA. In contrast, top-probability position selection (TPP), which has often been conflated with LCR under the shared label confidence-based decoding, avoids this diversity loss. TPP first selects the position whose most likely token has the highest probability, then samples directly from that position's distribution. Replacing LCR with TPP restores diversity and yields Pass@$k$ comparable to left-to-right decoding, suggesting that the reported diversity loss stems largely from LCR's filtering rather than from generating high-confidence positions first. To further exploit order flexibility, we introduce Entropy-Guided Initialization (EGI), which samples the first token at the highest-entropy position and then follows TPP. This simple modification further improves rollout diversity and solution coverage beyond left-to-right decoding, with gains extending to downstream policy optimization, highlighting the potential of arbitrary-order generation for diverse rollouts.
Chinese Translation
掩码扩散语言模型支持任意阶生成,这暗示了一种产生多样化输出的自然方式。然而,近期工作认为,这种灵活性会通过延迟那些可能导向不同生成路径的高不确定性 token 来降低多样性。我们将这种多样性损失归因于——并非任意阶生成本身,而主要是低置信度重掩码(LCR),一种被广泛使用的解码规则。在每一步,LCR 会在每个被掩码的位置上采样一个 token,但只确认其中采样概率最高的那个 token,其余则被过滤掉。我们表明,随着更多位置参与竞争,该机制会以指数方式抑制概率较低的 token,并在 LLaDA 中观察到同样的抑制现象。相比之下,最高概率位置选择(TPP)——它常因与 LCR 共享“基于置信度的解码”这一标签而被混为一谈——则避免了这种多样性损失。TPP 首先选出其最可能 token 具有最高概率的位置,然后直接从该位置的分布中采样。用 TPP 替换 LCR 可以恢复多样性,并取得与从左到右解码相当的 Pass@$k$,这表明所报告的多样性损失主要源于 LCR 的过滤,而非源于先生成高置信度的位置。为了进一步利用顺序灵活性,我们引入了熵引导初始化(EGI),它在熵最高的位置采样第一个 token,然后遵循 TPP。这一简单修改在从左到右解码之外进一步提升了 rollout 多样性与解的覆盖范围,其收益还延伸到下游策略优化,凸显了任意阶生成在多样化 rollout 方面的潜力。
cs.LG / 159 / 2609.34585
Compute Time Scaling with Recursive Models for Combinatorial Optimization
使用递归模型进行组合优化的计算时间缩放
Zhengxi Zhang, Paul Swoboda
cs.LG
diffusion
扩散模型相关
Abstract
We propose Tiny Recursive Models for Combinatorial Optimization (\ours{}), a general neural method for combinatorial optimization that scales both depth (how often we recursively invoke our network) and width (how much we sample in parallel). Both are fundamental for combinatorial optimization: hard instances demand a large amount of compute, while a small network is essential to avoid overfitting and capture the algorithmic essence of optimization. In particular, our method consists of a graph-aware tiny recursive model that iterates on a latent state with adaptive halting and needs only a lightweight problem-specific decoder. Compared with previous heatmap-based general neural solvers, it achieves a better balance between solution quality and inference speed on both the Traveling Salesman Problem~(TSP) and the Maximum Independent Set~(MIS) problem, and remains competitive with hybrid methods that combine neural components with heuristics specific to each problem. With the same backbone architecture for both tasks, \ours{} outperforms every diffusion-based solver on TSP from 500 to 10,000 cities at a lower inference cost, and on the standard Erdős--Rényi-[700-800] MIS benchmark it surpasses all neural solvers except those that only work well on MIS. We then explore self-relabeling for self-supervised training. We periodically replace the current set of training labels with the model's own better solutions, as an alternative training signal. Self-relabeling can, while forgoing supervision from near-optimal solutions, still result in on-par quality.
Chinese Translation
我们提出用于组合优化的微型递归模型(\ours{}),这是一种通用的组合优化神经方法,它同时扩展深度(我们递归调用网络的频率)和宽度(我们并行采样的数量)。两者对组合优化都是根本性的:困难实例需要大量计算,而小型网络对于避免过拟合和捕捉优化的算法本质至关重要。特别地,我们的方法由一个图感知的微型递归模型组成,该模型在潜在状态上迭代并具有自适应停止,且仅需要一个轻量级的、问题特定的解码器。与之前基于热图的通用神经求解器相比,它在旅行商问题(TSP)和最大独立集(MIS)问题上都在解质量和推理速度之间取得了更好的平衡,并且与将神经组件与每个问题特有的启发式方法相结合的混合方法相比仍具竞争力。在两个任务使用相同主干架构的情况下,\ours{} 在 500 到 10,000 个城市的 TSP 上以更低的推理成本优于每一个基于扩散的求解器,并且在标准 Erdős--Rényi-[700-800] MIS 基准上,它超越了所有神经求解器,除了那些仅在 MIS 上表现良好的求解器。然后,我们探索用于自监督训练的自重新标记。我们定期用模型自身更好的解替换当前的训练标签集,作为一种替代训练信号。自重新标记虽然放弃了来自近最优解的监督,但仍能产生同等质量。
cs.LG / 160 / 2609.34642
Tilted Schrödinger Bridge Matching
倾斜薛定谔桥匹配
Sergei Kholkin, Evgeny Burnaev, Alexander Korotin
cs.LG
diffusion
扩散模型相关
Abstract
Schrödinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion models but underexplored for Schrödinger bridges. We introduce Tilted Schrödinger Bridge Matching (TSBM), a post-training method for fine-tuning a learned bridge $P$ between source $p_0$ and target $p_1$ toward a reward-tilted target $p_1^r\propto p_1e^r$, while preserving source $p_0$. We formulate this adaptation as alternating optimization initialized from $P$, provide theoretical justification, and derive a practical algorithm based on Adjoint Matching. We evaluate TSBM on unpaired image-to-image translation targeting digit properties in MNIST and facial attributes in CelebA.
Chinese Translation
薛定谔桥提供了一个熵正则化的框架,并为非配对域翻译给出了一种有原则的解决方案。在实践中,预训练好的桥可能需要通过奖励来适应人类偏好或物理约束,这一问题与扩散模型中的奖励倾斜密切相关,但在薛定谔桥中尚未得到充分探索。我们提出了倾斜薛定谔桥匹配(Tilted Schrödinger Bridge Matching,TSBM),这是一种后训练方法,用于将源 $p_0$ 与目标 $p_1$ 之间学习到的桥 $P$ 微调至奖励倾斜的目标 $p_1^r\propto p_1e^r$,同时保持源 $p_0$ 不变。我们将这种适配形式化为从 $P$ 初始化的交替优化,给出了理论论证,并推导出一种基于伴随匹配(Adjoint Matching)的实用算法。我们在非配对图像到图像翻译任务上评估了 TSBM,其目标分别为 MNIST 中的数字属性以及 CelebA 中的人脸属性。
cs.LG / 161 / 2609.34677
Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
学习回忆什么:面向世界模型的自适应多线索情景记忆
Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar, Jong Chul Ye, Yuki Mitsufuji
cs.LG · cs.AI · cs.CV
diffusion
扩散模型相关
Abstract
World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current prediction, and which available retrieval cues should be trusted to find them? This is challenging because fixed criteria based on recency, pose overlap, or visual similarity can be unreliable across environments and queries. We propose Future-Aware Recall (FAR), a framework that learns episodic recall from future-aware predictive supervision and adaptive multi-cue scoring. During training, FAR measures predictive utility by the conditional log-likelihood of the realized future given recalled context, approximated by negative diffusion prediction loss, and uses it to train a retriever that remains future-blind at inference. The retriever learns cue-specific relevance and automatically determines which available retrieval cues, such as time, pose, vision, and audio, to trust for each query when selecting memories. Across three complementary settings, FAR outperforms hand-designed recall even with the same retrieval cues, automatically adapts which available cues to trust, and recalls the right history as the world changes. Together, these results establish FAR as a flexible, principled approach to episodic memory access in world models.
Chinese Translation
世界模型从当前经验和动作预测未来的观测,然而预测可能依赖于很久以前看到的观测。情景记忆保存过去的观测以供后续回忆;然而,随着记忆不断累积,它引出了一个根本性的问题:哪些记忆对当前的预测有用,以及应当信任哪些可用的检索线索来找到它们?这具有挑战性,因为基于新近程度、位姿重叠或视觉相似性的固定准则在不同环境和查询下可能并不可靠。我们提出未来感知回忆(Future-Aware Recall, FAR),一个从未来感知的预测监督和自适应多线索评分中学习情景回忆的框架。在训练期间,FAR 通过给定被回忆上下文下真实发生的未来的条件对数似然来衡量预测效用,并用负扩散预测损失来近似该似然,进而用它训练一个在推理时对未来保持盲视的检索器。该检索器学习线索特定的相关性,并在选择记忆时自动确定为每个查询信任哪些可用的检索线索,例如时间、位姿、视觉和音频。在三种互补的设置下,即便使用相同的检索线索,FAR 也优于人工设计的回忆方法,能够自动调整信任哪些可用线索,并随着世界的变化回忆出正确的历史。这些结果共同表明,FAR 是一种灵活且有原则的、用于世界模型中情景记忆访问的方法。
cs.LG / 162 / 2609.34679
From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling
从偏好到互惠:具有实证基础的基于 LLM 智能体建模的去中心化匹配
Wangxuan Fan, Xiaoyu Nie, Zhoutian Shi, Xiangcheng Meng, Shipei Zeng, Pin Gao, Yan Hu, Zhongxiang Dai
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under limited information. We propose a dynamic bipartite matching framework that combines large language model (LLM) agents with contextual bandits. In a simulated Chinese marriage market, economically grounded LLM agents evaluate locally encountered candidates, while agent-specific Logistic-UCB models learn reciprocal acceptance from realized proposal outcomes. The mechanism therefore separates two decisions---\emph{whom do I like?} and \emph{who is likely to like me back?}---without requiring ex ante market-wide preference rankings. We first validate LLM-induced mate preferences against the empirical conditional-logit reference across multiple LLM backbones. In the $50\times50$ matching experiment, Bandit-UCB achieves the highest mean mutual welfare (56.01 versus 54.87 for Gale--Shapley), a smaller gender rank gap than the classical baselines, and the fewest blocking pairs among the LLM-ABM policies. Learned acceptance models show economically interpretable gender-differentiated associations, while counterfactual setups reveal no systematic unilateral advantage from prior search knowledge. Overall, these results support the advantages of decentralized matching with LLM-based behavioral modeling and online learning under incomplete information for economic simulation and computational social science research.
Chinese Translation
二分匹配是博弈论和市场设计中的一个基本问题。诸如 Gale--Shapley 之类的经典方法假设完全偏好和集中式计算,而许多现实世界的匹配过程是去中心化的、异步的,并且受有限信息下的序贯交互塑造。我们提出一个动态二分匹配框架,将大语言模型(LLM)智能体与上下文老虎机相结合。在一个模拟的中国婚姻市场中,具有经济学基础的 LLM 智能体评估局部遇到的候选人,而针对特定智能体的 Logistic-UCB 模型从已实现的提议结果中学习互惠接受。因此,该机制将两个决策分开——“我喜欢谁?”和“谁可能也喜欢我?”——而不要求事前全市场的偏好排序。我们首先在多个 LLM 主干模型上,将 LLM 诱导的择偶偏好与实证条件 Logit 参照进行验证。在 $50\times50$ 匹配实验中,Bandit-UCB 实现了最高的平均互惠福利(56.01,而 Gale--Shapley 为 54.87)、比经典基线更小的性别排名差距,以及在 LLM-ABM 策略中最少的阻塞对。学习得到的接受模型显示出具有经济可解释性的性别差异化关联,而反事实设置则表明,来自先验搜索知识不存在系统性的单边优势。总体而言,这些结果支持了在不完全信息下,采用基于 LLM 的行为建模与在线学习的去中心化匹配在经济模拟和计算社会科学研究中的优势。
cs.LG / 163 / 2609.34680
QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
QuantForge:为 MXFP4 训练后量化发现残差分解
Qiulin Shang, Zhoutong Wu, Jie Hu, Kun Yuan
cs.LG
large language model
大语言模型相关
Abstract
Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algorithmic decomposition is therefore not fully known before search. LLM-driven program evolution offers a way to explore these choices, but performance scores alone do not explain which design should change next. We introduce QuantForge, a PTQ discovery system that records competing explanations, selects controls that distinguish them, and checks that successor code implements the resulting conclusions. This residual compilation guides program revisions while retaining useful programs even when their original explanations are rejected. Remeasuring the revised program reveals the next error to address. This process discovers HiRes, a fixed MXFP4 quantizer that shapes coordinates, refines legal code assignments, and recovers errors along attention and MLP paths. Each stage acts on residuals measured after the preceding stage has executed. Across seven tasks, HiRes achieves the lowest seven-model Robust Fit (0.09300) and the lowest quantized Fit-7 at 32B. In matched-budget comparisons of LLM-driven program evolution, each with 240 evaluator calls, QuantForge reaches a held-out transfer target in six of eight runs, compared with three each for textual memory and reflection memory, and one for score-only evolution, despite evaluating fewer new programs. These results show that QuantForge improves the discovery of transferable PTQ algorithms by turning controlled evidence into subsequent program changes.
Chinese Translation
四比特训练后量化可以降低大语言模型的内存需求,但在严格的 MXFP4 W4A4 下保持精度需要协调若干设计选择。坐标变换会改变块编码误差,而块编码误差又会影响通过网络传播的残差。因此,有用的算法分解在搜索之前并非完全已知。LLM 驱动的程序演化提供了一种探索这些选择的方式,但仅凭性能分数并不能解释下一步应该改变哪个设计。我们介绍 QuantForge,一个 PTQ 发现系统,它记录相互竞争的解释,选择能够区分它们的控制变量,并检查后继代码是否实现了由此得出的结论。这种残差编译指导程序修订,同时即使有用程序的原始解释被否定,也保留这些有用程序。重新测量修订后的程序会揭示下一个需要处理的误差。这一过程发现了 HiRes,一种固定的 MXFP4 量化器,它塑造坐标、细化合法码字分配,并沿注意力与 MLP 路径恢复误差。每个阶段都作用于在前一阶段执行后测得的残差。在七个任务上,HiRes 取得了最低的七模型 Robust Fit(0.09300),以及在 32B 上最低的量化 Fit-7。在 LLM 驱动程序演化的等预算比较中,每次比较有 240 次评估器调用,QuantForge 在八次运行中的六次达到留出迁移目标,而文本记忆和反思记忆各为三次,仅分数演化为一次,尽管它评估的新程序更少。这些结果表明,QuantForge 通过将受控证据转化为后续程序变更,改进了可迁移 PTQ 算法的发现。
cs.LG / 164 / 2609.34681
SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
SOLAR:一种用于LLM预训练的状态驱动在线学习率调度器
Qiulin Shang, Binyu Wang, Yongqi Qiao, Songde Rao, Zhoutong Wu, Kun Yuan
cs.LG
large language model
大语言模型相关
Abstract
Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.
Chinese Translation
学习率(LR)调度在大型语言模型(LLM)预训练中起着核心作用,然而当前的实践仍然在很大程度上依赖于人工设计的启发式方法,例如 Warmup-Cosine-Decay 和 Warmup-Stable-Decay。由于这些调度方案是预先固定的,它们无法适应不断演变的优化动态。在 Learning to Optimize(L2O)框架内进行在线学习调度提供了一种动态的替代方案,但由于信号噪声、反馈延迟以及灾难性发散的风险,它在 LLM 规模上仍然十分脆弱。我们提出 SOLAR(State-driven Online Learning rAte scheduleR),一个用于可靠在线学习率自适应的稳定化框架。SOLAR 使用一个基础调度作为参考,并为各个参数组学习有界的、依赖于状态的残差修正。每个修正量在每一步都重新锚定到基础调度上,从而使策略能够自适应地调整学习率,而无需重新学习预热-衰减曲线。一种轻量级的状态表示与进度感知的奖励引导在线学习,同时一个断路器(Circuit-Breaker)可在罕见的不安全动作之后恢复训练。在自回归语言模型预训练中,对于从 60M 到 1B 的稠密模型、AdamW 与 Muon,以及两种最高至 3B 的 MoE 设置,SOLAR 相较于调优后的静态调度和自动学习率调优器均改善了最终困惑度。匹配的 130M 对照实验表明,加入基础锚定与动作边界将全局 PPO 控制器的最终 PPL 从 27.09 改善到 23.74,而在相同的两个随机种子上,分组控制达到 22.87。在 60M 代理模型上训练的残差策略也可以被冻结,并在更大的稠密模型规模上复用,而无需针对目标进行 PPO 更新,且在四倍的基础学习率范围内仍然有效。这些结果确立了 SOLAR 作为 LLM 预训练中一种实用的可学习学习率控制器。
cs.LG / 165 / 2609.34831
Structured Neural SDEs for Functional Calibration
用于泛函校准的结构化神经随机微分方程
Francesco Piatti, Andrea Iannucci, Thomas Cass
cs.LG · math.PR
diffusion
扩散模型相关
Abstract
Neural Stochastic Differential Equations (Neural SDEs) provide flexible continuous-time generative models, but generic neural drift and diffusion networks are costly to simulate on long horizons and can give unstable gradients when the training signal is a path functional rather than a pointwise observation. We introduce SLiSDE, a family of Neural SDE models built from structured linear stochastic layers. Parallel-in-time simulation is obtained at the layer level, while expressivity is recovered by gated in-flow stacking: previous-layer paths modulate the next layer's latent flow through learned gates. For functional calibration tasks in which rare paths dominate the loss, we add an optional Girsanov tilt that acts as a learned importance sampler with an exact likelihood-ratio correction. We prove well-posedness, a discretisation error bound, validity of the change of measure, and a universality result: the terminal laws of the gated stack are dense in the space of square-integrable laws. Experiments on functional calibration benchmarks show that the structured model outperforms fully neural SDE baselines while retaining parallel-time simulation and stable importance weights.
Chinese Translation
神经随机微分方程(Neural SDEs)提供了灵活的连续时间生成模型,但通用的神经漂移和扩散网络在长时间范围上模拟代价高昂,并且当训练信号是路径泛函而非逐点观测时,可能给出不稳定的梯度。我们引入 SLiSDE,这是一族由结构化线性随机层构建的神经随机微分方程模型。在层级别上获得了时间并行模拟,而表达能力通过门控式流入堆叠得以恢复:前一层的路径通过学习到的门控调节下一层的潜在流。对于稀有路径主导损失的泛函校准任务,我们添加一个可选的 Girsanov 倾斜,它充当一个具有精确似然比校正的学习型重要性采样器。我们证明了适定性、离散化误差界、测度变换的有效性,以及一个普适性结果:门控堆叠的终端分布律在平方可积分布律空间中稠密。在泛函校准基准上的实验表明,结构化模型优于完全神经随机微分方程基线,同时保持时间并行模拟和稳定的重要性权重。
cs.LG / 166 / 2609.34965
Cyclostationary Phase Conditioning for Medical Time Series Diffusion
用于医学时间序列扩散的循环平稳相位条件化
Samuel Ruiperez-Campillo, Michele Copetti, Jorge da Silva Goncalves, Sonia Laguna, Julia E. Vogt
cs.LG · cs.AI · eess.SP
diffusion
扩散模型相关
Abstract
Many physiological time series, such as cardiac and brain recordings, exhibit cyclostationarity: their statistics vary periodically with an underlying cycle phase. Corruption from motion, poor contact, and physiological interference obscures morphology needed for diagnosis, making signal restoration essential. Existing diffusion approaches condition on corrupted observations alone and must learn cyclic structure implicitly. We instead propose two inductive biases which encode cyclostationarity: a shift-covariant wavelet representation and dense per-sample phase conditioning inferred from the corrupted input. We further introduce a training-free cyclostationarity index that quantifies phase structure and predicts when phase conditioning will help. Finally, we propose antithetic coupling of reverse trajectories to reduce sampling variance while achieving comparable performance with fivefold fewer network evaluations. Across modalities, our results show that explicitly encoding measurable cyclic structure improves physiological time-series restoration.
Chinese Translation
许多生理时间序列,例如心脏和大脑记录,表现出循环平稳性:其统计特性随一个潜在的周期相位而周期性变化。由运动、接触不良和生理干扰造成的损坏会掩盖诊断所需的形态,使得信号恢复至关重要。现有的扩散方法仅以受损坏的观测为条件,并且必须隐式地学习循环结构。我们转而提出两种编码循环平稳性的归纳偏置:一种平移协变小波表示,以及从受损坏输入中推断出的稠密逐样本相位条件化。我们进一步引入一个无需训练的循环平稳性指数,它量化相位结构并预测相位条件化何时会有帮助。最后,我们提出反向轨迹的对偶耦合,以减少采样方差,同时在仅需五分之一网络评估次数的情况下达到相当的性能。跨模态地,我们的结果表明,显式编码可测量的循环结构能够改善生理时间序列恢复。
cs.LG / 167 / 2609.34990
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
MaPP:一种用于数据高效 RLVR 的统一边缘化后验预测框架
Yangyang Ren, Haodong Zhu, Sheng Xu, Yanjing Li, Nikolai Yu. Zolotykh, Wentao Zhang, Baochang Zhang
cs.LG
large language model
大语言模型相关
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.
Chinese Translation
带可验证奖励的强化学习(RLVR)提升了大语言模型的推理能力,但会因 rollout 和策略更新而产生巨大成本。在线提示选择通过使用每个提示的贝叶斯后验来预测难度并优先选择信息量大的提示,从而提升效率。然而,现有方法忽视了从采样响应中提取学习信号的可靠性。在 GRPO 中,一个响应的优势既取决于其自身结果,也取决于通过组归一化得到的其同伴随机采样的结果。我们的理论和实验分析表明,组构成中的不确定性会引入构成噪声,这是一种不会消失的方差成分,会对梯度估计误差施加不可约的下界,并损害下游提示选择。我们提出 MaPP(Marginalized Posterior-Predictive),一种用于数据高效 RLVR 的统一框架,它使用共享的 Beta 后验对响应级优势估计进行去噪,并改进提示选择。对于每个响应,MaPP 通过闭式 Beta-Binomial 边缘化,将标准组相对优势替换为构成不变的内在优势。由此得到的后验预测估计器的误差可证明会随着后验集中而减小。使用同一后验,MaPP 推导出一个不确定性感知的提示选择得分,以提升数据效率且无需额外 rollout 成本。在五个模型骨干上进行的数学、规划和视觉几何实验表明,MaPP 始终优于 GRPO 和强选择基线,在相同 rollout 预算和设置下,相较于最强基线实现了最高 +2.45 的平均准确率提升,并创造了新的最优水平。
cs.LG / 168 / 2609.34992
Composable Decoding on the Probability Simplex: Theory and Implementation
概率单纯形上的可组合解码:理论与实现
Xiaotong Ji, Ahmed Khaled Khamis, Rasul Tutunov, Matthieu Zimmer, Haitham Bou-Ammar
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Decoding for large language models is typically treated as a collection of isolated sampling strategies, with limited theoretical understanding of the behaviours they induce and how their underlying objectives relate. We formulate decoding as an optimisation problem over next-token distributions on the probability simplex, balancing expected model score against regularisation under support constraints. This view recovers familiar decoding methods through choices of regularisers and support constraints; more importantly, it enables new decoders to be constructed by composing distributional preferences within a single optimisation problem without external rewards, learned critics, or model parameter updates. We introduce CompoSimplex, a library with configurable support rules, regularisation primitives, and simplex solvers for constructing and evaluating compositional decoders. We evaluate standard samplers, individual regularisers, and compositions across multiple models and reasoning tasks. Our results show that compositions can realise trade-offs between single-sample quality, multi-sample quality, and diversity that are not attained by individual decoding objectives.
Chinese Translation
大语言模型的解码通常被视为一组孤立的采样策略,而人们对它们所引发的行为以及其底层目标如何相互关联的理论理解有限。我们将解码形式化为概率单纯形上关于下一词元分布的优化问题,在支撑约束下平衡期望模型得分与正则化。这一视角通过选择正则化器和支撑约束,恢复了熟悉的解码方法;更重要的是,它使得可以通过在单个优化问题中组合分布偏好来构建新的解码器,而无需外部奖励、学习到的评论器或模型参数更新。我们介绍了 CompoSimplex,一个具有可配置支撑规则、正则化原语和单纯形求解器的库,用于构建和评估可组合解码器。我们在多个模型和推理任务上评估了标准采样器、单个正则化器以及组合。我们的结果表明,组合能够实现单样本质量、多样本质量和多样性之间的权衡,而这些权衡无法由单个解码目标达到。
cs.LG / 169 / 2609.35028
VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation
VEX-Bench:对LLM生成的错误信息的验证复杂性进行基准测试
Hanxun Huang, Yutao Wu, Qizhou Wang, Silvia Montaña-Niño, Yige Li, Xiang Zheng, Elif Buse Doyuran, Phoebe Matich, Xiao Liu, Xingjun Ma, Sarah Erfani, Christopher Leckie
cs.LG · cs.AI · cs.CL · cs.CY · cs.IR
large language model
大语言模型相关
Abstract
Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content to verify. We introduce VEX-Bench, a unified benchmark for evaluating the verification complexity of LLM-generated misinformation, as perceived during screening, across models and generation methods. Verification complexity is assessed along multiple dimensions derived from journalistic and fact-checking practices, capturing checkability, harm potential, source credibility signals, imposter legitimacy, and expected verification effort. We define the VEX score as an integrated measure combining elicitation yield and verification complexity to quantify how generated content consumes limited verification capacity. We construct a benchmark spanning two misinformation categories, 6 high-stakes domains, and 60 real-world topics, and evaluate 7 frontier LLMs and 7 generation methods, yielding 5{,}880 articles. We employ an LLM-as-judge for scalable evaluation and validate it using content-analysis methodology, including ordinal Krippendorff $α$ for inter-annotator reliability, complemented by fact-checking agents for verification. Our findings show that no single method dominates all dimensions, underscoring the need for multi-dimensional evaluation. LLMs can generate high-VEX misinformation at 3$\times$ to 169$\times$ lower cost than agent-based verification. Such content is often prioritized during screening, consuming scarce verification resources and introducing a systematic risk of misallocation in resource-constrained verification systems. The code is publicly available in our \href{https://github.com/HanxunH/VEX-Bench}{GitHub repository}.
Chinese Translation
大型语言模型(LLM)使错误信息的生产成本变得低廉,但验证成本并未降低,从而在信息生态系统中造成了日益加剧的不对称性。在时间、人力和预算紧张的限制下,媒体组织、平台和事实核查人员依赖筛选来确定哪些内容应优先验证。我们提出 VEX-Bench,这是一个统一的基准,用于评估在筛选过程中所感知到的、跨模型和生成方法的 LLM 生成错误信息的验证复杂性。验证复杂性沿多个维度进行评估,这些维度源自新闻与事实核查实践,涵盖可核查性、危害潜力、来源可信度信号、冒充者合法性以及预期验证工作量。我们将 VEX 分数定义为一种综合度量,结合诱导产出与验证复杂性,以量化生成内容如何消耗有限的验证能力。我们构建了一个基准,涵盖两个错误信息类别、6 个高风险领域和 60 个现实世界主题,并评估了 7 个前沿 LLM 和 7 种生成方法,产生了 $5{,}880$ 篇文章。我们采用 LLM 作为评判者以进行可扩展评估,并使用内容分析方法论对其进行验证,包括用于标注者间信度的序数 Krippendorff $α$,并由事实核查智能体进行验证作为补充。我们的研究结果表明,没有任何单一方法能在所有维度上占优,这凸显了多维度评估的必要性。LLM 能够以比基于智能体的验证低 3$\times$ 至 169$\times$ 的成本生成高 VEX 错误信息。此类内容在筛选过程中往往被优先处理,消耗稀缺的验证资源,并在资源受限的验证系统中引入系统性错配风险。代码已在我们 \href{https://github.com/HanxunH/VEX-Bench}{GitHub 仓库} 中公开可用。
cs.LG / 170 / 2609.35060
Beyond Gradient Flow: Identifiability and Recovery from Distribution Snapshots
超越梯度流:从分布快照中的可辨识性与恢复
Nam D. Nguyen, Valeriya Malysheva
cs.LG
diffusion
扩散模型相关
Abstract
Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift $F$ only through its score-weighted divergence $\nabla\cdot F+F\cdot\nabla\logρ$, leaving a $ρ$-solenoidal gauge invisible to any single-time constraint. Time-indexed transport formulations cannot resolve this ambiguity: every admissible marginal path admits a curl-free explanation, minimum-action reconstruction selects it, and marginal fit alone cannot distinguish dynamically inequivalent explanations. Requiring one autonomous field to explain several marginals instead makes part of the hidden circulation visible as $\nabla\logρ$ changes across marginals. Separating instantaneous Fokker-Planck source constraints from the snapshot experiment, we show that the source constraints identify the field modulo the kernel of a stacked score-weighted divergence operator. For generic Gaussian shape variation, source constraints at $K\ge m$ time points in intrinsic dimension $m$ eliminate every polynomial gauge direction, whereas finitely many density snapshots alone admit aliasing; we give the obstruction explicitly. At a Gaussian anchor, for Sobolev smoothness $s$ and $n$ samples per time point, we derive a conditional lower rate $(nK)^{-2s/(2s+m+1)}$ for the tangent snapshot experiment, with a matching upper rate in a degreewise benchmark. Strong-form fitting is non-orthogonal to score error and cannot be repaired by spectral filtering. Instead, we estimate using smooth test functions while retaining the known diffusion term, and derive a finite-sample bound that separates sampling error from fixed-grid quadrature bias. Planted-circulation experiments confirm the predicted gauge contraction and expose a design tension between cross-slice information and covariance-aware whitening.
Chinese Translation
从演化分布的快照推断动力学本质上是欠定的:Fokker-Planck 方程仅通过其得分加权散度 $\nabla\cdot F+F\cdot\nabla\logρ$ 约束漂移 $F$,留下一个对任何单时刻约束都不可见的 $ρ$-螺线管规范。按时间索引的输运表述无法消除这种歧义:每一条可容许的边缘路径都允许一种无旋解释,最小作用量重建会选择它,而仅凭边缘拟合无法区分动力学上不等价的解释。要求一个自治场解释多个边缘分布,反而会使部分隐藏环流变得可见,因为 $\nabla\logρ$ 在不同边缘分布之间发生变化。将瞬时 Fokker-Planck 源约束与快照实验分离,我们证明这些源约束在模去堆叠得分加权散度算子的核的意义下辨识该场。对于一般高斯形状变化,在内蕴维数 $m$ 中,$K\ge m$ 个时间点处的源约束会消除每一个多项式规范方向,而仅有限多个密度快照本身则允许混叠;我们明确给出该障碍。在高斯锚点处,对于 Sobolev 光滑度 $s$ 和每个时间点的 $n$ 个样本,我们推导出切向快照实验的条件下界速率 $(nK)^{-2s/(2s+m+1)}$,并在逐阶基准中给出匹配的上界速率。强形式拟合与得分误差非正交,且无法通过谱滤波修正。相反,我们使用光滑检验函数进行估计,同时保留已知的扩散项,并推导出一个有限样本界,该界将采样误差与固定网格求积偏差分离开来。植入环流实验证实了所预测的规范收缩,并揭示了跨切片信息与协方差感知白化之间的设计张力。
cs.LG / 171 / 2609.35084
GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents
GraphHCA:面向长时程 LLM 智能体的闭式后见之明信用分配
Haodong Zhu, Yangyang Ren, Changbai Li, Sheng Xu, Linlin Yang, haiguang liu, Baochang Zhang
cs.LG
large language model
大语言模型相关
Abstract
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes' rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.
Chinese Translation
基于分组的强化学习(RL)推动了大语言模型(LLM)的发展,并正日益扩展到智能体任务中,在这类任务中,稀疏的终端奖励使得步骤级信用分配至关重要。现有方法根据采样 rollout 中某一动作之后所发生的内容来分配信用,但并未显式地刻画其与已实现结果之间的回溯性关系。后见之明信用分配(HCA)则通过后见之明概率与行为策略概率之比来归因信用,但估计后见之明分布需要辅助模型或额外的一次前向传递。为解决这一估计瓶颈,我们提出 GraphHCA,一种无需模型的 HCA 实现,它消除了对后见之明分布的显式估计。对于具有确定性转移的终端目标任务,贝叶斯法则将后见之明比率化简为相邻状态处行为策略成功概率之比。取对数可得到逐状态的成功势能,其跨越一次转移的增量提供了步骤级信用。GraphHCA 通过在诱导转移图上进行折扣递归,从汇总的 rollout 中估计该势能,该递归在任意有向图上都允许唯一的不动点。由此得到的步骤级信号与轨迹级优势相结合,既不需要学习得到的后见之明模型,也不需要额外的前向传递,并且在步骤级权重为零时退化为 GRPO。在所有比较的基线中,GraphHCA 在两种 LLM 规模下的 ALFWorld 和 WebShop 上,以及在配备视觉-语言智能体的 Sokoban 上,均取得了最先进的结果。例如,在 ALFWorld 上,它相较于 GRPO 将总体成功率提升最多 24.6 个百分点,相较于最强的步骤级基线提升最多 4.7 个百分点。
cs.LG / 172 / 2609.35086
Retrieval-Augmented Diffusion Modeling for Stochastic Discount Factor Portfolios
面向随机折现因子投资组合的检索增强扩散建模
Kelvin J. L. Koa, Xinyang Li, Ke-Wei Huang
cs.LG · q-fin.CP · q-fin.PM
diffusion
扩散模型相关
Abstract
In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial markets exhibit non-stationary dynamics with shifting regimes, multimodal inputs such as price and news data often contain stochastic noise, and existing diffusion-based approaches, while effective for modeling stochastic dynamics, rely on assumptions such as isotropic Gaussian noise that fail to capture the state-dependent nature of financial uncertainty. To address these challenges, we introduce RADAR, a retrieval-augmented diffusion framework that learns market representations by conditioning on similar historical regimes. RADAR leverages retrieval to construct context-dependent noise distributions, applies conditional diffusion to denoise multimodal representations, and initializes the diffusion process using empirical statistics to reflect state-dependent uncertainty. Experiments show that RADAR achieves state-of-the-art performance on key risk-adjusted metrics while producing economically meaningful signals on asset returns and correlations.
Chinese Translation
在这项工作中,我们通过学习能够捕捉金融数据潜在风险结构的市场状态表示,研究随机折现因子(SDF)框架下的投资组合优化。这之所以具有挑战性,是由若干因素造成的:金融市场表现出具有区制转换的非平稳动态,诸如价格和新闻数据之类的多模态输入往往包含随机噪声,并且现有的基于扩散的方法虽然对于建模随机动态是有效的,却依赖于诸如各向同性高斯噪声之类的假设,而这些假设无法捕捉金融不确定性的状态依赖性质。为应对这些挑战,我们提出了 RADAR,一种检索增强扩散框架,它通过以相似的历史区制为条件来学习市场表示。RADAR 利用检索来构建依赖上下文的噪声分布,应用条件扩散来对多模态表示去噪,并使用经验统计量初始化扩散过程,以反映状态依赖的不确定性。实验表明,RADAR 在关键的风险调整指标上达到了最先进的性能,同时在资产收益和相关性上产生了具有经济意义的信号。
cs.LG / 173 / 2609.35166
Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion
学习重新起草:面向离散扩散的变分斯塔克尔伯格博弈
Dmitrii Moor, Federico Tomasi, Paul N. Bennett, Alice Wang, Mounia Lalmas
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser's token embeddings, while the follower optimizes a variational denoising objective with the corruption process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily the current denoiser can reconstruct them. We measure this improvement under a fixed reference corruption process, approximate the follower's response with a one-step gradient update, and optimize the leader using a score-function estimator. We evaluate VSDD across molecular, text, and playlist generation. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizable improvements in offline playlist recommendation metrics.
Chinese Translation
离散扩散模型提供了重新起草的能力,在生成过程中重新审视并修正较早的 token。这种能力取决于前向损坏过程,该过程定义了去噪器学习去纠正什么。掩码扩散模型一旦 token 被解除掩码就将其固定,而均匀扩散允许修订,但依赖于均匀随机的 token 替换。我们转而学习哪些替换对于训练去噪器进行重新起草最有用。我们提出变分斯塔克尔伯格离散扩散(VSDD),一个用于学习语义感知损坏过程的框架。VSDD 将训练表述为一个领导者-跟随者博弈:领导者定义一个由去噪器的 token 嵌入参数化的马尔可夫损坏过程,而跟随者在损坏过程保持固定的情况下优化一个变分去噪目标。领导者根据去噪器在从这些损坏中学习后提升了多少来奖励这些损坏,而不是根据当前去噪器能够多容易地重建它们来奖励。我们在固定的参考损坏过程下衡量这种改进,用一步梯度更新近似跟随者的响应,并使用得分函数估计器优化领导者。我们在分子、文本和播放列表生成上评估 VSDD。VSDD 在分子有效性方面显著优于均匀扩散和掩码扩散,相对于均匀扩散降低了文本困惑度,同时与掩码扩散保持竞争力,并在离线播放列表推荐指标上取得了可观提升。
cs.LG / 174 / 2609.35224
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
TANGO:在词元对中为掩码扩散语言模型添加水印
Kasra Arabi, Nir Weinberger, Micah Goldblum, Niv Cohen
cs.LG · cs.CL · cs.CR
diffusion
扩散模型相关
Abstract
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token frequencies in watermarked and unwatermarked text can recover the list and forge text that the provider's own detector accepts. We present TANGO, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked. A secret key splits the vocabulary into color classes, and TANGO biases the new token toward a color determined by the key and the nearby token's color. The watermark is therefore embedded in pairs of tokens. Because the favored color changes from position to position, token frequencies stay much closer to those of unwatermarked text than under a fixed green list. Detection needs only the text and the key, and it does not assume any unmasking order. On two masked-diffusion models, TANGO detects nearly all unedited watermarked texts and most edited ones, and frequency attacks that forge the fixed green list fail against it.
Chinese Translation
掩码扩散语言模型并行地、不以固定顺序填充掩码位置。大多数实用的文本水印假设从左到右生成。它们将每个词元与它之前的词元进行键控关联,而在扩散模型中,那些词元可能仍处于掩码状态。固定绿色列表不需要这种上下文,但它在每个位置都偏向相同的词元,因此这些词元在水印文本中出现得更频繁。攻击者通过比较水印文本和未加水印文本中的词元频率,可以恢复该列表,并伪造出提供方自己的检测器会接受的文本。我们提出 TANGO,一种用于掩码扩散语言模型的水印,它将每个新词元与一个已经解除掩码的邻近词元进行键控关联。一个秘密密钥将词汇表划分为颜色类别,TANGO 使新词元偏向由该密钥和邻近词元的颜色所确定的一种颜色。因此,该水印被嵌入到词元对中。由于偏好的颜色随位置变化,词元频率会比在固定绿色列表下更接近未加水印文本的词元频率。检测仅需要文本和密钥,并且不假设任何解除掩码的顺序。在两个掩码扩散模型上,TANGO 检测出几乎所有未经编辑的水印文本和大多数经过编辑的水印文本,而伪造固定绿色列表的频率攻击对其无效。
cs.LG / 175 / 2609.35269
eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models
eval-unlearn:对文本到图像扩散模型中的遗忘进行基准测试
Mansi, Nikhil Raghavan, Zixia Huang, Kai Sheng Ong, Ji Shen Lim, Brandon Siao Xiang Ling, Francesco Leofante
cs.LG · cs.AI · cs.CV
diffusion
扩散模型相关
Abstract
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models. eval-unlearn integrates twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention, alongside nine complementary evaluation metrics covering erasure efficacy, adversarial robustness, generative quality, and concept retention. Its plugin architecture lets third-party techniques and metrics self-register without modifying the core framework, and its streaming, batched pipeline supports efficient evaluation of both standard NSFW concepts and arbitrary general concepts. As a further contribution, we release a public leaderboard on HuggingFace along with an interactive tool for real-time evaluation of unlearning techniques. The leaderboard compares nudity concept erasure case study across all twelve techniques, exposing significant accuracy-quality trade-offs that are obscured by heterogeneous evaluation. eval-unlearn is released under the MIT license; the package, code, leaderboard, and documentation are all available at https://eval-unlearn.readthedocs.io.
Chinese Translation
用于文本到图像(T2I)扩散模型的概念遗忘技术数量不断增长,造成了碎片化的评估格局。方法在异构实验条件下被评估,使得有原则的跨方法比较变得困难。我们提出 eval-unlearn,一个开源 Python 库,为 T2I 扩散模型中的概念遗忘提供统一、可复现的基准测试框架。eval-unlearn 集成了十二种已发表的遗忘技术,涵盖微调、闭式模型编辑和推理时干预,以及九个互补的评估指标,覆盖擦除有效性、对抗鲁棒性、生成质量和概念保留。其插件架构允许第三方技术和指标在无需修改核心框架的情况下自行注册,其流式、批处理流水线支持对标准 NSFW 概念和任意通用概念进行高效评估。作为进一步贡献,我们在 HuggingFace 上发布了一个公共排行榜,以及一个用于实时评估遗忘技术的交互式工具。该排行榜比较了所有十二种技术中的裸体概念擦除案例研究,揭示了被异构评估所掩盖的显著准确率-质量权衡。eval-unlearn 在 MIT 许可证下发布;该包、代码、排行榜和文档均可在 https://eval-unlearn.readthedocs.io 获取。
cs.LG / 176 / 2609.35274
Multi-Attractor GNNs: Set-Valued Expressivity Beyond Unique Equilibria
多吸引子图神经网络:超越唯一平衡点的集合值表达能力
Jialin Liu
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Recurrent and equilibrium graph neural networks (GNNs) often enforce a unique fixed point or use one training target per graph. Yet many combinatorial and scientific problems admit multiple valid solutions, with no preferred one. A designated target can then impose an arbitrary selection rule. For tasks invariant to node relabeling, a symmetric graph may have a symmetric solution set but no symmetric solution. We show that multiple equilibria enable one weight-tied message-passing GNN to represent set-valued equivariant maps: different initializations approach different valid solutions. Under stated regularity assumptions, we first construct globally Lipschitz, permutation-equivariant dynamics that converge almost surely to valid solutions and reach every solution branch with positive probability. We then establish approximate realization by recurrent message passing with continuous component maps, with arbitrarily small update and limiting errors and arbitrarily high probability. This goes beyond standard universality arguments: although message passing alone cannot distinguish symmetric nodes, the evolving state keeps nodes distinguishable at every finite step without auxiliary node identifiers. Such dynamics can be learned without solution labels using problem-specific energies. On Ising ground states, structural module detection in protein graphs, and chemical reaction steady states, the learned updates produce multiple high-quality predictions with high numerical convergence rates. They achieve better average solution quality than the tested unique-equilibrium, single-target, and feedforward baselines, while remaining competitive with much larger diffusion-based solvers.
Chinese Translation
循环和平衡图神经网络(GNN)通常强制唯一不动点,或为每张图使用一个训练目标。然而,许多组合问题和科学问题允许多个有效解,且没有哪一个更受偏好。此时,一个指定的目标会强加一种任意的选择规则。对于对节点重标记不变的任务,一个对称图可能具有对称的解集,却没有对称解。我们表明,多重平衡态使一个权重绑定的消息传递 GNN 能够表示集合值等变映射:不同的初始化会趋近不同的有效解。在所述正则性假设下,我们首先构造全局 Lipschitz、置换等变的动力学,它们几乎必然收敛到有效解,并以正概率到达每一个解分支。然后,我们证明可用具有连续分量映射的循环消息传递来近似实现,其更新误差和极限误差任意小,且概率任意高。这超越了标准的通用性论证:尽管仅靠消息传递无法区分对称节点,但演化状态在每一个有限步骤都能保持节点可区分,而无需辅助节点标识符。此类动力学可以在没有解标签的情况下,利用问题特定的能量来学习。在 Ising 基态、蛋白质图中的结构模块检测以及化学反应稳态上,学习到的更新产生多个高质量预测,并具有高数值收敛率。它们取得了比所测试的唯一平衡、单目标和前馈基线更好的平均解质量,同时与规模大得多的基于扩散的求解器相比仍具竞争力。
cs.LG / 177 / 2609.35315
Collaborative Principle Evolution via Evidence Transfer for Scientific Discovery
面向科学发现的通过证据迁移实现的协作式原理演化
Yingming Pu, Hongyu Chen, Tao Lin
cs.LG
large language model
大语言模型相关
Abstract
Large Language Model (LLM)-based agents promise to automate scientific discovery, yet exploring the vast hypothesis space remains costly. Existing principle-evolution methods accelerate this loop, but operate sequentially, which caps exploration breadth and wastes wall-clock time on challenging problems. To address this, we formulate collaborative scientific discovery as evidence transfer between parallel principle-evolution branches. We present COEVOLVE, which realizes this transfer through a coordination core over parallel branches. By integrating value-of-information-gated routing and context-discounted likelihood injection, COEVOLVE enables branches to collaborate through shared measurements while keeping their principle posteriors separate. Across six scientific-discovery tasks under a matched evaluation budget, COEVOLVE attains a mean solution quality of 66.5% versus 57.0% for single-branch principle evolution, with a 1.80x mean wall-clock speedup on the GPT-5.6-Terra backbone; on five auto-research tasks delegated to an autonomous research harness, it is the only arm whose mean stays above the published SOTA anchor on every task. These results establish when evidence sharing accelerates parallel discovery and when transfer safeguards are necessary to limit negative or inert transfers
Chinese Translation
基于大语言模型(LLM)的智能体有望实现科学发现的自动化,然而探索庞大的假设空间仍然代价高昂。现有的原理演化方法加速了这一循环,但其以串行方式运行,这限制了探索广度,并在困难问题上浪费了墙钟时间。为解决这一问题,我们将协作式科学发现形式化为并行原理演化分支之间的证据迁移。我们提出 COEVOLVE,它通过作用于并行分支之上的协调核心来实现这种迁移。通过整合信息价值门控路由与上下文折扣的似然注入,COEVOLVE 使各分支能够通过共享测量结果进行协作,同时保持各自的原理后验相互分离。在匹配评估预算下的六个科学发现任务中,COEVOLVE 取得了 66.5% 的平均解质量,而单分支原理演化为 57.0%,并在 GPT-5.6-Terra 主干模型上实现了 1.80 倍的平均墙钟加速;在委托给自主研究工具链的五个自动研究任务上,它是唯一一个在所有任务上平均值均保持在已发表 SOTA 基准之上的分支。这些结果确立了证据共享何时会加速并行发现,以及何时需要迁移防护机制来限制负向迁移或惰性迁移。
cs.LG / 178 / 2609.35335
Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation
用于自动化跨域机器学习任务类型识别的大型语言模型:基准数据集与评估
Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.
Chinese Translation
机器学习任务类型识别对于构建有效的 ML 流水线至关重要,但在实践中通常由人工指定。我们研究大型语言模型(LLMs)是否能够在用户仅提供目标特征时,直接从数据集级信息中推断数据领域和下游预测任务。结合我们基于 LLM 的系统,我们还发布了一个包含 625 个公开表格和时间序列数据集的带标注基准。我们在三种设置下评估所提出的方法:(i) 表格数据集,与已有的 AutoML 启发式方法进行比较,(ii) 跨表格和时间序列数据集的跨域评估,以及 (iii) 使用更小的本地模型的实际部署场景。结果表明,基于 LLM 的任务类型识别具有一致优势,而在异构和资源受限设置中难度不断增加。基于 LLM 的方法在表格设置中优于 AutoGluon,达到 0.98 的 F1 macro,而后者为 0.93。在跨域设置中,最佳模型达到 0.90 的 F1 macro,而更小的可本地部署模型达到 0.75,表明在部署可行性与准确性之间存在权衡。
cs.LG / 179 / 2609.35362
d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
d-OPD:面向块扩散语言模型的未来感知同策略蒸馏
Ruitao Liu, Qinghao Hu, Song Han
cs.LG · cs.AI
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to $4.0$ points over OPDLM and reduces training time by $1.35$-$1.58\times$. The code is available at https://github.com/mit-han-lab/d-OPD.
Chinese Translation
大语言模型(LLM)通常以自回归(AR)方式生成文本,每次预测一个 token。块扩散语言模型(dLLM)则按顺序生成块,同时在每个块内并行去噪多个 token,为加速生成提供了一种有前景的方式。近期工作不是从头训练此类模型,而是通过蒸馏将强大的预训练 AR 模型适配为块 dLLM。同策略蒸馏(OPD)已被广泛用于 LLM 训练,因为它基于学生模型当前策略生成的状态对学生模型进行监督,而不仅仅基于固定的离线轨迹。通过在学生模型实际访问的状态上训练,它减少了训练与生成之间的不匹配,并且随着学生模型演进,能够提供更相关的监督。近期工作已将这一思想扩展到 AR 到块扩散的转换。然而,这种设置在监督中引入了根本性的不匹配:块扩散学生模型和因果 AR 教师模型在同一训练状态下所依据的信息不同。学生模型从整个部分去噪的块中进行预测,包括可见的未来上下文,而标准 AR 教师目标仅由因果前缀定义。因此,用于蒸馏的教师分布并未与学生模型可获得的信息完全对齐。因此,我们提出 d-OPD,一种未来感知的同策略蒸馏方法,它通过纳入每个块内可见的未来信息来校正 AR 教师分布,使其更好地与对学生可见的状态对齐,从而提供与学生所用信息更匹配的监督。在从 0.6B 到 8B 的 Qwen3 模型上,d-OPD 相比 OPDLM 将六项基准的平均值提升了最高 $4.0$ 分,并将训练时间缩短了 $1.35$-$1.58\times$。代码可在 https://github.com/mit-han-lab/d-OPD 获取。
cs.LG / 180 / 2609.35377
First Learn, Then Memorize: The Spectral Bias of Diffusion Models
先学习,后记忆:扩散模型的谱偏置
Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard
cs.LG · cond-mat.dis-nn
diffusion
扩散模型相关
Abstract
Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample ($m$ noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for $m=1$. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size $n$. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear ($n \asymp d$) and polynomial ($n \asymp d^k$) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank $r$ tunes the generalization--memorization transition, and an $L_2$ penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.
Chinese Translation
在有限数据集上训练的扩散模型首先学会生成新颖、高质量的样本,而直到很久之后才坍缩到其训练集上。我们揭示了这种时间尺度分离背后的机制以及探测它的对象。得分函数的训练动力学——精确地,并且在任意宽度下——由在含噪训练数据上求值的神经正切核(NTK)的 Gram 矩阵所支配,因此泛化与记忆的时间尺度必然编码在其谱中。我们表明它们确实如此,并且负责这一结构在标准核设置中没有类似物。在得分匹配损失中,对每个样本使用多个噪声实现($m$ 个固定噪声水平下的加噪副本)正是将 Gram 矩阵谱重构为两个不同部分的原因。第一个部分由大特征值构成,承载目标分布的全局特征,并且对于 $m=1$ 已经存在。第二个部分由重复加噪产生,由最小特征值构成,并支撑在与样本特定噪声方向对齐的特征向量上;它设定了一个在训练集大小 $n$ 上参数化地更大的记忆时间尺度。我们在两个方面确立了这一图景。在解析上,我们求解了惰性高维极限中的谱,涵盖线性($n \asymp d$)和多项式($n \asymp d^k$)样本复杂度,并通过偏差--方差分解证明,第一个主体最小化近似误差,而第二个主体驱动与记忆相关的误差。在经验上,我们在 CelebA 上的卷积 NTK 中,以及在远超惰性区域训练的有限宽度 U-Net 中,展示了相同的双主体结构,并且我们使这种联系具有因果性:将 Gram 矩阵截断到秩 $r$ 会调节泛化--记忆转变,而针对第二个主体的 $L_2$ 惩罚会抑制特征学习 U-Net 中的记忆。
cs.LG / 181 / 2609.35426
Frontier Learning: Training LLM Reasoners at the Edge of Capability
前沿学习:在能力边缘训练 LLM 推理器
Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi
cs.LG · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
Chinese Translation
基于强化学习的大型语言模型(LLM)后训练已被成功应用于提升其推理能力。现有流程主要使用 GRPO 损失,在训练前指定的固定问题池上对 LLM 进行微调。这具有根本性的局限,因为只有当策略 rollout 混合了成功与失败时才会产生学习信号,这导致任何固定问题池中的有用部分都会随着模型的改进而迅速过时。为解决这一问题,我们提出了前沿学习,一种开放式后训练方法,其中程序化生成器被在线使用,以持续产生具有信息量的训练问题。它将生成器的任务特定参数视为一个搜索空间,并使用遗憾信号来优先排序并探索前沿难度水平,从而将训练聚焦于模型不断演进的推理能力的边缘。在多个推理任务和模型系列上,我们的方法相较于固定问题池基线持续取得更高的相对增益,表明有效的后训练不仅需要选择有用的问题,还需要在能力的边缘持续生成它们。
cs.LG / 182 / 2609.35440
SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
SOLO:以共享输出局部学习预训练十亿参数语言模型
Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li
cs.LG · cs.DC
large language model
大语言模型相关
Abstract
Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness, since they leave each module without information from deeper modules. We propose Shared-Output LOcal learning (SOLO), which replaces them with a shared, read-only copy of the final module's readout, the only one trained on the output of the whole network. Taken from the previous step, the copy transmits information from the final module without passing gradients between modules or reintroducing update locking. SOLO approaches backpropagation on Transformers of 340M to 2B parameters pretrained on 15B tokens, staying within one point in average zero-shot accuracy with a perplexity gap that narrows with scale. Readout ablations attribute SOLO's improvement over private readouts to sharing. Without update locking, each of p pipeline stages holds activations for O(1) micro-batches instead of O(p). The freed memory permits larger micro-batches, which reach up to 1.44x the best measured throughput of pipeline backpropagation on the same partition. To our knowledge, SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining. Local learning thus becomes a practical alternative to backpropagation for large-scale pretraining.
Chinese Translation
大型语言模型通过反向传播进行训练,其全局梯度协调所有层,但迫使每一层都保留其激活值,并等待梯度回传穿过每一个更深的层。传统的局部学习通过训练每个模块经由其自身的读出层来预测目标,从而消除了这种更新锁定,但尚未扩展到十亿参数的预训练。我们认定这些私有读出层是一个关键弱点,因为它们使每个模块都无法获得来自更深层模块的信息。我们提出共享输出局部学习(SOLO),它用最终模块读出层的一个共享、只读副本取代这些私有读出层,该读出层是唯一在整个网络输出上训练的读出层。该副本取自前一步,它在不于模块之间传递梯度、也不重新引入更新锁定的情况下,传递来自最终模块的信息。SOLO 在 340M 到 2B 参数的 Transformer 上、基于 15B token 预训练时逼近反向传播,平均零样本准确率差距保持在一个点以内,且困惑度差距随规模增大而缩小。读出层消融实验将 SOLO 相较私有读出层的改进归因于共享。由于没有更新锁定,p 个流水线阶段中的每一个都只需为 O(1) 个微批次保留激活值,而非 O(p)。被释放的内存允许更大的微批次,其在同一划分上最高达到流水线反向传播最佳实测吞吐量的 1.44 倍。据我们所知,SOLO 是首个在十亿参数语言模型预训练中展现出此类内存与吞吐量收益的局部学习方法。因此,局部学习成为大规模预训练中反向传播的一种实用替代方案。
cs.LG / 183 / 2609.35445
NeuronSifter: Intervention Planning in CNS Microenvironments
NeuronSifter:中枢神经系统微环境中的干预规划
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
cs.LG · q-bio.NC
diffusion
扩散模型相关
Abstract
Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar exposure, discarding where and when the target is engaged; handing a point estimate to a separate planner then discards the joint uncertainty that makes a measurement worth running. We therefore treat decision quality as a property of the intervention interface, not of controller placement. NeuronSifter compiles regimens into state-conditional target-occupancy fields with support masks, propagates them through microenvironment dynamics with an occupancy-conditioned diffusion operator, and selects measurements by their expected reduction in intervention loss, assimilating typed outcomes into the same posterior. In a declared synthetic Alzheimer's disease (AD) evaluation over 64 paired scenario blocks, occupancy conditioning lowers trajectory continuous ranked probability score from 0.165 to 0.110 and raises intervention ordering accuracy from 0.760 to 0.880, and every paired benchmark contrast remains separated after Holm correction. Decision-directed acquisition attains terminal risk 0.160 against 0.166 for a matched numerical Bayesian experimental design planner, and reaches the target risk at 0.796 $[0.732,0.873]$ of an earlier design control's cost, while the corresponding ratio against the matched planner, 0.963 $[0.907,1.025]$, is not separated from equality; point-state and dependence-ablated interfaces instead raise risk to 0.220 and 0.199, and a full-posterior external controller ties exactly. Published AD trials supply a separate retrospective endpoint bridge.
Chinese Translation
对中枢神经系统(CNS)干预进行优先级排序,需要预测剂量、给药途径和给药时程如何作用于部分观测到的微环境,然后选择会改变决策的测量。以动作为条件的预测器将方案简化为一个身份标记或一个标量暴露,丢弃了靶标在何处、何时被作用;将点估计交给单独的规划器,又会丢弃使某项测量值得执行的那种联合不确定性。因此,我们将决策质量视为干预接口的一种属性,而不是控制器放置方式的属性。NeuronSifter 将方案编译为带有支持掩码的状态条件靶标占据场,用占据条件扩散算子将它们通过微环境动力学传播,并根据其对干预损失的期望减少量来选择测量,同时将类型化结果同化到同一后验中。在一项预先声明的、覆盖 64 个配对情景块的合成阿尔茨海默病(AD)评估中,占据条件化将轨迹连续排序概率评分从 0.165 降至 0.110,并将干预排序准确率从 0.760 提高到 0.880,而且在 Holm 校正后,每一项配对基准对比仍保持分离。以决策为导向的采集达到终端风险 0.160,而匹配的数值贝叶斯实验设计规划器为 0.166;并且在较早设计对照成本的 0.796 $[0.732,0.873]$ 处达到目标风险,而相对于匹配规划器的相应比值 0.963 $[0.907,1.025]$ 与相等未能分离;相反,点状态接口和依赖消融接口将风险分别提高到 0.220 和 0.199,而全后验外部控制器则完全打平。已发表的 AD 试验提供了一个单独的回溯性终点桥接。
cs.LG / 184 / 2609.35499
Physics-Guided Conditional Diffusion Model for Rare Event Synthesis and Diagnosis for the Water-Gas Shift Reaction
用于水煤气变换反应稀有事件合成与诊断的物理引导条件扩散模型
Md Abrar Rafid Siddique, Bibek Aryal, Qiugang Lu
cs.LG
diffusion
扩散模型相关
Abstract
As the world moves towards sustainable energy sources, hydrogen (H2) can be treated as an eco-friendly alternative to fossil fuels due to its high energy density and zero carbon emissions. The water-gas shift (WGS) reaction is a widely used industrial process for hydrogen production by converting carbon monoxide and steam into hydrogen and carbon dioxide. However, occurrences like severe fouling, catalyst deterioration, and thermal runaway can hamper the reaction kinetics/process safety and decrease the yield of H2. These incidents are rare, and gathering process data under such abnormal conditions is challenging. In this work, we propose a physics-guided conditional diffusion model to generate realistic rare-event trajectories for the WGS reaction. The proposed model integrates a conditional denoising diffusion probabilistic model (CDDPM) with governing laws of the reaction to generate physically consistent process trajectories. The conditioning features allow the model to produce high-quality synthetic profiles for rare-event domains that are typically beyond the training regimes. The generated rare-event trajectories then augment the raw dataset for a balanced distribution between normal and abnormal conditions. We further propose a hazard score to assess the risk severity of the operating condition based on the operating trajectory. Deep learning models are trained with the augmented dataset to diagnose the health status of the reaction. Simulation results show that the proposed physics-guided diffusion model outperforms data-driven models in terms of the quality of synthetic data and diagnosis performance for rare events.
Chinese Translation
随着世界向可持续能源转型,氢气(H2)因其高能量密度和零碳排放,可被视为化石燃料的环保替代品。水煤气变换(WGS)反应是一种广泛使用的工业过程,通过将一氧化碳和蒸汽转化为氢气和二氧化碳来制氢。然而,严重结垢、催化剂劣化和热失控等事件会阻碍反应动力学/过程安全,并降低H2产率。这些事件很少发生,并且在此类异常条件下收集过程数据具有挑战性。在这项工作中,我们提出了一种物理引导的条件扩散模型,用于生成WGS反应的逼真稀有事件轨迹。所提出的模型将条件去噪扩散概率模型(CDDPM)与反应的控制定律相结合,以生成物理上一致的过程轨迹。条件特征使模型能够为通常超出训练范围的稀有事件域生成高质量的合成曲线。然后,生成的稀有事件轨迹对原始数据集进行增强,以在正常和异常条件之间实现平衡分布。我们进一步提出了一种危害评分,以基于操作轨迹评估操作条件的风险严重程度。使用增强后的数据集训练深度学习模型,以诊断反应的健康状态。仿真结果表明,所提出的物理引导扩散模型在合成数据质量和稀有事件诊断性能方面优于数据驱动模型。
cs.LG / 185 / 2609.35553
Simplex Diffusion Models
单纯形扩散模型
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
cs.LG · stat.ML
diffusion
扩散模型相关
Abstract
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).
Chinese Translation
扩散模型通过逐步细化信念状态,已经彻底改变了连续数据的生成建模。这种迭代细化尚未延续到离散扩散模型,后者通过类别采样在中间步骤丢弃不确定性(信息坍缩)。我们提出单纯形扩散模型(Simplex Diffusion Models, SDMs),这是一个将扩散过程提升到概率单纯形以表示类别上的信念的框架。SDMs 允许具有闭式反向转移的概率路径,并可以用简单的交叉熵损失进行训练。与先前如狄利克雷流匹配(Dirichlet Flow Matching)等需要积分一个常微分方程的方案相反,我们引入了一种类似 DDIM 的采样器,具有可调的随机性水平。因为 SDMs 对单纯形上的样本进行操作,它们能够在去噪步骤间携带不确定性,从而缓解信息坍缩。在 OpenWebText 上,SDMs 与强离散扩散基线相比具有竞争力,在 64 个采样步骤中达到 $17.0$ GenPPL 和 $5.46$ unigram 熵,接近真实验证数据。即使没有自条件(Self-Conditioning, SC),SDMs 在代码生成(TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$)上也优于掩码扩散和均匀扩散(使用 SC 或预测器-校正器采样)。蒸馏到 8 步后,SDMs 解决了 $32.1\%$ 的 GSM8K 问题,超过了具有 128 步的蒸馏离散扩散模型($21.4\%$)。
cs.LG / 186 / 2609.35579
Output-aware Residual Stream Pruning for Large Language Models
面向大语言模型的输出感知残差流剪枝
Chayne Thrash, Kevin Chen, Soheil Kolouri
cs.LG
large language model
大语言模型相关
Abstract
Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.
Chinese Translation
残差流剪枝方法通过收缩模型的隐藏维度来降低推理成本,但现有方法通常通过最小化激活重构误差来选择这些维度。该准则隐含地将所有扰动方向视为同等重要,忽略了下游层的敏感性。我们引入一种敏感性感知的残差流剪枝方法,直接刻画这种依赖于方向的敏感性。利用对输出 KL 散度的二阶近似,我们通过残差流扰动的激活协方差以及模型输出的局部敏感性这两者来刻画该扰动的影响。由此得到的子空间选择目标耦合了这两个量,但难以直接优化。我们推导出一个可处理的谱上界,将子空间选择归结为对敏感性加权协方差矩阵的特征分解,从而保持了基于旋转的剪枝方法的高效性与结构简洁性。在多个指令微调的语言模型系列上,相较于仅基于激活的剪枝,我们的方法一致地降低了校准 KL 散度,并在多种压缩程度下改善了困惑度与下游任务性能。我们的结果表明,仅保留激活能量对于残差流剪枝而言并不充分,而显式地考虑扰动如何传播至模型输出,则为选择要移除的维度提供了更有效的准则。
cs.LG / 187 / 2609.35609
Twist, Don't Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models
扭转,而非倾斜:面向掩码扩散模型的轨迹精确约束解码
Aditya Thimmaiah, Lara Marinov, Jayanth Srinivasa, Haris Vikalo, Junyi Jessy Li, Milos Gligoric
cs.LG · cs.AI · cs.CL
diffusion
扩散模型相关
Abstract
Constrained decoding for Masked Diffusion Language Models (MDLMs) aims to ensure that generated outputs satisfy a specified structure or syntax constraint. MDLMs generate outputs by repeatedly unmasking masked positions present in their current state. Recent strategies for constrained decoding constrain the model's per-step mean-field posterior (which factorizes over masked positions) by enforcing the desired constraint with an automaton. The resulting chain-structured factor graph allows exact constrained sampling via dynamic programming. However, despite each draw being exact and constraint-satisfying, we prove that their composition, in general, tilts away from the model's relative probabilities over valid trajectories, thus leading to trajectory bias. We derive an exact expression for this bias as a product of ratios measuring how valid continuation mass changes when the denoiser is reconditioned, and characterize when the bias vanishes. We then correct the bias by introducing TWISTER, the first automaton-twisted Sequential Monte Carlo decoder for MDLMs, using the step-exact decoder as the proposal. We show that for regular language constraints, the Feynman-Kac correction is exactly computable, with the twists obtained efficiently using quantities pre-computed for step-exact sampling. We prove that the resulting Feynman-Kac model targets the unbiased Doob h-transformed path law conditioned on constraint satisfaction.
Chinese Translation
掩码扩散语言模型(MDLMs)的约束解码旨在确保生成的输出满足指定的结构或语法约束。MDLM 通过反复对其当前状态中存在的被掩码位置进行解掩码来生成输出。近期的约束解码策略通过用自动机强制施加所需约束,来约束模型每步的平均场后验(该后验在被掩码位置上因子分解)。由此得到的链式结构因子图允许通过动态规划进行精确的约束采样。然而,尽管每次抽取都是精确且满足约束的,我们证明它们的组合一般而言会偏离模型在有效轨迹上的相对概率,从而导致轨迹偏差。我们为该偏差推导出一个精确表达式,其形式为若干比值之积,这些比值衡量当去噪器被重新条件化时有效延续质量如何变化,并刻画该偏差何时消失。随后,我们通过引入 TWISTER 来校正该偏差,TWISTER 是首个用于 MDLM 的自动机扭转式序列蒙特卡洛解码器,其使用步精确解码器作为提议分布。我们表明,对于正则语言约束,Feynman-Kac 校正可以精确计算,其中扭转可通过为步精确采样预计算得到的量高效获得。我们证明,所得的 Feynman-Kac 模型所针对的是以约束满足为条件的无偏 Doob h-变换路径律。
cs.LG / 188 / 2609.35615
Behavioral Foundation Models for Quality Diversity
面向质量多样性的行为基础模型
Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all tested parameter-space QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.
Chinese Translation
行为基础模型(Behavioral Foundation Models, BFMs)是强化学习中的一种新兴范式,其作用类似于自然语言处理中的大语言模型:它们已展现出显著的多功能性,能够实现零样本性能、快速模仿和在线适应,而这一切都通过利用潜在空间的结构来实现。在这项工作中,我们研究由 BFMs 诱导的潜在行为空间是否能够作为一个有效的搜索空间,以通过质量-多样性(Quality-Diversity, QD)方法发现大量行为多样且高性能的策略。尽管 QD 方法通常直接在高维策略参数空间中进行搜索,但在本文中,我们提出 BFM-QD,这是一个在 BFM 的紧凑潜在空间中执行 QD 搜索的框架。我们进一步表明,BFM-QD 框架提供了一种闭式、无梯度的策略改进算子,它近似于策略梯度更新,但不需要 critic 训练,也不需要反向传播。在涵盖密集运动、稀疏导航和接触丰富操作的连续控制基准上,BFM-QD 始终优于参数空间基线,在稀疏和欺骗性设置中增益尤为显著,在这些设置中,所有测试过的参数空间 QD 方法都崩溃至接近零的性能。这些结果显示了 BFM-QD 框架的有效性,它受益于搜索空间降维与来自多样行为数据的离线预训练之间的协同作用。这将 BFMs 定位为 QD 优化的通用骨干,将其用途从零样本任务求解扩展到发现多样化的行为库。
cs.LG / 189 / 2609.35621
Cartridges++: KV Cache Compression without Off-Context Derailment
Cartridges++:避免上下文外脱轨的 KV 缓存压缩
Sonia Laguna, Joao Monteiro, Marco Cuturi, Pierre Ablin, Eleonora Gualdoni
cs.LG
large language model
大语言模型相关
Abstract
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference time. Methods to obtain CKVs range from drop mechanisms that reduce their number of columns, to learned approaches. Among the latter, Cartridges have emerged as a leading compression method, learning compact KV representations through distillation on relevant Q/A pairs. While existing evaluations focus primarily on whether Cartridges and other CKVs yield approximately similar responses to document-related, on-context queries, we investigate the crucial deployment question of whether they can handle off-context queries, something the native KV representation is particularly good at, thanks to the mechanics of attention. We observe a fundamental trade-off: while Cartridges perform better for on-context queries, heuristic-variants preserve better the original LLM's ability to operate off-context. We measure this through their capability to avoid context contamination in their response, retain general knowledge, and follow instructions. We propose Cartridges++, simple modifications to cartridges that retain off-context abilities at small or negligible cost. The router variant decides at inference time whether the query should use the learned long-context memory, while the data-mixing variant allocates a small fraction of training Q/As to queries outside the reference long document. Our study shows that assessing CKVs on document utility alone can mask substantial degradation in broader model capabilities, yet those issues can be fixed with benign changes to CKV inference or training.
Chinese Translation
反复向大型语言模型(LLM)提供长文档的服务成本高昂:计算量随上下文长度增长,而键值(KV)缓存的内存占用急剧膨胀。压缩 KV(CKV)表示旨在模拟文档的缓存,并且通常在推理时间之前一次性计算完成。获取 CKV 的方法从减少其列数的丢弃机制,到学习式方法不等。在后一类方法中,Cartridges 已成为一种领先的压缩方法,通过在相关 Q/A 对上进行蒸馏来学习紧凑的 KV 表示。尽管现有评估主要关注 Cartridges 和其他 CKV 是否能对与文档相关的、上下文内查询产生大致相似的响应,我们研究了一个关键的部署问题:它们能否处理上下文外查询,而由于注意力机制,原生 KV 表示尤其擅长这一点。我们观察到一种根本性的权衡:虽然 Cartridges 在上下文内查询上表现更好,启发式变体则更好地保留了原始 LLM 在上下文外运行的能力。我们通过它们在响应中避免上下文污染、保留通用知识以及遵循指令的能力来衡量这一点。我们提出 Cartridges++,这是对 Cartridges 的简单修改,以很小或可忽略的成本保留上下文外能力。路由变体在推理时决定查询是否应使用学习到的长上下文记忆,而数据混合变体则将一小部分训练 Q/A 分配给参考长文档之外的查询。我们的研究表明,仅根据文档效用评估 CKV 可能会掩盖更广泛模型能力的大幅退化,然而这些问题可以通过对 CKV 推理或训练进行良性修改来修复。
cs.LG / 190 / 2609.35634
DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising
DR-net-Mamba:用于长程ECG时间序列去噪的选择性状态空间建模
Basile Morel, Samuel Ruiperez-Campillo, Andreas P. Streich, Julia E. Vogt, Thomas Hofmann
cs.LG · cs.AI
diffusion
扩散模型相关
Abstract
Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically with sequence length, and diffusion-based approaches incur prohibitive inference cost. We propose a Mamba-augmented model that inserts selective state-space blocks at the convolutional bottleneck, combining local feature extraction with long-range temporal modeling at linear complexity. We comprehensively evaluate the proposed model with respect to reconstruction fidelity, noise robustness, recording-length scaling, and downstream diagnostic classification across over 40 pathology classes. On synthetic and real datasets, our model achieves the highest SNR and lowest RMSE, with the Mamba advantage increasing with sequence length and in low-SNR regimes. On classification with two independent classifiers, the proposed Mamba-based models achieve the best macro AUROC among all denoisers and improve over their convolutional base models. Calibration is more nuanced and classifier-dependent: denoising improves Binary Cross-Entropy and Brier score on Inception1D but often fails to beat the noisy input on ResNet1D-Wang, and the lead-specific Mamba variant is the only denoiser to improve both calibration metrics over the noisy baseline on both classifiers. Per-class analysis reveals a morphology-dependent benefit: Mamba substantially improves ST/T-change diagnoses, which depend on broad, context-sensitive waveforms.
Chinese Translation
心电图(ECG)记录受到非平稳噪声源的破坏,这些噪声源会降低诊断可靠性,尤其是在动态和长时间记录中。深度学习去噪器已经存在,但卷积架构受限于其感受野,基于Transformer的模型随序列长度呈二次方扩展,而基于扩散的方法则带来难以承受的推理成本。我们提出一种Mamba增强模型,该模型在卷积瓶颈处插入选择性状态空间块,以线性复杂度将局部特征提取与长程时间建模相结合。我们针对重建保真度、噪声鲁棒性、记录长度缩放以及跨越40多个病理类别的下游诊断分类,对所提出的模型进行了全面评估。在合成和真实数据集上,我们的模型实现了最高SNR和最低RMSE,其中Mamba的优势随序列长度增加以及在低SNR情形下而增大。在使用两个独立分类器进行分类时,所提出的基于Mamba的模型在所有去噪器中取得了最佳宏AUROC,并优于其卷积基础模型。校准更加微妙且依赖于分类器:去噪在Inception1D上改善了二元交叉熵和Brier分数,但在ResNet1D-Wang上常常无法优于含噪输入,而导联特定的Mamba变体是唯一一种在两个分类器上相对于含噪基线同时改善两个校准指标的去噪器。逐类分析揭示了一种依赖形态的收益:Mamba显著改善了ST/T改变诊断,而这些诊断依赖于宽泛且对上下文敏感的波形。
cs.LG / 191 / 2609.35695
Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models
重新思考个性化生成:通过因子化排序模型进行测试时对齐
Qiyao Ma, Junshan Zhang, Zhe Zhao
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.
Chinese Translation
将大型语言模型(LLMs)对齐到多样化的用户偏好,从根本上受到以单一化用户为优化目标的标准对齐范式的阻碍。在这项工作中,首先使用实证研究揭示,通过测试时对齐,个性化生成存在巨大的、尚未挖掘的性能提升空间。我们证明,个性化生成特别适合像 Best-of-N(BoN)这样的测试时扩展方法,因为它主要可以被视为一个候选匹配问题,而不是生成器能力瓶颈。尽管奖励模型原则上可以利用这一提升空间,但它们在个性化方面校准不佳,而且其十亿参数规模使对大型候选池进行评分变得成本高得令人望而却步。为克服这一局限,我们提出一个参数高效的框架,采用百万参数规模的多层感知机(MLP)排序模型。我们的个性化排序模型以最小开销直接复用基础生成器的内部嵌入。通过扩展训练时数据以提供细粒度的个性化偏好,这个百万参数排序模型能够准确地对大型候选池进行评分,并能无缝地引导生成,以降低将 N 个候选实际实例化的成本。在涵盖三种个性化生成设置的九个数据集上进行的大量实验表明,我们的个性化排序模型有效利用了所发现的提升空间,在每个数据集上都优于十亿参数的通用奖励模型,且参数量不足其 0.4%,评分延迟低四个数量级。
cs.LG / 192 / 2609.35699
Distillation Defenses Easily Break After Reinforcement Learning
蒸馏防御在强化学习后轻易失效
Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin Gal, Yonatan Gideoni
cs.LG · cs.AI · cs.CR
large language model
大语言模型相关
Abstract
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
Chinese Translation
蒸馏攻击会复制闭源大语言模型的推理能力,使恶意行为者能够以低成本复现最先进的性能。攻击者会系统性地收集大量前沿模型的推理轨迹,然后在这些轨迹上训练(即“蒸馏”)自己的模型。现有的针对蒸馏攻击的防御通常是在蒸馏之后立即进行评估,这隐含地假设攻击者不会对其模型进行任何进一步训练。在本文中,我们认为一个更现实的威胁模型包括在蒸馏之后使用强化学习进行进一步训练。一个被错误设定的威胁模型可能会带来虚假的安全感——一些在蒸馏之后看似有效的防御,在随后的强化学习之后可能会被攻破。在实践中,强化学习降低了蒸馏攻击取得成效的门槛。我们表明,简单攻击可以使用当前 API 中易于获取的数据,从现有闭源语言模型中窃取推理能力,其带来的推理提升等价于那些提取完整隐藏轨迹的更复杂攻击。结果表明,任何泄露了足够信息以重建近似推理轨迹的蒸馏防御都可能是无效的。最后,我们讨论了更广泛的启示,以及可能在效果上更好的批次级蒸馏防御。
cs.LG / 193 / 2609.35701
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
MeqMuon:用于 LLM 预训练的矩阵均衡 Muon
Chang-Wei Shi, Xu Wang, Wu-Jun Li
cs.LG
large language model
大语言模型相关
Abstract
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
Chinese Translation
大型语言模型(LLMs)的成功伴随着模型规模和预训练成本的持续增长。Muon 在 LLM 预训练中提供了高精度与高训练效率。近期的工作将行归一化引入 Muon,以平衡更新幅度并提升预训练性能。然而,仅靠行归一化无法适应更新矩阵中不同的不平衡模式。在本文中,我们提出了一种改进的 Muon 优化器,称为 \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon),用于 LLM 预训练。MeqMuon 通过归一化同时平衡行与列的幅度,该归一化能够自动适配不同的不平衡模式,而无需人工干预。此外,MeqMuon 无需存储 AdamW 的二阶矩估计,从而减少了优化器状态的内存占用。实证结果表明,在 LLM 预训练中,MeqMuon 相比 AdamW、Muon 及其他基线取得了更好的收敛性能。
cs.LG / 194 / 2609.35760
TokenCast: Forecasting Token Consumption During LLM Agent Execution
TokenCast:预测 LLM 智能体执行期间的 Token 消耗
Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
cs.LG · cs.AI · cs.SE
large language model
大语言模型相关
Abstract
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.
Chinese Translation
当大型语言模型(LLM)智能体执行同一任务时,不同运行之间的 token 消耗可能相差超过一个数量级。智能体根据工具反馈和中间结果选择其后续步骤,而不断增长的上下文会持续膨胀每一次后续调用的输入规模。因此,一项任务的总消耗在执行前难以预测,并且随着运行的展开,预测必须被修正。在本文中,我们提出 TokenCast,它为每个执行片段学习一种可组合的成本表示,记录其自身的消耗以及它引入的上下文增长。组合相邻片段会产生一个累积估计,该估计捕捉了当来自较早片段的上下文被之后每一次调用重新读取时所产生的额外输入成本。随着执行展开,新观察到的证据会刷新预测,不需要额外的 LLM 调用,并且在 SWE-bench Verified 上每次运行带来 32.8 ms 的平均累积预测时间。在 4 个任务套件和 6 个智能体模型上,在 96 个评估组合中,TokenCast 相对于最强比较方法的平均绝对误差降低平均为 14.5%。在离线预算控制回放中,在轨迹完成度匹配的情况下,TokenCast 平均比固定预算策略少使用 21.3% 的 token。代码可在 https://github.com/DEFENSE-SEU/TokenCast 获取。
cs.MA / 195 / 2609.33871
Population Physics, Population Problems: Safety and Emergence in LLM Societies
群体物理,群体问题:LLM 社会中的安全与涌现
Adrian de Wynter
cs.MA · cs.CL
large language model
大语言模型相关
Abstract
The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic systems, however, understanding these systems is paramount. For that we introduce a framework for measuring self-organisation in LLM social systems and apply it to three such systems: a Schelling grid, a social network (Moltbook), and a Twitter-like misinformation simulation ('Rogue'). All three exhibit statistically significant self-organisation. Moreover, their relaxation dynamics vary with the environmental information available to the agents, with open-ended systems (Moltbook, Rogue) exhibiting sharp, phase-transition-like dynamics. Further results show that population-level pathologies can emerge even when the LLMs are safety-tuned or monitored, being primarily driven by the coordinated activity of a population subset. We also show when self-organisation does \textit{not} emerge under two additional scenarios (a commons dilemma, GovSim, and a LLM-as-a-judge deliberation scheme, ChatEval). We argue that measuring signatures of this kind offers a lightweight, agent-agnostic diagnostic layer for detecting coordinated collective behaviour in deployed multi-agent systems without relying on natural language or model versioning.
Chinese Translation
大型语言模型(LLM)社会的集体行为并不是其个体输出的总和。它会产生统计上不同、有时不可预测的现象,而我们用来研究单个智能体的工具可能无法扩展到这些现象。然而,由于近期涉及自主智能体系统的事件,理解这些系统至关重要。为此,我们引入一个用于测量 LLM 社会系统中自组织的框架,并将其应用于三个此类系统:一个谢林网格、一个社交网络(Moltbook),以及一个类似 Twitter 的虚假信息模拟(“Rogue”)。三者都表现出统计显著的自组织。此外,它们的弛豫动力学随智能体可获得的环境信息而变化,其中开放式系统(Moltbook、Rogue)表现出尖锐的、类相变的动力学。进一步的结果表明,即使 LLM 经过安全微调或受到监控,群体层面的病态现象也可能出现,并且主要由群体中一部分的协同活动驱动。我们还展示了在另外两个场景(公地困境 GovSim,以及 LLM 作为评判者的审议方案 ChatEval)中,自组织并不出现的情况。我们认为,测量这类特征提供了一种轻量级、与智能体无关的诊断层,用于在已部署的多智能体系统中检测协同的集体行为,而无需依赖自然语言或模型版本管理。
cs.MA / 196 / 2609.33885
Prospective Interpretation Risk: Principled Communication Control Between LLMs
前瞻性解释风险:LLM 之间有原则的通信控制
Wanrong Yang, Rehan Deen, Julian Ma, Yuheng Fan, Yaoyu Jin, Taher Jafferjee, Ziquan Liu, Dominik Wojtczak, Yalin Zheng, David Henry Mguni
cs.MA · cs.LG
large language model
大语言模型相关
Abstract
Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers can reconstruct different tasks from the same message. We model this as a sender-receiver problem with a latent receiver type and define prospective interpretation risk (PIR): the probability that a receiver reconstructs a task other than intended. Rather than model an LLM's full input-output behaviour, we use black-box probes relating messages, intended tasks, and receiver-specific reconstructions, yielding scalable supervision while separating interpretation from downstream capability failure. Offline, heterogeneous frozen receivers provide supervision for receiver-conditioned risk and the effects of predefined mutable message features. At deployment, history induces a posterior over receiver types, guiding message revision and selection. We introduce value of interpretation information (VoII), querying for receiver information only when its expected communication benefit exceeds its cost. Our theory characterises when receiver information has decision value and bounds such queries. Empirically, interpretation-failure rates vary by 4-13x across receivers. Receiver information reduces PIR calibration error by 68% relative to a receiver-agnostic predictor, largely by correcting receiver-specific risk levels. PIR-guided revision reduces interpretation failure by 44% relative to the original message and 40% relative to a generic rewrite, mostly through a repair that helps every receiver. VoII outperforms information-gain and random querying at matched cost on the interpretation objective it optimises, lowering interpretation failure from 3.84% to 3.79% while querying 18.2% of episodes.
Chinese Translation
大型语言模型(LLM)智能体系统日益依赖模型彼此通信,然而现有的不确定性和多智能体方法很少在消息发送前估计特定接收者将如何解释该消息。这在异构系统中很重要,其中能力较强的接收者可以从同一条消息中重构出不同的任务。我们将其建模为一个具有潜在接收者类型的发送者-接收者问题,并定义前瞻性解释风险(PIR):接收者重构出的任务不同于预期任务的概率。我们没有建模 LLM 的完整输入-输出行为,而是使用将消息、预期任务和接收者特定重构联系起来的黑盒探针,从而产生可扩展的监督,同时将解释与下游能力失败区分开来。离线时,异构的冻结接收者为以接收者为条件的风险以及预定义可变消息特征的影响提供监督。在部署时,历史会诱导出关于接收者类型的后验,指导消息修订和选择。我们引入解释信息价值(VoII),仅当其预期通信收益超过成本时,才查询接收者信息。我们的理论刻画了接收者信息何时具有决策价值,并对这类查询给出界。经验上,解释失败率在不同接收者之间变化达 4-13 倍。相对于与接收者无关的预测器,接收者信息将 PIR 校准误差降低了 68%,这主要是通过校正接收者特定的风险水平实现的。PIR 指导的修订相对于原始消息将解释失败降低了 44%,相对于通用重写降低了 40%,这主要是通过一种对每个接收者都有帮助的修复实现的。在其优化的解释目标上,VoII 在同等成本下优于信息增益查询和随机查询,将解释失败率从 3.84% 降至 3.79%,同时仅查询 18.2% 的回合。
cs.OS / 197 / 2609.34727
Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs
动态流,静态图:面向移动 NPU 高效 LLM 服务的 KV 缓存复用
Zhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun, Zhaode Wang, Zeyu Zhao, Chengfei Lv, Fan Wu, Guihai Chen
cs.OS · cs.AI · cs.DB
large language model
大语言模型相关
Abstract
On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily optimized for cloud GPUs with dynamic execution environments and abundant memory bandwidth. These architectural assumptions do not hold on mobile NPUs, where computation graphs must be statically compiled and both memory capacity and I/O bandwidth are severely constrained. In this work, we present a compute-storage co-design for mobile-centric prefix and non-prefix KV reuse. We first propose an intra-graph mechanism that maps selective KV recomputation onto static NPU graphs, reconciling algorithmic dynamicity with NPU staticity. We further develop an inter-graph scheduler to optimize chunk merging and minimize padding with dynamic programming. To address mobile bandwidth limitations, we introduce a hierarchical KV manager featuring a tree-hash-semantic hybrid structure, along with cost-aware prefetching and eviction policies. We also build a two-dimensional pipeline that overlaps KV loading, rerotation, and storage with NPU execution, hiding data-movement latency. Experiments across representative on-device workloads and LLMs show that our design reduces time-to-first-token (TTFT) by $40-60\%$ compared with no reuse and prefix-only caching.
Chinese Translation
设备端大语言模型(LLM)服务是本地优先个人智能的基石,为用户提供数据主权、强隐私保障,并摆脱云 API 的延迟和成本。尽管 KV 缓存被广泛用于降低长上下文推理中的延迟,但现有设计主要针对具有动态执行环境和充足内存带宽的云 GPU 进行了优化。这些架构假设在移动 NPU 上并不成立,因为在移动 NPU 上,计算图必须被静态编译,并且内存容量和 I/O 带宽都受到严重限制。在这项工作中,我们提出了一种面向移动端的计算-存储协同设计,用于以移动端为中心的前缀和非前缀 KV 复用。我们首先提出一种图内机制,将选择性 KV 重计算映射到静态 NPU 图上,从而协调算法动态性与 NPU 静态性。我们进一步开发了一种图间调度器,利用动态规划来优化块合并并最小化填充。为了解决移动带宽限制,我们引入了一种分层 KV 管理器,其具有树-哈希-语义混合结构,以及成本感知的预取和驱逐策略。我们还构建了一个二维流水线,将 KV 加载、重旋转和存储与 NPU 执行重叠,从而隐藏数据移动延迟。在代表性设备端工作负载和 LLM 上的实验表明,与不复用和仅前缀缓存相比,我们的设计将首 token 时间(TTFT)降低了 $40-60\%$。
cs.AI / 198 / 2609.34268
SAGE: Symbolic Action-Gating and Editing for LLM Task Planners
SAGE:面向 LLM 任务规划器的符号动作门控与编辑
Trung Minh Bui, JongSul Moon, YoungOuk Kim, Quang-Ngoc Phung, Se-Woong Jun, Dongin Shin
cs.RO · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be separated from another. We present SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate (~250 lines of Python, zero tokens, $O(|π|)$) that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact; a hybrid seed+live memory store supports cold-start coverage. We evaluate under a leak-free protocol (leave-one-out retrieval) over five open-weight models and a 75-task AI2-THOR benchmark. On the standard benchmark goal-completeness saturates (52% of instances trivially solved) and SAGE ties strong hierarchical baselines. On a harder, method-agnostic multi-goal composition, SAGE's completeness lead re-emerges large (+0.06 to +0.23 across four models). Under injected mid-execution failures, SAGE recovers as reliably as whole-plan replanners at 2.4-3.3x fewer LLM calls. As a verify-before-execute gate, the symbolic monitor blocks unsafe actions before actuation and raises simulator-reported step-success for every planner tested (up to +0.11), a signal the verifier never sees (non-circular). Because the gate calls no model (0.008 ms/plan), it is a safety layer that runs essentially free on the edge: SAGE planning reproduces its quality on a Jetson AGX Orin, where small-model verification helps most. We release the benchmark, the leak-free protocol, the recovery and safety-gate harnesses, and a verifier-portability study (auto-induced on ALFWorld, 0.89 held-out).
Chinese Translation
大型语言模型(LLM)现已成为具身家庭智能体的默认认知核心,然而它们输出的计划在执行前很少会对照一个有环境依据的模型进行检查,并且它们所报告的任务成功率往往是在如此饱和的基准上测量的,以至于没有任何方法能够与另一种方法区分开来。我们提出 SAGE(符号动作门控与编辑),一个由两个轻量级机制构建的单一 LLM 规划器:一个领域无关的符号门控(约 250 行 Python、零 token、$O(|π|)$),它作为运行时安全监视器,以带类型的原因阻止违反前置条件的动作;以及一种局部编辑,它仅重新生成失败子目标的后缀,同时保持已完成和未触及的工作不变;一个混合的种子+实时记忆存储支持冷启动覆盖。我们在一个无泄漏协议(留一法检索)下,对五个开放权重模型和一个包含 75 个任务的 AI2-THOR 基准进行了评估。在标准基准上,目标完整度趋于饱和(52% 的实例被轻易解决),SAGE 与强层次化基线持平。在更困难、与具体方法无关的多目标组合上,SAGE 的完整性领先优势重新大幅显现(在四个模型上为 +0.06 到 +0.23)。在注入的执行中途失败下,SAGE 在 LLM 调用次数少 2.4-3.3 倍的情况下,恢复得与完整计划重规划器一样可靠。作为一个先验证后执行的门控,该符号监视器在执行动作之前阻止不安全动作,并提升所测试的每个规划器的模拟器报告的单步成功率(最高 +0.11),这是验证器从未看到的一种信号(非循环)。由于该门控不调用任何模型(0.008 ms/计划),它是一个基本上可在边缘端几乎免费运行的安全层:SAGE 规划在 Jetson AGX Orin 上复现了其质量,而在该平台上小模型验证最有帮助。我们发布了该基准、无泄漏协议、恢复与安全门控测试框架,以及一项验证器可移植性研究(在 ALFWorld 上自动诱导,留出集为 0.89)。
cs.AI / 199 / 2609.34270
Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning
交互式机器人规划中意图消歧的贝叶斯主动学习
Huao Li, Carson Sobolewski, Augustinos Saravanos, William Tan, John Karigiannis, Chuchu Fan
cs.RO · cs.AI · cs.HC
large language model
大语言模型相关
Abstract
Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.
Chinese Translation
交互式机器人规划要求机器人从往往模糊、不完整或表述不充分的自然语言指令中推断并执行人类意图。尽管大语言模型(LLM)为澄清提供了强大的接口,但依赖生成模型来驱动多轮对话可能会引入系统性失败。我们提出一个贝叶斯框架,将澄清视为在基于具体环境的信号时序逻辑(STL)任务规范上的主动学习问题。我们的方法使用LLM初始化候选形式化规范,并将具有信息量的对比转化为自然语言澄清问题,同时贝叶斯优化维持对用户意图的不确定性估计,并选择能够最大化信息增益的查询。在收敛之后,推断出的STL规范被传递给形式化规划器,以合成可验证的机器人轨迹。在四个仿真与真实世界任务领域中,我们的方法总体上都比LLM基线取得了更高的任务满意度,并且需要更少的澄清轮次,同时帮助较小的模型缩小与较大推理模型之间的性能差距。
cs.AI / 200 / 2609.34608
Efficient World Action Model Inference with Adaptive Intermediate States
基于自适应中间状态的高效世界动作模型推理
Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma, Tao Ma, Zheng Liu, Yifeng Chen, Yunquan Zhang, Ting Cao, Yunxin Liu, Kun Li
cs.RO · cs.AI
diffusion
扩散模型相关
Abstract
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ achieves 1.47-3.05$\times$ speedups in observation-to-action latency and 2.23-3.27$\times$ speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Chinese Translation
世界动作模型(WAMs)通过联合建模动作与环境动力学,实现具有未来感知能力的控制。然而,迭代式的扩散或流推理会带来显著的去噪延迟。先前的推理状态为加速提供了天然的契机,但不断变化的规划上下文、观测以及中间表示会迅速使保留的状态变得过时。因此,要保留有用的计算,就需要对推理状态进行自适应调整,而不是将其原样复用。为此,我们提出了 $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$,这是一个免训练框架,它通过保留并自适应调整推理状态,在控制回路不断演进的过程中实现高效且准确的延续,从而加速 WAM 推理。在闭环重规划之间,轨迹重映射(Trajectory Remapping)将来自前一次重规划的重规划状态进行重映射,以初始化下一次重规划,从而减少冗余的轨迹生成。在去噪步骤之间,观测重绑定(Observation Rebinding)在执行动作期间进行预期性推理,并在一致性检查通过时将保留的去噪状态重新绑定到真实观测以继续推理,从而减少暴露给控制回路的延迟。在 Transformer 层之间,残差重缩放(Residual Rescaling)选择性地对保留的层状态进行重缩放,并在探针检查失败时通过中间层的完整计算对其进行刷新,从而减少重复的 Transformer 计算。在 LIBERO 和 RoboTwin 2.0 上对三种代表性 WAM 架构的评估表明,$\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ 在观测到动作的延迟上实现了 1.47-3.05$\times$ 的加速,在每次重规划的 GPU 推理时间上实现了 2.23-3.27$\times$ 的加速,同时保持了原生 WAM 任务成功率的 96.69-99.54%。
cs.AI / 201 / 2609.35469
Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching
重新思考流匹配中采用条件退火的因果动作标记化
Chenyu Zhang, Yuhang Cao, Daru Du, Yingxi Lu, Jing Shao, Ruoqu Chen, Jiajun Liu, Liu Cao, Yicheng Liu, Hang Zhao, Mengdi Xu
cs.RO · cs.AI · cs.LG
diffusion
扩散模型相关
Abstract
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
Chinese Translation
自回归视觉-语言-动作(VLA)模型为机器人学习提供了一条可扩展的路径,然而现有的动作标记器将标记化视为一个压缩问题,所产生的表示在语义上与自回归骨干网络不对齐。我们提出了 CATok,一种因果动作标记器,它将标记化重新表述为一个因果结构化的生成过程。CATok 引入了一种条件退火机制,该机制通过逐步退火一个流匹配过程来提取动作标记:每个标记以所有先前标记为条件,并编码特定噪声水平下的残差重建信号,从而建立一个由粗到细的因果标记空间,其生成语义在结构上与自回归建模对齐。一个基于多模态扩散 Transformer(MMDiT)构建的、以标记为条件的流匹配解码器,以混合扩散头架构的精度从这些离散标记中重建连续动作片段。这种离散瓶颈在设计上强制实现知识隔离,干净地将高层语义推理与低层运动执行分离,而无需显式注意力掩码。在三个仿真基准和真实世界机器人操作任务上的广泛评估表明,CATok 在重建保真度-压缩权衡和推理效率方面均持续超越现有的标记化方法,同时提高了 VLA 任务成功率和训练效率,为纯自回归 VLA 系统建立了一个高性能、可扩展的基础。
cs.MA / 202 / 2609.35651
Denoising Multi-Robot Trajectories
多机器人轨迹去噪
Yuhao Zhang, Keisuke Okumura, Ajay Shankar, Amanda Prorok
cs.RO · cs.MA
diffusion
扩散模型相关
Abstract
Multi-robot trajectory planning is a fundamental problem in multi-robot coordination but remains computationally challenging due to its nonconvex, multimodal, and high-dimensional nature. This work builds upon D4orm, a dynamics-aware diffusion-denoising framework, and develops a family of planning architectures for diverse operational requirements. Unlike conventional numerical optimization methods, D4orm employs sampling-based optimization to generate solution trajectories through massively parallel sampling, leveraging modern computing architectures such as GPUs. Its diffusion-denoising structure iteratively optimizes \textit{deformations} to candidate control trajectories, providing an efficient and versatile paradigm for generating kinodynamically feasible and conflict-free trajectories. Using D4orm as the building block for advanced planners, we present a decoupled planner for improved scalability, an online receding-horizon planner with feedback control, and a distributed planner for resource-constrained settings. Evaluations with differential-drive and holonomic robots in 2D and 3D environments demonstrate that D4orm-based approaches find high-quality solutions faster and more reliably than other sampling-based optimization methods, such as MPPI, as well as a learned diffusion-model-based method. We further demonstrate zero-shot deployment on ten real quadrotors with obstacles, large-scale deconfliction with 100 simulated robots, and fully onboard distributed `lifelong' operation with six ground robots. Overall, these results establish diffusion denoising as a scalable and reliable framework for multi-robot coordination. Code and video: https://github.com/proroklab/d4orm
Chinese Translation
多机器人轨迹规划是多机器人协调中的一个基本问题,但由于其非凸、多模态和高维特性,在计算上仍然具有挑战性。这项工作建立在 D4orm 这一动力学感知的扩散-去噪框架之上,并针对多样化的操作需求开发了一系列规划架构。与传统数值优化方法不同,D4orm 采用基于采样的优化,通过大规模并行采样生成解轨迹,并利用 GPU 等现代计算架构。其扩散-去噪结构迭代地优化对候选控制轨迹的 \textit{形变},为生成满足运动动力学可行性且无冲突的轨迹提供了一种高效且通用的范式。以 D4orm 作为高级规划器的构建模块,我们提出了一种用于提高可扩展性的解耦规划器、一种带有反馈控制的在线滚动时域规划器,以及一种用于资源受限环境的分布式规划器。在 2D 和 3D 环境中对差速驱动和全向机器人进行的评估表明,基于 D4orm 的方法比其他基于采样的优化方法(如 MPPI)以及一种基于学习到的扩散模型的方法更快、更可靠地找到高质量解。我们进一步展示了在十架带障碍的真实四旋翼上的零样本部署、在 100 个仿真机器人上的大规模冲突消解,以及在六台地面机器人上完全机载的分布式 `终身` 运行。总体而言,这些结果确立了扩散去噪作为一种可扩展且可靠的多机器人协调框架。代码和视频:https://github.com/proroklab/d4orm
cs.AI / 203 / 2609.34030
Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
通过发现时序解释中反复出现的概念来揭示音频分类器中的捷径学习
Cecilia Bolaños, Luciana Ferrer, Magdalena Fuentes
cs.SD · cs.AI
large language model
大语言模型相关
Abstract
Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that explain classifier decisions, caption them with an ensemble of Large Audio-Language Models, and use a Large Language Model to extract recurring concepts. The resulting concepts can be audited by humans to uncover potential shortcut learning. We evaluate our framework using datasets curated from AudioSet Strong, controlling for the presence or absence of spurious correlations. Results show that this approach reliably uncovers learned shortcuts, such as the model relying on the presence of "laughter" to predict "applause".
Chinese Translation
机器学习数据集中事件之间的相关性可能导致捷径学习,其中模型学习基于相关事件的存在来预测目标事件。当这些相关性是虚假的——源于数据收集过程中的人为因素——时,模型很可能在实践中表现不佳。我们提出了一种流水线,通过发现音频分类器的时序解释中反复出现的概念,来揭示音频分类器中的捷径学习。具体而言,我们分离出解释分类器决策的音频片段,使用大型音频-语言模型集成对它们进行描述,并使用大型语言模型提取反复出现的概念。由此得到的这些概念可由人类审计,以揭示潜在的捷径学习。我们使用从 AudioSet Strong 中整理而来的数据集来评估我们的框架,控制虚假相关性的存在与否。结果表明,这种方法能够可靠地揭示已学习到的捷径,例如模型依赖“笑声”的存在来预测“掌声”。
cs.AI / 204 / 2609.34347
SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former
SAIL:通过解耦声学-空间编码与双流 Q-Former 的大语言模型空间音频智能
Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou, Woon-Seng Gan, Wenwu Wang
cs.SD · cs.AI · eess.AS
large language model
大语言模型相关
Abstract
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Chinese Translation
空间音频大语言模型(LLMs)使具身智能体、可穿戴助手和沉浸式系统能够识别声音事件、定位声源并推理它们的空间关系。然而,现有的空间音频 LLM 往往依赖于声学特征和空间特征的早期融合以及源无关的 token 表示。这些设计使得难以保持单个声音事件与其空间属性之间的对应关系,尤其是在多声源场景中。为了解决这一局限,我们提出了 SAIL,一个基于 LLM 的空间音频智能框架,它从音频编码到 LLM 对齐过程中保持声学-空间结构和声源级对应关系。SAIL 引入了一种解耦空间音频 Transformer,它将梅尔频谱图和双耳相位差特征表示为独立的声学流和空间流。源判别性任务查询进一步学习每个声源的事件、方向和距离信息。然后,一个双流 Q-Former 使用按声源槽位组织的声学查询和空间查询,将这两个流与 LLM 对齐。与早期融合基线相比,SAIL 在双声源声音事件检测、方向和距离估计以及空间推理方面取得了持续改进。这些结果证明了结构化、声源判别性音频表示对于多声源空间理解和推理的重要性。
cs.SE / 205 / 2609.33918
Green AI: Cost of LLM-Based Code Completion
绿色 AI:基于 LLM 的代码补全成本
Negar Alizadeh, Nishant Saurabh, Fernando Castor
cs.SE
large language model
大语言模型相关
Abstract
Code completion is one of the most widely used applications of large language models (LLMs) in software development. Open-weight LLMs are increasingly adopted for locally deployed code completion systems, partly due to privacy concerns. Despite advances in LLM accuracy, the energy cost of inference remains underexplored, particularly under large-context workloads and across programming languages. This study investigates the trade-off between accuracy and energy consumption in LLM-based code completion and how workload characteristics, context size, and model scale influence inference energy usage. We evaluate 25 open-weight LLMs on two workloads: repository-level next-line completion with varying context sizes on RepoBench, and fill-in-the-middle (FIM) code completion across Python, Java, and Rust on McEval. We analyze the influence of input tokens, output tokens, model size, and their interactions on energy consumption using correlation analysis and cluster-robust linear regression. Our findings show that the dominant drivers of energy consumption depend strongly on task structure. In RepoBench, energy consumption is primarily influenced by input context size and its interaction with model scale, whereas in McEval, output generation and its interaction with active parameter count dominate. Output generation is more energy-intensive per token than prompt processing. Across both benchmarks, smaller and heavily quantized models frequently achieve Pareto-optimal trade-offs, often providing accuracy comparable to larger FP16 models while consuming substantially less energy. Increasing model size or context length does not necessarily lead to proportionally better completion quality, while quantization can substantially improve energy efficiency with limited accuracy degradation. These findings support more energy-aware deployment strategies for sustainable AI-assisted software development.
Chinese Translation
代码补全是大型语言模型(LLM)在软件开发中最广泛使用的应用之一。开放权重 LLM 正越来越多地被用于本地部署的代码补全系统,部分原因是隐私方面的顾虑。尽管 LLM 准确性取得了进展,但推理的能耗成本仍未得到充分探索,特别是在大上下文工作负载下以及跨编程语言的情况下。本研究考察了基于 LLM 的代码补全中准确性与能耗之间的权衡,以及工作负载特征、上下文大小和模型规模如何影响推理能耗。我们在两个工作负载上评估了 25 个开放权重 LLM:在 RepoBench 上进行具有不同上下文大小的仓库级下一行补全,以及在 McEval 上进行跨越 Python、Java 和 Rust 的中间填充(FIM)代码补全。我们使用相关性分析和聚类稳健线性回归,分析输入 token、输出 token、模型规模及其交互作用对能耗的影响。我们的研究发现表明,能耗的主要驱动因素在很大程度上取决于任务结构。在 RepoBench 中,能耗主要受输入上下文大小及其与模型规模的交互作用影响,而在 McEval 中,输出生成及其与活跃参数数量的交互作用占主导地位。输出生成在每 token 上比提示处理更耗能。在两个基准测试中,较小且高度量化的模型经常实现帕累托最优权衡,通常提供与更大的 FP16 模型相当的准确性,同时消耗显著更少的能量。增加模型规模或上下文长度并不一定会带来按比例更好的补全质量,而量化可以在准确性下降有限的情况下大幅提高能效。这些发现支持更注重能源的部署策略,以实现可持续的 AI 辅助软件开发。
cs.SE / 206 / 2609.34661
Codoku: Renewable Program-Reasoning Challenges for Frontier Coding Agents
Codoku:面向前沿编码智能体的可再生程序推理挑战
Cong Li, Hao Sun, Zenan Li, Zhendong Su
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input. Coding agents break two assumptions on which these benchmarks rest: an agent can recover the answer by executing the program instead of reasoning about it, and fixed task sets drawn from existing programs are increasingly exposed to contamination, yet costly to renew. We introduce Codoku (code sudoku), a renewable benchmark in which a solver fills typed cells in a partial program to satisfy global static and dynamic constraints, such as a prescribed control-flow graph and execution path. Because a partial program cannot be executed and valid fillings are sparse in an exponentially large space of interdependent choices, neither tool use nor enumeration can substitute for program reasoning. Puzzles are synthesized from scratch via semantic reification, so fresh puzzles of controllable complexity can be generated on demand, each with a witness that guarantees solvability. We evaluate five frontier models on 300 puzzles through a coding agent free to use any tool within a fixed budget. Small puzzles already challenge open-weight models, whereas even proprietary models solve only about half of the large ones. Codoku thus offers a renewable testbed for program reasoning that can keep pace with rapidly improving coding agents. GitHub: https://github.com/connglli/Codoku.
Chinese Translation
现有的程序推理基准要求大语言模型预测程序在给定输入下的行为。编码智能体打破了这些基准所依赖的两个假设:智能体可以通过执行程序而非对程序进行推理来得到答案,并且,从现有程序中抽取的固定任务集越来越容易受到污染,而更新它们的成本又很高。我们提出 Codoku(代码数独),这是一种可再生的基准,其中求解器填充部分程序中的带类型单元,以满足全局静态与动态约束,例如规定的控制流图和执行路径。由于部分程序无法执行,且有效填充在相互依赖选择的指数级大空间中十分稀疏,工具使用和枚举都无法替代程序推理。谜题通过语义具体化从零合成,因此可以按需生成复杂度可控的新谜题,每个谜题都带有一个保证可解性的见证。我们通过一个可在固定预算内自由使用任何工具的编码智能体,在 300 个谜题上评估了五个前沿模型。小型谜题已经对开放权重模型构成挑战,而即使是专有模型也只能解出大约一半的大型谜题。因此,Codoku 为程序推理提供了一个可再生的测试平台,能够跟上快速进步的编码智能体。GitHub:https://github.com/connglli/Codoku。
cs.SE / 207 / 2609.35281
Is there a future for models in the LLM era?
在LLM时代,模型还有未来吗?
Marco Calamo, Massimo Mecella, Monique Snoeck
cs.SE
large language model
大语言模型相关
Abstract
In an era where software development is deeply tied with Large Language Models, does Model Driven Engineering (MDE) still make sense? This raises the question of the extent to which MDE can be successfully combined with an LLM approach to address the downsides of each approach separately. In this paper, we try to answer that question by investigating a central Research Question: Can Agents powered by Large Language Models improve code generated from UML diagrams and text specifications by Model Driven Engineering (MDE) tools? To investigate this problem, we developed a novel LLM-powered Multi-Agentic approach called ARTHUR (Architecture Refactoring Through Hybrid UML Reasoning), a framework that aims at combining the reliability of MDE and the ease of use of LLMs. ARTHUR is designed to refactor legacy Java code produced by traditional, rule-based, MDE code generators from UML diagrams. To asses the answer to our question, we have refactored the legacy code of several projects from our custom dataset crafted for this purpose. \name{} which made it possible to add support for modern frameworks like Spring Boot, while ensuring compliance with Model-Based Testing techniques to verify that the code still corresponds to the initial model's specifications. We then measured the results obtained in terms of time and cost, passing test rate, and \texttt{compile@k}, \texttt{pass@k} and \texttt{pass$^k$} metrics. We also observed the effect of generating code directly from the conceptual model without the refactoring. Our preliminary test results show that MDE could not be more far from retirement, after all.
Chinese Translation
在软件开发与大型语言模型深度关联的时代,模型驱动工程(MDE)仍然有意义吗?这引出了一个问题:MDE能够在多大程度上成功地与LLM方法相结合,以解决每种方法各自的缺点。在本文中,我们试图通过研究一个核心研究问题来回答这个问题:由大型语言模型驱动的智能体能否改进由模型驱动工程(MDE)工具从UML图和文本规范生成的代码?为了研究这一问题,我们开发了一种新颖的、由LLM驱动的多智能体方法,称为ARTHUR(Architecture Refactoring Through Hybrid UML Reasoning,通过混合UML推理进行架构重构),这是一个旨在结合MDE的可靠性与LLM易用性的框架。ARTHUR旨在重构由传统的、基于规则的MDE代码生成器从UML图生成的遗留Java代码。为了评估我们问题的答案,我们重构了为此目的构建的自定义数据集中的多个项目的遗留代码。\name{} 使得可以添加对Spring Boot等现代框架的支持,同时确保符合基于模型的测试技术,以验证代码仍然符合初始模型的规范。随后,我们从时间和成本、测试通过率以及 \texttt{compile@k}、\texttt{pass@k} 和 \texttt{pass$^k$} 指标方面衡量了所得结果。我们还观察了不进行重构而直接从概念模型生成代码的效果。我们的初步测试结果表明,毕竟,MDE远未到退休之时。
cs.CL / 208 / 2609.34217
Explainable and Generalisable LLM-based Cognitive Decline Detection with Spontaneous Speech
基于自发语音的可解释且可泛化的LLM认知衰退检测
Ziyun Cui, Wen Wu, Chuan Shi, Shuguang Yang, Xueying Gui, Yan Zheng, Qiong Yang, Haiyan Zhao, Wei-Qiang Zhang, Ji Wu, Yelei Li, Nan Li, Chao Zhang
eess.AS · cs.CL
large language model
大语言模型相关
Abstract
Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system directly processes raw speech to learn joint acoustic-semantic representations, preserving critical prosodic cues often lost in transcription. Utilising our newly collected PUTH-AD dataset alongside multiple open-source corpora, we implemented a multi-task learning objective that simultaneously performs cognitive status classification and generates clinician-understandable natural language explanations. Our system achieved the highest average accuracy and AUROC across six dataset/task conditions, comparing three representative baselines. The system demonstrated cross-task transfer to held-out PUTH-AD task subsets, maintaining classification accuracy on an entirely unseen cognitive task without task-specific fine-tuning. Furthermore, clinician evaluation confirms that the generated explanations are both clinically relevant and largely consistent with the underlying speech evidence, supporting their potential utility in clinical interpretation. This study provides a scalable, objective, and explainable framework for speech-based cognitive screening, combining cognitive status classification with natural language explanations that clinicians can assess and verify, bridging the gap between advanced AI and clinical utility.
Chinese Translation
阿尔茨海默病(AD)和轻度认知障碍(MCI)(后者可能先于AD出现)通过细微的语言和声学改变在早期表现出来。然而,传统诊断往往资源密集,缺乏大规模筛查的可扩展性。为了应对这些挑战,我们提出了一种新颖的双语语音大语言模型框架,用于自动化、可解释的认知筛查。与依赖易出错自动语音识别的传统流程不同,我们的系统直接处理原始语音以学习联合的声学-语义表示,保留了转写中常常丢失的关键韵律线索。利用我们新收集的PUTH-AD数据集以及多个开源语料库,我们实现了一个多任务学习目标,该目标同时执行认知状态分类并生成临床医生可理解的自然语言解释。我们的系统在六个数据集/任务条件下,与三个代表性基线相比,取得了最高的平均准确率和AUROC。该系统展示了对留出的PUTH-AD任务子集的跨任务迁移,在完全没有见过的认知任务上保持了分类准确率,而无需针对特定任务进行微调。此外,临床医生评估证实,生成的解释既具有临床相关性,又在很大程度上与底层语音证据一致,支持其在临床解读中的潜在效用。本研究提供了一个可扩展、客观且可解释的基于语音的认知筛查框架,将认知状态分类与临床医生可以评估和验证的自然语言解释相结合,弥合了先进AI与临床实用性之间的差距。
cs.AI / 209 / 2609.35581
QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
QC-Stark:一个多任务基准,揭示在量子计算任务上评估的大语言模型中的能力分离
Pranav Gupta
quant-ph · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.
Chinese Translation
我们提出 QC-Stark,一个用于在 11 项量子计算(QC)任务上评估大语言模型(LLMs)的基准,这些任务涵盖电路构建、调试、编译、纠错和模拟。在 2,750 次评估中(10 个模型 $\times$ 11 项任务 x 5 个难度级别 x 5 个种子),我们发现总体排名掩盖了各任务之间相当大的差异。对于该基准中包含的 11 项任务中的 4 项,总体排名与各任务排名之间的 Spearman 相关系数在统计上不显著。一个 2 参数项目反应理论(IRT)模型验证了测量质量,提示敏感性分析确认了排名在不同提示条件下的稳健性。所有任务均可通过执行自动验证,因此不需要任何人工评估。我们在 Huggingface 上公开提供代码和数据。
cs.AI / 210 / 2609.34188
AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning
AlphaPareto:基于LLM引导的多目标强化学习的公式化Alpha发现
Yingbo Zhao, Zeyu Yang, Zhoufan Zhu
stat.ML · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.
Chinese Translation
公式化 alpha 发现是量化交易中的核心挑战,因为识别能够良好协同工作的 alpha 仍然很困难。近期的强化学习(RL)方法将该任务表述为马尔可夫决策过程(MDP),但仍有两个重要问题尚未解决。首先,随着 alpha 池的演化,奖励函数也随之变化,使得该 MDP 本质上是非平稳的。其次,大多数现有方法优化单一目标,通常是预测能力,而忽略了高质量 alpha 池的其他重要性质。受这些挑战的启发,我们提出 AlphaPareto,一种用于公式化 alpha 发现的 RL 方法。为了解决非平稳性,AlphaPareto 扩展状态,使其同时包含正在构建的 alpha 和当前 alpha 池,并应用大语言模型(LLM)对 alpha 池进行编码。这种设计使智能体能够适应不断演化的搜索环境。为克服单目标奖励设计的局限,AlphaPareto 用多目标向量值奖励替代标量奖励,该奖励同时刻画预测能力、时间稳定性、扰动鲁棒性和多样性,并通过 Pareto 正则化学习过程优化这些目标。在真实世界数据集上的实证应用表明,我们的 AlphaPareto 方法优于其竞争对手。
人工智能 (cs.AI)
217
cs.AI / 1 / 2609.33816
Dual-Vocabulary Language Model for Cross-Tokenizer Distillation
Kedi Chen, Chen Lin, Yutao Sun, Wei Zhang
cs.AI
Abstract
On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
cs.AI / 2 / 2609.33822
Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory
Jayant Parashar, Eugene F. Douglass, William C. Bastian, Suchendra M. Bhandarkar
cs.AI
Abstract
An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model. Its organizing overhypothesis is that tasks of a shared kind may exhibit recurring failures whose remedies transfer within that kind. Vestrum expresses failures as recognizable classes, proposes changes across verification, retrieval, decomposition, and knowledge synthesis, and screens their scope before evaluating them as a bundle. A persistent lessons file informs subsequent proposals. Across five settings and two baseline harnesses, the frozen harnesses improve held-out performance: UltraHorizon rises from 47.6 to 59.8 over GAM, Terminal-Bench 4 Hard from 63.7% to 70.3% of checks passed over Claude Code on eight held-out tasks at 1.03x test cost, and cell-type annotation agreement from 67.5% to 77.8% on held-out sections of one slide, alongside gains on LoCoMo and AMA-Bench. Across our searches, verification grounded in evidence helped both intermediate steps and final answers, at lower cost at intermediate steps, while critics asked to rebuild finished answers broke more than they repaired. On the three memory benchmarks, Vestrum also scores above the evaluated GEPA configurations in every paired evaluation.
cs.AI / 3 / 2609.33843
Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation
Gowthamkumar Nandakishore
cs.AI · cs.CL · cs.LG
Abstract
The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, $0.214$. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy ($0.767$ vs. $0.766$). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at $q=0.05$ (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.
cs.AI / 4 / 2609.33845
How code helps different tasks? A decompositional lens on LLM post-training
Zheng Yu, Yiwei Li, Yishen Chen, Xiang Li, Jiale Han, Benyou Wang, Jingbang Chen
cs.AI
Abstract
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decompositional lens for studying these effects in LLM post-training. We first decompose an execution-verified code corpus into interpretable categories based on the computational patterns of its solutions. Through controlled fine-tuning experiments, we compare individual categories with a balanced mixture across instruction-tuned models on question answering, mathematics, and code generation. The resulting response maps reveal recurring gains in average question-answering performance, while the same category can improve one model or task and degrade another. The best-performing category also varies with the starting model and target task. We then compose compact mixtures guided by these results and examine whether benefits observed in individual categories persist under joint training. On selected model--task pairs, mixtures whose constituents each improve the target task outperform both their best constituent and full-corpus training while using roughly 10--15\% of the full corpus. These exploratory findings illustrate a \emph{less is more} pattern and highlight how the value of code data in post training depends on which categories are combined for which model and task.
cs.AI / 5 / 2609.33867
R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution
Mingda Zhang, Qiang Huang, Yanjin Li, Zijia Wang, Qika Lin, Xiaoying Tang, Tiesunlong Shen
cs.AI
Abstract
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.
cs.AI / 6 / 2609.33870
When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur, Ece Kamar, Asli Celikyilmaz
cs.AI
Abstract
LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.
cs.AI / 7 / 2609.33910
When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
Zhihao Zhang, Chao Wang, Rujia Li, Qingze Wang, Xiaoyan Sun, Jun Dai
cs.AI
Abstract
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can outlive the context that originally justified the approval, creating residual authority reusable without renewed consent. We expose this failure mode through a longitudinal attack that starts from a target security-sensitive action, identifies the authority required to execute it, induces benign interactions that legitimately obtain that authority, and later replays the residual authority during adversarial execution. Across controlled and live settings, we demonstrate that residual-authority replay arises in practice and substantially increases the success of prompt-injection and context-rebinding attacks. We evaluate 508 AgentDojo attack cases across six LLM families using production-derived authorization semantics. With residual authority, attack success rate (ASR) increases by up to 35.1 percentage points compared with a fresh authorization state. In live context-rebinding attacks on 55 Terminal-Bench cases across three real-world production coding agents, residual-authority replay increases ASR by 24.9 percentage points on average. These findings expose a fundamental mismatch between persistent authorization and the contextual nature of user consent in long-lived LLM agents.
cs.AI / 8 / 2609.33955
Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents
Kaiwen Luo, Ming Gao
cs.AI
Abstract
Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an integrated methodology spanning evaluation specification, modular LLM judges, intent-preserving user simulation, and human-in-the-loop governance. The specification defines conversation-level end states and actionable failure ownership. Atomic judges share versioned evidence and feed an explicit aggregation graph. The simulator is released only after task-preservation and stability checks. Independent human audits estimate measurement fidelity, renew tiered reference sets, and route disagreements to label correction, guideline revision, or judge improvement. Production studies show that system-level fidelity improved across repeated audits, that human reviewers and automated judges improved together under the shared feedback loop, and that their combined workflow had the strongest descriptive performance in both reported task-completion settings. Because the studies are observational and the human reference itself required revision, these findings demonstrate operational usefulness rather than causal or universal superiority. The contribution is a practical framework for making multi-turn agent measurement reliable, actionable, and maintainable as the evaluated system and its evidence evolve.
cs.AI / 9 / 2609.34007
EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
Andre R Goncalves, Vincent Liu, Priyadip Ray
cs.AI
Abstract
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients' records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.
cs.AI / 10 / 2609.34015
A Computer Vision Approach to Visual Fraud Detection in Phishing Websites Using YOLOv8
Basil Sajid Shaikh, Hajar Homayouni
cs.AI · cs.CV
Abstract
Phishing remains one of the most common vectors for financial and identity fraud, and most detection systems still rely on inspecting a page's URL, HTML markup, or domain registration history. These signals are easy for an attacker to rotate or obfuscate, and they say very little about what actually convinces a victim to hand over a password or a card number: the way the page looks. This paper describes a visual, image-based approach to phishing detection that treats a rendered webpage the same way a human eye would, as a picture that either matches a trusted brand or doesn't. A YOLOv8 convolutional neural network was trained to classify full-page website screenshots as phishing or legitimate based on layout, logo placement, color scheme, and login-form structure, rather than on text extracted from the page. The system reached 92% classification accuracy on a held-out test set, processed a single screenshot in roughly 100 milliseconds, and, after a round of data augmentation aimed specifically at lighting, compression, and scaling variation, cut the false-positive rate by 11% relative to the pre-augmentation baseline. The paper walks through the dataset construction, the augmentation strategy, the model architecture and training setup, and the resulting performance, and closes with a discussion of where this kind of visual detector fits alongside, rather than instead of, existing URL- and content-based defenses.
cs.AI / 11 / 2609.34024
Jev in Medicine: A Benchmark Evaluation. Preliminary Results
Alfredo Madrid-García, Beatriz Merino-Barbancho
cs.AI · cs.LG
Abstract
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
cs.AI / 12 / 2609.34049
Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
Timothy DeLise, Seth Cromelin
cs.AI
Abstract
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry information from beyond the current window and transmit it to later states. This study presents a series of experiments using five open-weight models spanning Qwen, Llama, Mistral, and Muse Glimmer. We investigate whether information originating outside the immediate context window can persist through a rolling KV cache and remain useful for retrieval. Initial results show that retaining previously computed states improves retrieval across the models tested compared with recomputing the final fixed window from raw tokens. We then measure how far this effect extends and find that Muse Glimmer and Mistral 7B show the strongest \emph{latent information relay}: they can recover information even after the relevant source tokens have left the cache. Both models incorporate sliding-window attention in their published architectures, an association that motivates testing whether training with sliding windows promotes more reliable information retention.
cs.AI / 13 / 2609.34103
A Differentiable Optimization Framework for Registering Sequential Bounding Boxes with Point Cloud Stream
Xuesong Li, Jinguang Tong, Jie Hong
cs.AI
Abstract
Refining a sequence of coarse 3D bounding boxes against a LiDAR point-cloud stream demands tracks that are geometrically accurate (high IoU) and temporally coherent (low roughness), preferably without training data. The usual recipe keeps the two concerns apart: register each frame independently, then smooth the trajectory afterwards with a Kalman~RTS or Savitzky--Golay filter. Smoothing displaces boxes from a geometric optimum and never re-optimises, so it trades accuracy for smoothness. We instead fold the temporal smoothness constraint into a training-free registration objective and solve for all poses jointly with L-BFGS. The payoff depends on how well the object is seen. On well-observed tracks it is large: within the low-roughness budget, the joint objective beats both post-hoc smoothers on paired multi-seed statistics and cuts roughness several-fold relative to frame-wise registration at matched accuracy. Treating visibility as an experimental variable exposes the limit. The advantage decays monotonically as views become one-sided, until it is indistinguishable from zero for near-edge-on objects and slightly negative under a ray-cast simulator with range-dependent density and ego motion, where the decoupled pipeline is in fact ahead at tight roughness budgets. We locate that boundary and trace it to one term: orientation alignment ties yaw to the estimated velocity and fails once that estimate is noisy. A ground-truth-free rule can choose the temporal scale and keep every track inside the roughness budget.
cs.AI / 14 / 2609.34111
SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction
Sejin Sim, Jinsoo Bae, Seoung Bum Kim
cs.AI · cs.CV
Abstract
The rapid advancement of artificial intelligence has observed increased application in predicting vehicle interior noise levels within the automotive industry. However, the collection of labeled data for training models in this context involves significant costs. Previous studies in semi-supervised regression (SSR) have effectively mitigated the reliance on labeled data by incorporating unlabeled data. Nonetheless, these approaches often introduce a high computational cost due to the training of multiple models and data sampling. This study introduces SpecRegMatch, a novel SSR method aimed at addressing the computational cost associated with training by leveraging a single model, thus eliminating the need for multiple data samplings. SpecRegMatch integrates consistency regularization and information maximization to robustly train the model, achieved through various augmentations applied to both the embedding vectors and predicted values. Experimental results demonstrate that SpecRegMatch achieves state-of-the-art performance across various scenarios, even when using a single model. It attains a remarkable performance, as indicated by an R^2 score of 0.434. This is especially noteworthy in scenarios where labeled data is scarce. You can access the code for our proposed method at https://github.com/sejin-sim/SpecRegMatch.
cs.AI / 15 / 2609.34113
GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
Shaoqing Zhang, Kehai Chen, Xuefeng Bai, Zhuosheng Zhang, Pengfei Zhang, Yang Xiang, Min Zhang
cs.AI
Abstract
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR
cs.AI / 16 / 2609.34132
From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
Mingxi Zou, Langzhang Liang, Zhuo Wang, Yiyang Zhao, Lizhen Qu, Zenglin Xu
cs.AI
Abstract
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
cs.AI / 17 / 2609.34134
StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents
Wenle Liao, Zhao Wang, Jingchao Zhang, Jiajie Jin, Yimeng Xu, Zhicheng Dou
cs.AI
Abstract
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, making it difficult for agents to track which analytical artifacts remain valid over increasingly long horizons and changing dependencies. Consequently, stale artifacts may be silently inherited, propagating errors to downstream stages. To address this challenge, we propose StateGuard, an analytical-state validity management framework for long-horizon data agents. StateGuard externalizes evolving analytical progress into a state graph containing constraints, versioned variables, intermediate conclusions, and cross-state relations, treating each state as an executable, verifiable, and traceable object rather than textual memory alone. StateGuard maintains state validity through evidence-grounded verification and hierarchical intervention. To equip StateGuard with these capabilities, we first introduce Manager-Oriented Counterfactual Supervision, which constructs 3K state-centric trajectories through counterfactual runtime synthesis to fine-tune StateGuard for state maintenance, verification, and repair. We then apply Validity-Guided Policy Optimization, using runtime validity evidence to provide fine-grained learning signals for protocol correctness, state grounding, and intervention quality. Experiments on three diverse long-horizon data-analysis benchmarks show that StateGuard consistently improves data-agent performance while reducing dependency-induced downstream error propagation, demonstrating the advantages of explicit analytical-state management for reliable long-horizon data analysis.
cs.AI / 18 / 2609.34135
Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment
Renxiang Wang, Jiaming Cui
cs.AI
Abstract
A skill bank that helps one multi-agent system may leave another's behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4--32 agents and GPT and Qwen model ladders. Source evolution meets a joint quality, cost, model-tier, and confirmation goal in 14 of 16 settings. We then evaluate Evo2Team, which selects, adapts, and confirms source skills for the target team, alongside six frozen selectors across 28 transfer directions. Evo2Team's target-side exploration cost is below that of evolving a new target bank in every direction, even when reused reference evaluations are charged once. Twenty of 28 held-out outcomes meet the positive-transfer criterion, including three saved diagnostic tests. Selection alone does not explain these outcomes: KNN and CORAL choose different banks in two AgentsNet directions but produce identical recorded executions. When Evo2Team changes execution, gains can reach many tasks, as in a Count-Frequency direction that improves 28 of 32 tasks over KNN. Seven positive AgentsNet outcomes save 6.1--14.6\% in deployment cost while using transferred skills on only three to six of fifteen tasks. In five earlier accepted directions, all 22 task records using transferred skills pass three fixed-graph confirmations, but four fail in recorded executions on new graphs. Graphs and model responses change together in this comparison. These results show that skill transfer must be assessed through the actions agents take, the tasks those actions reach, and the quality and cost of the final deployment.
cs.AI / 19 / 2609.34136
Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms
Mingxi Zou, Wei Zhu, Zhuo Wang, Langzhang Liang, Zhiwen Tang, Yinghui Xu, Zenglin Xu
cs.AI
Abstract
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable local law that can be shared across interchangeable agents and adapt coordination as populations or interaction conditions change, without redefining a global organization. We introduce Waggle, a shared anonymous policy over bounded local views that jointly selects task actions, semantic communication, and local commitment updates. Repeated execution of the same law allows coordination to form, persist, and reorganize online without explicit roles or global topology. To learn this law across interchangeable agents and evolving coordination, we develop Swarm-Consistent Distillation (SCD), combining anonymous-orbit consistency with rollout-grounded prediction of the next local coordination field, with no added inference-time components. Across diverse coordination settings, the same learned law remains effective as populations and interaction budgets change, retains over 96% of substrate-specific oracle quality, and transfers without retraining; SCD further improves reorganization after counterevidence. Together, these results show that LLM-agent organization can emerge and adapt through repeated execution of a learned local law.
cs.AI / 20 / 2609.34139
Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?
Tien Tran, Namho Koh, Daiki E. Matsunaga, Ayush Jain, Kee Eung Kim
cs.AI
Abstract
Mobile GUI agents deployed in real settings must work across different applications that support the same functionality. Most existing benchmarks test each task in only one app, so a high score can mean the agent understands the task, or only that it knows that particular app. We introduce AnyAppBench, a category-controlled live Android benchmark that evaluates cross-application generalization while keeping the user goal fixed. It spans 10 functional categories, 100 task templates, and 520 task--application pairs over 52 applications. Agents run from raw instructions and with app-independent sub-goals, and a VLM judge labels every failed run under a fixed failure taxonomy whose reliability is measured by human annotation. We find that, across 13 agents, success on the original application does not transfer reliably to new applications with the same goal. Furthermore, providing high-level sub-goal decomposition produces only small, category-dependent changes that do not close the gap, and the mix of failure types changes with the target interface. Based on those insights, we believe the AnyAppBench benchmark provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Our code, data and the leaderboard can be found at the project website https://anyappbench.github.io/.
cs.AI / 21 / 2609.34151
Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization
Xuanqi Zhang, Ruinan Jin, Running Yang, Yuxuan Zhang, Minghui Chen, Wenlong Deng, Xiaoxiao Li
cs.AI
Abstract
Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution, where each recurring task type maintains a persistent tool space which is constructed from accumulated output experience. We identify three limitations of existing methods: (1) output-unaware selection: they rely primarily on tool descriptions or model priors rather than observed tool outputs; (2) statelessness across request: they select tools independently for each request without consolidating prior output experience into persistent task-specific state; (3) inference cost: they repeatedly search, rank, or reason over candidate tools for subsequent requests of the same task. We address these limitations through output-aware tool scoring, persistent task-specific tool spaces, amortized tool selection, and reusable configurations across models. We introduce LOTS (Likelihood-Only Tool Scoring), which evolves an agent's tool space from accumulated output experience while keeping model parameters fixed. After each request, LOTS holds the model's generated answer and estimates each tool's contribution by measuring how much the answer likelihood changes when its observed output is removed. These contributions are aggregated within each recurring task to rank tools and update its persistent space. Across three benchmarks, LOTS improves task performance while substantially reducing tool context. More importantly, sequential experiments demonstrate that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.
cs.AI / 22 / 2609.34157
TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora
Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang, Lihua Yu, Zujie Ren, Gang Chen, Junbo Zhao
cs.AI
Abstract
Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depends on surrounding schema and table context. Heterogeneous schemas, value formats, and serializations further weaken one-shot matching. We present TableSeek, a structure-preserving agentic search framework for heterogeneous table corpora. Instead of ranking tables once, an LLM agent iteratively follows sparse clues, inspects schema-preserving previews, identifies schema- and value-level mismatches, and refines its investigation. TableSeek uses cells and schemas as evidence anchors while retaining complete tables as evidence units, enabling fine-grained localization without losing the context required for interpretation and answerability checking. Without relying on retriever training or a precomputed semantic index, TableSeek produces transparent evidence-seeking trajectories and achieves competitive end-to-end performance against strong retrieval-and-reranking pipelines on heterogeneous table benchmarks. These results suggest that active, structure-preserving evidence seeking is a promising paradigm for open-domain table retrieval.
cs.AI / 23 / 2609.34160
RoutePrism: Tracing Construction Order Effects in Agent Memory
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu, Jiangqiang Li, Jun Zhang, Junkai Ji
cs.AI
Abstract
Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces which sources, compiled contexts, and answers differ. Because record content, timestamps, policy, and the answer model all stay fixed, any observed difference is localized to the memory construction step. A matched four-condition intervention tests whether a record displaced by reordering actually carried task-relevant evidence: restoring that single record recovers over 60 percentage points of lost accuracy, while substituting a non-supporting record of equal length does not. We evaluate the protocol on PersonaMem-32K (63 primary queries, 29 users) and 470 LongMemEval-S questions with histories spanning 38 to 62 sessions, replicating the core intervention across five answer models. Survivor selection, defined as the choice of which record a cluster retains, drives most source-level changes, while different memory policies (compaction, bounded recency, MemoChat-style summarization, A-MEM) produce distinct failure signatures at the source, context, and metadata layers.
cs.AI / 24 / 2609.34177
ReplayLens: Auditing Agents' Use of Outcomes
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu, Jiangqiang Li, Jun Zhang, Junkai Ji
cs.AI · cs.LG · cs.MA
Abstract
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.
cs.AI / 25 / 2609.34179
RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints
Richard Krueger, Lucas Krause, Zach Pocquette
cs.AI · cs.CL · cs.LG
Abstract
Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a constrained evidence decision rather than a leaderboard choice. RAGWarrant normalizes evaluator outputs and operational telemetry, applies predeclared quality and hard-risk gates, assigns evidence-class claim ceilings, preserves negative outcomes, and emits auditable PROMOTE, BLOCK, REJECT, or INCONCLUSIVE decisions. We evaluate the framework across T2-RAGBench, MultiHop-RAG, CRAG, HotpotQA, synthetic reproduction, and bounded local generative experiments. On HotpotQA, operational savings were blocked because answer quality fell beyond the declared margin. A bounded CRAG study selected a lower-cost quality-tied policy, but related generative gains were unstable and a held-out guardrail failed closed. We claim an auditable promotion-control abstraction, not optimizer superiority, human validation, or production readiness. The tagged artifact reproduces from a fresh clone, runs as a hardened Docker job, accepts external evaluator exports, and verifies artifact integrity.
cs.AI / 26 / 2609.34180
Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen
Xukui Qin, Youting Wang, Xinjie He, Ziyang Luo, Runxiong Wu, Yan-Syuan Chen, Zhongyao Chu
cs.AI · cs.CV
Abstract
How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study's strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.
cs.AI / 27 / 2609.34181
Efficient Reasoning via Constrained Optimization in Latent Space
Zhinan Hou, XingChen Li, Keyou You
cs.AI · math.OC
Abstract
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.
cs.AI / 28 / 2609.34184
CASS: Contribution-Aware Structured Sparsity for Model Merging
Yan Li, Guiping Cao, Meng Xu, Tao Jiang, Yaguang Song, Ming Tao, Yaowei Wang, Dongmei Jiang
cs.AI
Abstract
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters'' and overlooking their inherent modularity. In this paper, we propose \textbf{C}ontribution-\textbf{A}ware \textbf{S}tructured \textbf{S}parsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.
cs.AI / 29 / 2609.34195
PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
Shane K. A. Dalumura Hettige, Jonas Oppenlaender
cs.AI · cs.CL · cs.HC
Abstract
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas through tool calls and observes the result after every turn. The canvas includes a starting shape which cannot be erased, and the agent's goal is to incorporate this shape into the most original drawing it can produce. The task is open-ended, and the agent itself decides when the drawing is finished. The benchmark tests incremental visual planning over a short horizon and the transfer of creative ability from pretraining to multi-turn tool use. We evaluate 14 multimodal language models from small to frontier scale. Across the primary study and six sensitivity analyses, we collect 2,700 drawings and crowdsource creativity and recognizability ratings for every drawing and for 300 human reference drawings. We also present ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings (r = 0.85 on random held-out test split). Figural divergent thinking varies widely across the 14 models, and GPT-6 Astra produces the most creative drawings. Relative to the human drawings, the agent drawings score higher in creativity but lower in recognizability. We release the final drawings, per-round canvas snapshots, tool call traces, stimulus bank, benchmark harness, crowdsourced ratings (N = 72,000), and ViDrA checkpoint.
cs.AI / 30 / 2609.34211
Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning
Linbo Shao, Huilin He, Yating Lou, Dawei Cheng
cs.AI
Abstract
In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semantics, but at the cost of behavioral realism; textual descriptions for contextual reasoning remain scarce. We address this gap through a semantic enrichment framework grounded in original transaction behavior to simulate multimodal financial data. We (1) propose a multi-agent semantic enrichment framework that generates interpretable financial semantics grounded in transaction behavior through role-specialized agents and consistency refinement, and (2) newly contribute a valuable multimodal financial fraud dataset, MS-FFSD, enriched with structured semantics and textual semantics while preserving real-data-grounded transaction behavior. Furthermore, we systematically analyze the quality and utility of semantic enrichment. Results demonstrate statistical fidelity and framework generalizability, while showing that richer semantics benefit fraud modeling and context-aware LLM reasoning. Overall, this work advances multimodal financial fraud research and bridges emerging LLM and multi-agent capabilities with operational anti-fraud practice. The framework and dataset are released at https://github.com/AI4Risk/MS-FFSD.
cs.AI / 31 / 2609.34214
GlyphBench: A Playground for Language-Model Reinforcement Learning
Roger Creus Castanyer, Marc-Alexandre Côté, Matthew James Sargent, Augustine N. Mavor-Parker, Glen Berseth, Pablo Samuel Castro
cs.AI
Abstract
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.
cs.AI / 32 / 2609.34215
Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu, Jiangqiang Li, Jun Zhang, Junkai Ji
cs.AI · cs.LG · cs.MA
Abstract
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure that records which action wins without reflecting the absolute level of performance. When all actions fail, they tie at zero reward, and independent runs produce the same tied set with high probability, creating an illusion of stability that masks near-zero recovery success. We formalize this limitation through a set-path symmetry result, proving that for equal-cost Bernoulli actions the success probabilities (0.9, 0.8) and (0.2, 0.1) yield identical best-action-set distributions at every sample size. No procedure based solely on which action wins can distinguish these two regimes. We further prove that certifying exact population ties is impossible in finite time, and that the assignment of outcomes to checkpoints carries information beyond marginal outcome distributions. The pooled success probability is the missing scalar that resolves the ordinal ambiguity. Experiments on 864 frozen RecoveryBench episodes and two planning cohorts totaling 3,456 responses confirm the theoretical predictions. Agreement and held-out quality can move in opposite directions, and permuting checkpoint-to-action bindings changes 8 to 13 percent of cell-level conclusions. Based on these findings, we propose reporting four diagnostic quantities (agreement, all-zero fraction, held-out success, and pooled success) that expose this failure mode with no additional data collection.
cs.AI / 33 / 2609.34227
When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
Rishabh Sharma, Rishika Lall
cs.AI · cs.CL · cs.IR
Abstract
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
cs.AI / 34 / 2609.34241
AdaGuard: An Adaptive Guard Model with User-defined Policies
Yunhao Feng, Yifan Ding, Yuxiang Xie, Zheng Li, Mingrui Lao, Zeyuan Wang, Yanming Guo
cs.AI
Abstract
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard
cs.AI / 35 / 2609.34242
Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
Chidera Biringa, Lucas Yannul, Xiaowen Wang, Marco Ayala, Nicholas Yi, Alex Moyse, Nishant Manchanda, Vivek Gupta
cs.AI · cs.LG
Abstract
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
cs.AI / 36 / 2609.34249
Evolving Support Priorities in Empathetic Reinforcement Learning
Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, Xun Yang
cs.AI
Abstract
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
cs.AI / 37 / 2609.34262
Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
Weijun Luo, Kelvin Luu, Xinyi Liu, Guangze Luo, Miguel Romero Calvo, Soham Dan, Daniel Yue Zhang, Ying Liu, Mohamed Elfeki
cs.AI · cs.LG · cs.SE
Abstract
Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless can become exploitable, making benchmark validity an ongoing maintenance problem. We introduce a process-verification framework that audits passing trajectories, distinguishes evidenced reward hacking from verifier weakness, and localizes exploitable surfaces for repair. Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations often increase with model generation but not monotonically. On SWEBench Pro V1.0, confirmed violation rates rise from 24% to 73% between Opus 4.7 and Fable 5 on matched tasks; later cohorts fall to 11% for Fable 5.1 and 0% for GPT-6 Astra. These comparisons are descriptive: configurations were not normalized, and the latest models also pass fewer exploitable tasks. Violations concentrate around a small set of recurring surfaces, especially unintended access to reference solutions through git history. Three repair case studies across two benchmarks show why blocking a recorded exploit is insufficient: the same protected information can remain accessible through another route. Therefore, we combine minimal patches with exploit replay and fresh agent evaluation, auditing new passes under the original standard. No evaluated attempt against the final patches reached the protected channel, and every post-patch pass was judged legitimate. Benchmark integrity requires ongoing maintenance: audit passing behavior, repair the enabling surface, and re-evaluate both exploit access and legitimate solvability.
cs.AI / 38 / 2609.34273
Query Expansion and Key Specialization in Transformer Attention Geometry
Vidit Gupta, Siddhesh Nadkarni, Mihik Chaudhari, Vinaya Sawant, Prachi Tawde
cs.AI · cs.LG
Abstract
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that $PR_Q - PR_K$ is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of $QK^\top$ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.
cs.AI / 39 / 2609.34274
BIABench: Evaluating AI agents on real-world bioimage analysis tasks
Zixuan Pan, Davide Panzeri, Lukas Johanns, Marilin Moor, Yu Zhou, Hedi Peterson, Yiyu Shi, Jianxu Chen
cs.AI · cs.CL
Abstract
Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth. The tasks span eleven analysis subtasks and modalities from H&E histology to single-molecule localization microscopy. Each submission receives an outcome score, which compares the output files with the ground truth using field-standard metrics, and a process score, in which a vision-language model judges method choice and quality control against an expert-written rubric. We evaluated general-purpose and biology-specific agents across several language models, with repeated runs of every task. Routine two-dimensional tasks were solved well, but on some tasks that added a third dimension or a time axis no agent scored above 0.19. Neither biological specialization, stronger models nor detailed expert instructions closed this gap. The agents were also unreliable, with scores varying more between repeated runs of one agent than between different agents, and without ground truth a correct run could not be told from a wrong one by its process score or by the time spent. Released openly with its data and code, BIABench provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
cs.AI / 40 / 2609.34280
MoSPR: Histology-to-Gene Expression Prediction with Morpho-Spatial Macrostates and Low-Rank Molecular Programs
Dongmyung Shin, Geongyu Lee, Yesung Cho, Park Jong Bae
cs.AI · cs.CV · cs.LG
Abstract
Predicting molecular profiles from histopathology remains challenging because whole-slide images contain spatially organized, heterogeneous tissue patterns, while gene expression comprises thousands of correlated targets. We introduce MoSPR (Morpho-Spatial Program Regression), a linear framework that couples an adjacency-informed histology representation with a low-rank molecular basis. MoSPR clusters frozen patch embeddings into morphology microstates, aggregates their spatial adjacencies across the training cohort, and groups microstates with similar adjacency patterns into shared macrostates. Each slide is then represented by global morphology and macrostate-specific deviations, which are linearly mapped to coefficients of a training-derived low-rank gene-expression basis. Across three cancer cohorts from The Cancer Genome Atlas, MoSPR achieves the highest mean gene-expression prediction scores among all evaluated methods. Without pathway-level supervision, pathway scores derived from its predicted expression profiles rank first in eight of nine comparisons across three pathway collections. Ablation studies on the breast cancer cohort show complementary gains from adjacency-derived macrostate representation and low-rank molecular prediction. Moreover, with half of the training data on this cohort, MoSPR exceeds the full-data gene-prediction score of the strongest competing baseline. Finally, its linear formulation enables exact decomposition of each predicted expression profile into global and macrostate-specific molecular contributions, providing an interpretable link between spatially coherent macrostate regions and their associated molecular programs. Our code is available at https://github.com/Radisen-Panthera/MoSPR.
cs.AI / 41 / 2609.34313
ControlScope: Workflow Revision and Reliability in LLM Agents
Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li
cs.AI · cs.CL · cs.SE
Abstract
How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the actions an agent selects. We evaluate one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld. Across two source programs per task and three reasoning-reviewer draws on 20 filesystem tasks, FULL completes 15-16 tasks versus 13 for KEEP; across four fast draws it completes 10-13 versus 13. Fresh student-record confirmation reproduces a batch-read repair. ALFWorld fast panels yield KEEP/ARG/FULL scores of 85/86/87 on 87 tasks across 52 scenes and 134/134/127 on 134 tasks across four scenes; reasoning on the 87-task cohort also yields 85/86/87 with substantial review cost. An AppWorld V1 official-test panel of 585 task instances from 195 scenario templates shows small net differences. Frozen replays expose viable agent-written replacements interrupted by later revision in two failed file-organization runs. An offline source-trajectory midpoint comparison shows later reviews completing an insufficient repair. Five-call protection saves 19.4% of logged model output and loses one success across 20 fresh source runs. An argument-only shortcut shows that the broader sampled policy can overlook a cheaper successful edit available in both operation sets. These outcomes tie repair access to actual choices and subsequent execution.
cs.AI / 42 / 2609.34316
Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models
Kang Yang, Gaofeng Dong, Liying Han, Mani Srivastava
cs.AI
Abstract
This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.
cs.AI / 43 / 2609.34322
Test-Time Scaling via Budgeted Multi-Attribute Verification
Bo Xue, Ji Cheng, Shen-Huan Lyu, Yuanyu Wan, Shuang Qiu
cs.AI
Abstract
Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated along several costly attributes, and the goal is to certify as many candidates as possible whose mean scores exceed the prescribed thresholds on all attributes. We propose \textsc{BMA-GAI}, an algorithm that combines cost-aware arm selection with adaptive sampling of attributes. Every observation serves both to guide adaptive allocation and to support anytime-valid certification, which removes the need for a separate confirmation stage. We establish an asymptotic coverage guarantee for \textsc{BMA-GAI} and derive a matching information-theoretic converse that characterizes the intrinsic complexity of the problem, thereby proving that \textsc{BMA-GAI} is first-order optimal away from critical budget levels. Experiments on synthetic benchmarks and an LLM answer-verification task show that \textsc{BMA-GAI} allocates the verification budget more efficiently and certifies more high-quality candidates than competing methods.
cs.AI / 44 / 2609.34327
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
cs.AI · cs.CL
Abstract
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
cs.AI / 45 / 2609.34342
SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing
Zhiwei Chen, Tianchun Wang, Zhongtao Rao, Haiming Zhu, Ding Cao, Tianxiang Zhao
cs.AI
Abstract
A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents' behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupported strategic assumptions, inconsistent opponent estimates, and harmful interference from irrelevant historical interactions. To address these issues, we propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate. SAGE first anchors reasoning to an equilibrium policy that provides a strategically valid prior. It then conditions deviations from this anchor on a soft belief over opponent behavioral tendencies, enabling opponent-specific exploitation. Finally, SAGE distills strategically related interactions into counterfactual hypotheses about previously missing considerations, allowing past experience to recalibrate the model's reasoning. We evaluate SAGE on three repeated imperfect-information games: Leduc Hold'em, Liar's Dice, and Goofspiel, against various opponent types in each game. Compared with reasoning-intensive LLM agents, including Suspicion-Agent, ReTA, Agent-Pro, EMO, and Hypothetical Minds, SAGE achieves up to a 127.6% payoff improvement in Liar's Dice while reducing input and output token usage by up to 80% and 90%, respectively. In direct match-up play, it attains non-negative mean payoff against 5/10, 8/10, and 8/10 evaluated opponents in Leduc Hold'em, Liar's Dice, and Goofspiel, respectively, while using relatively fewer tokens. Code is available at https://github.com/chenzhwsysu57/SAGE.
cs.AI / 46 / 2609.34349
Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation
Michael Vasilakakis, Dimitris K. Iakovidis
cs.AI
Abstract
Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic and deep generative models, where their interpretability remains limited. This paper proposes a novel fuzzy distribution modeling methodology for synthetic tabular data generation based on fuzzy sets theory. Feature distributions are represented using fuzzy sets and feature dependencies are modeled through Fuzzy Cognitive Maps, resulting in a low-parameter, and an interpretable data representation. Synthetic samples are generated by sampling fuzzy concepts rather than raw values, enabling native support for mixed data types, missing values, and domain constraints. The methodology further supports linguistic queries and IF-THEN reasoning, facilitating transparent simulation of decision-making processes. Experimental results on benchmark datasets demonstrate competitive performance with respect to utility, fidelity and privacy compared to state-of-the-art methods, while offering substantially improved interpretability. These results establish fuzzy distribution modeling as a principled and effective approach for synthetic tabular data generation in fuzzy systems and decision support applications.
cs.AI / 47 / 2609.34353
SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception
Hu Xu, Chun Li, Siyuan Qiu, Zeyan Li, Jianfeng Xu
cs.AI · cs.IT
Abstract
Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open: which remote evidence is indispensable given the receiver's own observation? We introduce a closure-fidelity perspective on ego conditioned remote perception. Under a finite deductive abstraction and explicit conditions, its rate--distortion function decomposes over an irredundant core, and the exact zero-distortion rate becomes $P_A H(π_A)$. This analysis suggests a concrete design principle: transmit compact evidence and recover derivable context with bounded receiver-side inference. Guided by this principle, SemRD-V2X is an operational neural proxy that combines exact-budget BEV support selection, pointwise channel compression, and masked shared-weight reconstruction before standard fusion. Experiments on simulated V2XSet and real-world DAIR-V2X validate the resulting design. In a controlled five-run V2XSet comparison against a locally reproduced V2X-ViT-v1 baseline on one Tesla V100, SemRD-V2X reduces the analytical feature payload by $26.6\times$ while improving AP@0.5/AP@0.7 by 4.13/8.57 points, with 3.81\% additional mean compute latency. These results position closure fidelity as both an analytical lens and an actionable design principle for communication-efficient cooperative perception.
cs.AI / 48 / 2609.34360
CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency
Junwoo Bae, Jin-Hyun Ahn
cs.AI · stat.ML
Abstract
Split Federated Learning (SFL) enables resource-constrained clients to participate in collaborative training, but vanilla SFL exchanges smashed data and gradients at every batch, which incurs significant communication overhead. Recent methods reduce this overhead with an auxiliary network at the client-side cut layer. However, we identify that this approach makes the client optimize a local objective that differs from the end-to-end objective, which fundamentally limits the collaborative training between the client and the server. We propose Compensated Feedback based SFL (CoeF-SFL), a communication-efficient framework that retains the end-to-end objective without any auxiliary network. In CoeF-SFL, the client and the server exchange the smashed data and the gradients once per round and reuse them during local training. Since this reuse makes the gradients stale on the client side, we compensate them with a curvature-based correction in the activation space and develop two variants. CoeF-D approximates the Hessian with a diagonal gradient outer product, while CoeF-J exploits the tractable Jacobian-based Hessian of a surrogate loss that upper-bounds the true loss. We provide the theoretical background of each method, characterizing its compensation. Across vision and language tasks, model capacities, cut layers, and data distributions, CoeF-SFL significantly outperforms auxiliary-network-based methods under the same communication frequency, and the improvement is most substantial on vision tasks. Code is available at https://anonymous.4open.science/r/CoeF-SFL-2686/README.md
cs.AI / 49 / 2609.34372
PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents
Hanzhong Zhang, Ziwei Xiang, Weicheng Xie, Shizhe Liu, Siyang Song
cs.AI
Abstract
The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory processing to be inconsistent with the pre-defined personality, and makes it difficult to validate whether agent behaviours follow this personality. In this paper, we propose Personality-Integrated Memory (PersMem), which integrates personality into the agent's memory processing pipeline, making it consistently personality-dependent. PersMem processes memory using four steps, where the personality is mapped to operation-specific parameters controlling: (i) affective appraisal annotating emotion states of the user input; (ii) retention of previously stored memories along with the current input; (iii) passive affect-driven memory retrieval exploring memories similar to user input in semantics and personality-guided emotions; and (iv) active goal-driven memory retrieval that refines and selects passively retrieved memories for the reply. Consequently, consistency with the pre-defined personality can be examined by inspecting memory-processing traces during human-agent interactions. We evaluate these personality-dependent differences in attachment and Big Five settings. PersMem exceeds the chance baseline for four-way attachment classification by 23.1 percentage points. In Big Five dialogue comparisons, PersMem achieves 67.5% accuracy, 6.7 percentage points above a baseline using uniformly sampled memories. On CoSER, PersMem achieves an average score of 66.13, with scores of 69.33 for Character Fidelity and 84.33 for Storyline Quality. Together, these results show that PersMem produces distinguishable personality-related memory-processing patterns.
cs.AI / 50 / 2609.34392
Org-Agent: Beyond Personal Assistants Towards Organizational Agents
Luyao Zhuang, Yujing Zhang, Zijin Hong, Yilin Xiao, Xiao Huang
cs.AI
Abstract
Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-user memory and knowledge use. Both capabilities are governed by organizational constraints across three aspects: user identity, authority, and access permissions; the attribution and temporal validity of information; and rules for resolving conflicting requirements across users and completion requirements for joint decisions. These constraints shape what information or decisions must be obtained before an action can proceed and what conditions must be satisfied during its execution. Motivated by this, we introduce Org-Agent, a unified constraint-centric reasoning framework that organizes task execution in three stages. Specifically, Org-Agent decomposes a task into atomic subtasks and constructs a task dependency graph whose edges encode the dependencies among them. Building on this graph, it schedules the subtasks in dependency order through topological sorting. It then executes each subtask while accounting for the task's constraints, supported by evidence-acquisition and memory-management tools. Experiments on MUSES-Bench and GroupMemBench demonstrate the effectiveness of Org-Agent on both capabilities, and ablations further support the contributions of dependency modeling and tool use.
cs.AI / 51 / 2609.34410
Mathematics for and by human cognition: A resource-rational search for bottlenecks in problem-solving
Sneha Aenugu
cs.AI
Abstract
Human cognitive constraints are generally viewed as limiting factors in problem-solving. We argue that these constraints can instead play a critical role in driving advances in mathematics and beyond. We propose a theory of mathematical abstraction as a resource-rational search for bottlenecks in problem-solving. Bottlenecks arising from cognitive constraints create pressure to restructure existing knowledge, potentially giving rise to novel formalisms with applications beyond the problems that originally motivated them. Drawing on episodes from the history of mathematics, we illustrate how such bottlenecks can drive the development of novel abstractions and examine how cognitive constraints and affective responses shape this process. Finally, we discuss the implications of this account for machine mathematical discovery and argue that incorporating human-like constraints may facilitate the discovery of useful mathematical abstractions.
cs.AI / 52 / 2609.34418
OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue
Rui Xu, Yikai Zhang, Aili Chen, Zicheng Zhao, Xu Yinghui, Libo Wu
cs.AI
Abstract
Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.
cs.AI / 53 / 2609.34419
Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation
Hanzhong Zhang, Jindong Wang, Siyang Song
cs.AI
Abstract
Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG methods rely on end-to-end architectures that directly map speaker behaviours to listener expressions without an explicit intermediate internal state. We reformulate FRG as generation mediated by a structured internal-state process and propose the \textbf{Intentional Agent}, which shifts FRG from direct stimulus-response mapping to stimulus-grounded generation through explicit intermediate states. To represent temporal internal-state evolution, we propose an internal dynamics model that integrates emotional drives with an iterative Inner Thought Flow (ITF) within a structured intermediate state used for subsequent generation. This state can continue to update during conversational silences. Furthermore, to bridge abstract internal states with physiological actions, we formulate FRG as a downstream affective mapping from this latent thought flow to facial expressions. Experiments on the REACT 2025 dataset show an FRDist of 72.39 and an FRDiv of 0.5057; perceptual plausibility is evaluated separately through blinded human ratings. A blinded human evaluation of 96 reactions found no significant difference in mean score between Full and ground truth ($5.527$ vs.\ $5.195$, $p_{\mathrm{Holm}}=.076$), while Full significantly outperformed Event-Triggered and Heuristic-Only (both $p_{\mathrm{Holm}}<.001$). The Reaction Quality Scorer (RQS) correlated strongly with human judgements (Pearson $r=.855$; Spearman $ρ=.821$, both $p<.05$), supporting its use as an automatic metric. These results underscore the immense potential of endogenous dynamics in building highly autonomous, human-like agents.
cs.AI / 54 / 2609.34444
Social Circuits behind Multi-agent Echo Chambers
Chuiyang Meng, Wenlu Yu, Ming Tang, Cheng Li
cs.AI
Abstract
Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent's internal activations and affects its decision. In this work, we introduce Social Circuits, a framework for tracing message effects through receiver activations. We compare the receiver's answers before and after changing a message. Then, we restore selected activations recorded under the original message to determine how much of the message effect these activations reproduce. Based on Social Circuits, we propose Circuit-Guided Deliberation (CGD), which learns to select useful messages using receiver activation changes. We establish when activation replacement preserves receiver decisions and bound the gap between CGD's task performance and the best achievable through message selection. Experiments show that receiver activation changes explain the message effects and guide message selection that improves the task performance. Across three models and four datasets, CGD achieves the highest or joint-highest average accuracy in our main comparisons while generating fewer tokens than multi-agent baselines.
cs.AI / 55 / 2609.34459
Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning
Yijie Sun, Sanquan Sun, Yanda Zhu, Yuanyang Zhu, Yaohua Hu, Chunlin Chen
cs.AI
Abstract
Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challenges, we propose a novel interpretable framework, called escaping local views (ELV), which introduces semantically structured latent concepts to render policy decisions transparent. Specifically, each agent extracts low-dimensional semantic concepts from its local observation and action-observation trajectory. These concepts are jointly encoded into a contextual latent variable via a variational autoencoder (VAE), which builds a bridge between local views and global semantics. To explicitly model the decision of each agent, we employ a dual-path attention mechanism in which one module estimates the salience of individual concepts relative to the global context, while the other captures higher-order cooperative patterns with pairwise concept interactions. Furthermore, we incorporate a concept prediction module that derives an intrinsic reward from next-concept prediction errors, which incentivizes agents to explore regions of semantic novelty. Experiments in multiple environments verify that ELV not only achieves competitive performance but also explicitly provides how agents reason about their decisions.
cs.AI / 56 / 2609.34460
When Does Structured Knowledge Help Neural Theorem Proving?
Sareh Nabi, Roland Vogl, Marzieh Nabi
cs.AI · cs.LG · cs.LO
Abstract
Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analogies, generalizations, cross-domain bridges) remains implicit. We introduce MathAgent, which builds this layer as a knowledge graph, MathKG, and uses it to augment LLM theorem provers. MathKG connects 364 Mathlib theorems and definitions by 9,434 typed semantic edges inferred via LLM-based relation extraction anchored to verified Mathlib declarations. We run a controlled ablation across four augmentation modes (no context, knowledge-graph context, Mathlib retrieval, both) and five models: Qwen3-8B/32B, their Lean-specialized derivatives Goedel-Prover-V2-8B/32B, and Claude Sonnet 4.6, on miniF2F, plus PutnamBench and MathOlympiadBench for Sonnet. Three findings emerge. (i) Specialization dominates augmentation: Lean fine-tuning adds 33-38 percentage points of solve rate in every mode, and a specialized 8B model beats a $4\times$ larger general one by 29-35 points, while no augmentation mode improves solve rate by more than 3 points. (ii) Augmentation is capability-conditioned: knowledge-graph context helps small models but hurts large ones, with the specialized model gaining more relative to its general base at every scale. (iii) Yet the augmentation modes solve different problems: an oracle selecting the best mode per problem solves 6% to 58% more than the unaugmented prover, a complementarity effect that strengthens on harder problems (32% more on PutnamBench). These results motivate adaptive strategies that select augmentation by model capability and problem. Code, data, and artifacts are available at https://github.com/sarehnabi/mathagent
cs.AI / 57 / 2609.34506
Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
Manya Singh, Arjun Pakrashi
cs.AI
Abstract
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement ($ρ= 0.24--0.55$), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.
cs.AI / 58 / 2609.34519
EOPSA: Efficient On-Policy Self-Distilled Safety Alignment
Qirui Liu, Yichen Sun, Yan Wang, Yu Mi, Wei Cao, Yue Shen, Zhixuan Chu, Kui Ren
cs.AI
Abstract
On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher's corrective efficacy degrades precipitously as the student's generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by $\sim$50% and backpropagates through merely $\sim$2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.
cs.AI / 59 / 2609.34526
PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use
Mingfei Lu, Mengjia Wu, Yi Zhang
cs.AI
Abstract
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores ($Δ$) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6\% to 18.3\% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.
cs.AI / 60 / 2609.34545
Remember Before You're Asked: MemDream for Self-Probing Memory Evolution
Mingfei Lu, Mengjia Wu, Runsong Jia, Zhe Luo, Yi Zhang
cs.AI
Abstract
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
cs.AI / 61 / 2609.34557
SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
Bingqing Jiang, Guoxi Zhang, Jasper Wang, Auric Wang, Bingning Wang, Tianyi Lin, Zichao Yu, Yujin Han, Ziye Ma, Difan Zou
cs.AI
Abstract
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.
cs.AI / 62 / 2609.34565
FlowState: Execution State as Memory for Long-Horizon LLM Agents
Minghao Li, Bangyan Li, Zifan Wang, Yulong Li, Hu Xu, Gan Zhang, Jingtong Wu, Wenqiang Xu
cs.AI
Abstract
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.
cs.AI / 63 / 2609.34572
Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence
Rohit Saxena, Utkarsh Upadhyay
cs.AI · cs.CL · cs.LG
Abstract
Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model's subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.
cs.AI / 64 / 2609.34577
Calibrated Uncertainty for Informative Path Planning in Aquatic Environmental Monitoring
Samuel Yanes Luis, Alejandro Casado Pérez, Alejandro Mendoza Barrionuevo, Dame Seck Diop, Sergio Toral Marín, Saniel Gutiérrez Reina
cs.AI · cs.IR
Abstract
Informative Path Planning for scalar field reconstruction uses predictive uncertainty to direct sensing vehicles toward maximally informative locations. Gaussian Processes provide this signal but their stationary isotropic kernels are misspecified for non-homogeneous phenomena such as oil spills, producing miscalibrated estimates that degrade planning. We investigate whether replacing the Gaussian Process with a well-calibrated Deep Ensemble improves path planning outcomes, and whether uncertainty quality interacts with the choice of planning algorithm. Five strategies ($ε$-Greedy, Value Greedy, Uncertainty Greedy, Monte Carlo Tree Search, and Receding Horizon Orienteering) share a common Deep Ensemble backbone trained on physics-based oil spill simulations. On held-out stochastic spill scenarios, the Deep Ensemble reduces normalised reconstruction error by $83\%$ relative to the Gaussian Process baseline. Crucially, well-calibrated uncertainty amplifies the importance of the planning strategy: the performance gap between algorithms is negligible under miscalibrated models but becomes substantial under the ensemble, where multi-step lookahead planners outperform greedy selection by up to $32\%$ in reconstruction error and achieve IoU above $0.85$. Monte Carlo Tree Search is the recommended planner, matching Orienteering in reconstruction quality at an order-of-magnitude lower computational cost.
cs.AI / 65 / 2609.34582
SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences
Mingyue Huo, Shivam Mehta, Bhavin Jawade, Yinghong Lan, Haoqi Li
cs.AI · cs.SD
Abstract
Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.
cs.AI / 66 / 2609.34591
TULIP: Targeted LLM Unlearning at Layers Identified Per-Input
Yejin Kim, William F. Shen, Seokwon Jung, Daeun Park, Seong Joon Oh
cs.AI
Abstract
Representation-level unlearning intervenes on the intermediate hidden states of LLMs. Although knowledge is distributed across layers, existing methods operate at a single fixed layer for the entire forget set. We ask whether such a fixed layer is sufficient. To answer this, we design a hijacking experiment that grafts hidden states of the target model into an oracle trained only on the retain set. The oracle cannot produce the forget answer on its own, yet it produces the answer from the grafted state. Thus, the answer is formed at an intermediate layer and merely read out afterward, so unlearning should focus on formation, not readout. Moreover, the layer where formation ends varies widely across inputs. Motivated by these findings, we propose Targeted Unlearning at Layers Identified Per-input (TULIP). For each input, TULIP uses the logit lens to locate the formation-readout boundary and removes the hidden state's alignment with the forget answer's unembedding vector there. TULIP consistently outperforms output- and representation-level baselines on TOFU, PISTOL, and WMDP across Llama, Qwen, and Zephyr models. It also remains robust to paraphrase and quantization attacks. Beyond standalone use, its per-input layer selection serves as a plug-and-play component that further improves existing methods.
cs.AI / 67 / 2609.34603
After the Fix: How Corrected Agent Histories Transfer to Related Tasks
Yanfei Zhang, Xu Lin
cs.AI · cs.CL · cs.SE
Abstract
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-state inheritance, under eleven conditions. ThinkingBox's Full/Skill/Hybrid correction gains are 44/29/32 percentage points, with corrected performance 25/22/18 points above independence; inference weakens at the task-family level. Yet 12 of Full's 15-point larger correction gap over Skill come from worse uncorrected performance, not better corrected memory. Moreover, 22 of Full's 46 upward transitions restore observed baseline success. Neither APEX regime establishes comparable aggregate correction benefits. Action evidence connects workflow gains with reusable obligations and convention conflicts with source-local choices. Text APEX's accepted execution reaches 52% versus its summary's 40%, without robust global/group-level superiority or an estab- lished advantage over independence. Smaller handoffs reduce input but increase calls. The value of repairing experience is therefore distinct from the value of reusing it: memory updates require both a previous-version reference and a fresh-start reference.
cs.AI / 68 / 2609.34634
A Persistent State for Auditable Mixture-of-Experts Routing
Abdurrahman Javat, Allan Kazakov
cs.AI
Abstract
Mixture-of-Experts (MoE) models repeatedly route tokens to sparse subsets of experts, but conventional routers expose no routing-specific record of how cross-layer influences accumulate. We introduce Scratchpad-Augmented Mixture-of-Experts (SA-MoE), which gives each router access to a low-dimensional persistent state that is not provided to the experts. Learned layerwise writes update this state, and their realized post-update changes exactly decompose the state-mediated contribution to any later routing margin, forming a routing ledger. Across sparsely upcycled SmolLM2- and Gemma-based models and three independent training seeds per architecture, this pathway adds less than 1% analytical forward compute and is strongly used by trained routers: local removal of its router contribution changes the selected Top-2 expert set in 87.6% and 69.9% of decisions, respectively. Relative to a matched latest-write-only control, persistent accumulation increases long-horizon future-routing accessibility by 19.4 and 12.2 percentage points, with positive effects in every seed. More than 90% of absolute ledger contribution comes from non-recent writes in both families, and full-forward suppression of ledger-selected writes changes later routing and output distributions. The ledger is an exact provenance object for the persistent-state pathway, not a complete causal explanation of routing. Sensitivity-aware scores better predict full-forward intervention effects, and post-hoc methods recover related cross-layer attribution without architectural modification. SA-MoE instead makes one routing-specific computational history explicit and directly inspectable within the model's natural forward computation.
cs.AI / 69 / 2609.34649
Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
Weiyuan Li, Jinghan Xu, Aili Chen, Xintao Wang, Shuang Liang, Jiaqing Liang, Deqing Yang
cs.AI
Abstract
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.
cs.AI / 70 / 2609.34653
OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
Zongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu, Xuanzhe Liu
cs.AI · cs.LG
Abstract
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.
cs.AI / 71 / 2609.34654
A General Harness for Protein Foundation Model Fitness Prediction
Yang Tan, Qijia Tian, Gangyu Sun, Bozitao Zhong, Mingchen Li, Yuanxi Yu, Nanqing Dong, Liang Hong
cs.AI · q-bio.QM
Abstract
Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.
cs.AI / 72 / 2609.34687
VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience
Siqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu, Shufan Shen, Xihui Liu, Zhihua Wei, Tai Wang, Jiangmiao Pang
cs.AI
Abstract
Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.
cs.AI / 73 / 2609.34701
ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems
Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta, Soujanya Soni, Sameep Mehta
cs.AI · cs.LG
Abstract
Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.
cs.AI / 74 / 2609.34710
FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management
Peiyu Zang
cs.AI
Abstract
Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark that studies five configurable factors through controlled comparisons on a single simulator, using paired seeds and a frozen calibration. We evaluate four foundation models and four agent scaffolds. In the Model Track, Coach points Z-scores span 0.19, while Manager points Z-scores span 0.68, with GPT-5.6 showing a sharp rise in passivity under responsibility expansion. Its responsibility ladder rises from 46.1 to 58.1 points with recruitment, then falls to 46.8 under full management, localizing the regression to the final responsibility boundary. Across that boundary, its skipped-decision rate rises from 1.1% to 57.9%. Within the Flash-Pro pair crossed across every scaffold, scaffold choice changes Manager points Z-scores by up to 0.48 relative to the fixed stateless scaffold. The 3Y cohort shows a directional reversal in mean ranking between years one and three, while a selected Claude Code+Pro configuration peaks in year three and remains below that peak, showing that responsibility scope and horizon expose behavior changes that a single headline score conceals.
cs.AI / 75 / 2609.34715
PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs
Zhentao Tan, Jianrong Zhang, Ruijie Quan, Yi Yang
cs.AI
Abstract
Physical trajectories contain more than snapshots of a system: they also reveal how its states evolve under governing conditions. However, representation learning for parametric partial differential equations (PDEs) has largely relied on reconstruction-based objectives that emphasize recovering observed physical fields. In this paper, we investigate predictive representation pretraining as an alternative to reconstruction-based learning. We find that predictive representations preserve rich physical information, yet this advantage alone does not ensure accurate field evolution. Based on these observations, we introduce PDE-JEPA for parametric PDE dynamics. Specifically, we first train an encoder using a masked-latent prediction to capture the underlying regularities of PDE dynamics. To explicitly adapt the pretrained representation toward a more dynamics-aligned state space, we then introduce a geometry projector that aligns latent trajectory geometry with the evolution geometry of physical fields. Finally, building on this geometry-aligned latent space, we further develop a physics-structured latent predictor that decomposes the dynamics into parameter-independent evolution and parameter-dependent response components. Extensive experiments on nine widely used PDE benchmarks demonstrate that our framework outperforms existing state-of-the-art methods by an average of 33.4\% in-distribution, while achieving an average improvement of 51.4\% when extrapolating to unseen governing parameters. The project page is available \href{https://tanpig-x.github.io/PDE-JEPA/}{here}.
cs.AI / 76 / 2609.34735
From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram
Josef Pavlíček, Petra Pavlíčková, Irena Štrausová
cs.AI · cs.SD
Abstract
Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human composers operate within personal and cultural contexts that influence harmonic decisions and deliberate departures from established patterns. This study investigates six narrative-driven popular songs by Bob Dylan, Johnny Cash, and Ritchie Valens. Original human harmonies are compared with outputs of an explainable computational harmonizer operating on the same melodies without access to the original chord progressions. We examine harmonic vocabulary, functional persistence, repetition, non-diatonic events, and tension-resolution patterns using Chord Wheel Diagrams and BPMN-based representations. Results show that high melody-chord compatibility does not necessarily imply preservation of the original human harmonic decision pattern. Some generated harmonizations retain the economical structure of the reference, while others alter harmonic diversity or suppress distinctive events while remaining compatible with the melody. Rather than quantifying artistic quality, the study introduces narrative-conditioned harmonic structure as a complementary perspective for computational music analysis. The findings suggest that generative systems may benefit from modeling not only harmonic correctness, but also structural identity, context, and human compositional intention.
cs.AI / 77 / 2609.34768
Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision
Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin
cs.AI · cs.CV
Abstract
Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-modal teacher-student framework that lifts commercial radar to full-body, per-frame, metric 3D mesh reconstruction with per-joint uncertainty. Three innovations: (1) a mesh-foundation-model teacher - SAM 3D Body produces whole-body MHR ground truth (70 joints, 18,439 mesh vertices) from a single RGB frame with zero training, slashing annotation cost by orders of magnitude; (2) StudentPoseFormer - set encoding with masked attention pooling, a temporal Transformer, and a CVAE multi-hypothesis head that outputs both the pose mean and per-joint variance, honestly reporting where the radar cannot see; and (3) a multi-stage ground-truth quality pipeline (confidence gating, depth validation, temporal smoothing, bone-length consistency, bad-frame rejection) plus systematic information-lever ablations. On the public MM-Fi benchmark (same TI IWR6843 sensor, cross-subject), our full configuration reaches 7.45 cm 12-joint MPJPE, with ablations proving the causal value of point accumulation (k = 3, -0.34 cm), Doppler (-0.85 cm; -2 cm at the wrist on fast actions), and velocity loss (-0.27 cm). On our own synchronized radar + RGB-D corpus with block-level held-out splits, the pipeline achieves 21.47 cm end-to-end (per-joint hierarchy from 4.8 cm at the hip to 34.7 cm at the wrist - matching physical information limits), could be improved to 15 cm with ~30k diverse samples, and a scaling law shows sample diversity, not volume, is the binding constraint. Deployment inference is radar-only - no camera, no image.
cs.AI / 78 / 2609.34771
When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
Tianyi Guan, Jianhui Chen, Liangming Pan
cs.AI · cs.CL · cs.LG
Abstract
Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.
cs.AI / 79 / 2609.34776
Page-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFs
Abdelhak kelious
cs.AI
Abstract
We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the complete system scores 0.4450 MRR@10 and 0.4013 Recall@10, compared with 0.3430 and 0.2994 for BM25. Dense retrieval alone and a simple lexical--dense fusion both underperform BM25. Reranking improves the hybrid system, whereas adding query decomposition produces the largest further gain, with higher latency and more detected output artifacts. The complete system slightly exceeds the reported anonymous overall mean on two answer metrics but falls below it on most page-retrieval metrics. These results identify accurate page selection, rather than semantic retrieval in isolation, as the main opportunity for improvement in this setting.
cs.AI / 80 / 2609.34785
BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
Yuheng Wu, Berk Gokmen, Sujeeth Jinesh, Lauren McLane, Aarav Wattal, Qi Yang Huang, Zhaozhuo Xu, Thierry Tambe
cs.AI · cs.AR · cs.LG
Abstract
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B's RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.
cs.AI / 81 / 2609.34799
STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
Wanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun, Marlene Wagner, Verena Zimmermann, Mennatallah El-Assady
cs.AI · cs.CY
Abstract
Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.
cs.AI / 82 / 2609.34805
SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL
Zenghuang Fu, Ningqi Chen, Mingda Jia, Xiaofeng Han, Zhaoyang Li, Qiuyuan Ai, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, Minghui Wu, Guannan He, Changwei Wang
cs.AI
Abstract
Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score--outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by $1.31$ and $1.07$ percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected--fresh value gap alongside a near-zero fresh--fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at https://github.com/Zenghuang-Fu/SIPO
cs.AI / 83 / 2609.34810
UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning
Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Xiaofeng Han, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, Minghui Wu, Guannan He, Changwei Wang
cs.AI
Abstract
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of $82.8\%$ and $83.6\%$, WebShop success rates of $75.0\%$ and $82.0\%$, and Search-QA aggregate accuracies of $45.3\%$ and $49.8\%$, respectively. On 3B WebShop, UniOPSD improves over SDAR by $7.0$ percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd
cs.AI / 84 / 2609.34840
Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
Wolfgang Maass
cs.AI
Abstract
An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-χ)$ of the per-body value with $χ$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.
cs.AI / 85 / 2609.34848
Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
Yugu Li, Zehong Cao, Peizhen Li, Yang Zhang, Siyi Hu, Jianglin Qiao
cs.AI
Abstract
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5$\times$ reduction in token credit magnitude.
cs.AI / 86 / 2609.34850
From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
Jiangtao Lin, Bangyang Wei, Siyi Liu, Yihang Ding, Yuhan Dong
cs.AI · cs.LG
Abstract
Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewards produced by different objectives? We introduce assignment geometry to study this interaction. Mean-matched smoothing controls target dispersion, while within-stratum reassignment changes correspondence and preserves the complete target distribution. Across five reward objectives, intact correspondence retains the largest clean preference margins among the compared soft targets within a common accuracy-equivalence budget. Attenuation orderings change with the reward objective, revealing different responses to the same target assignments. Independent reassignments and a related source construction reproduce the retention direction. An attenuation-retention profile compares these combinations through margin magnitude, edit response, and accuracy. Against independently calibrated scaling, APLOT uniform targets deliver additional attenuation on both aggregate and presentation edits. These findings establish a joint design space in which target placement and reward objective shape reward properties beyond preference accuracy.
cs.AI / 87 / 2609.34854
AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces
Zhijie Deng, Ling Li, Junhao Ji, Siwei Lyu, Zhipeng Xu, Zulong Chen, Rongyao Fang, Shuai Bai, Xuming Hu, Jiaheng Wei
cs.AI
Abstract
Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.
cs.AI / 88 / 2609.34864
On the Limits of Metacognitive Monitoring in LLMs
Dongqi Han, Yifan Yang, Dongsheng Li
cs.AI · q-bio.NC
Abstract
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.
cs.AI / 89 / 2609.34886
Fewer Assumptions by Design: A Reusable Skill for LLM-Assisted Verus Verification
Andrada-Livia Antoneac, Dorel Lucanu, Dragoş Teodor Gavriluţ
cs.AI · cs.PL · cs.SE
Abstract
LLM-assisted Verus verification is a less tedious method to verify Rust implementations, but paired with self-referential structures, e.g., Doubly Linked Lists (DLLs)—notoriously difficult to formalise for verification—it becomes a substantially more demanding verification task. Moreover, a specification weakness can arise when verification relies on unproven or invalidated assumptions, such as axiomatic lemmas and assume statements. We investigate whether LLM agents can synthesize strong DLL specifications while minimizing these trusted base. The analysis follows three different approaches: manual verification, property-specific verification, and a defined skill for the specific case of DLLs and certain properties of this type of data structure. The skill encodes domain knowledge and a task-decomposition strategy. We show that an LLM agent equipped with a carefully designed verification skill can generate strong, low-trust specifications for DLLs in Verus.
cs.AI / 90 / 2609.34896
DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models
Qirui Liu, Yichen Sun, Yan Wang, Zhixuan Chu, Linbo Jiang, Jianan Lin, Kui Ren
cs.AI
Abstract
Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.
cs.AI / 91 / 2609.34902
Dual-Stream Simultaneous Translation via 2D Grid Attention
Yu Pu, Wei-Qiang Zhang
cs.AI
Abstract
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.
cs.AI / 92 / 2609.34913
DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
Jeremy Canale
cs.AI · cs.CR
Abstract
Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for. Dossiers are generated from canonical facts under 61 executable rules, with 42 authoritative records and 32 narrative documents; every gate is certified decidable from those records. Attacks never change an authoritative value, so an attacked dossier keeps the reference decisions of its clean copy. A success is attributable only when the agent receives the injection and takes the exact injected action, which it does not take on the paired clean dossier; the DGF score is the share of applicable fixed attacks a model blocks. Reading documents and records themselves, five of six models were outcome-strict (disposition, findings, actions and authorization all correct) on 82 to 85 of 85 gates. Over 2,622 attacked gate runs, seven direct-order, false-data and false-authority attacks obtained one attributable success against these five, whereas task-aligned attacks imitating the organization's own process passed against four of them: a record note citing a fake review procedure lowered GPT-6 Luna Pro from 34 to 6 outcome-strict gates and DeepSeek V4 Pro from 33 to 7. DGF scores ranged from 96.2 to 26.9, and a policy-aware adaptive attacker writing in records succeeded against five of six models. The approval tool executed no forged approval, yet deceived agents submitted approvals that the rules forbid. The open-source package dgf-bench computes the DGF score with one command.
cs.AI / 93 / 2609.34920
RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models
Dmitrii Kharlapenko, Sergei Bratchikov, Konstantin Korolev, Aleksandr Nikolich
cs.AI
Abstract
On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.
cs.AI / 94 / 2609.34948
Proactive Dialogue Policy Optimization via Cognitive-State Transition
Minghui Ma, Mengqi Chen, Bin Guo, Jingqi Liu
cs.AI
Abstract
Proactive dialogue requires agents to continually adapt their policies to user feedback while progressing toward task objectives over multiple turns. To move beyond imitation learning on static datasets, recent approaches use user simulators to collect interactive data for policy optimization. However, many simulators do not explicitly model the evolution of user cognition, limiting the consistency and state dependence of feedback across turns. Moreover, representing each action only by a high-level strategy label overlooks the large utterance space and cannot distinguish alternative realizations of the same strategy. To this end, we jointly design a $\textbf{Cog}$nitive User $\textbf{Sim}$ulator $\textbf{(Cog-Sim)}$ and $\textbf{C}$ognitive-$\textbf{S}$tate $\textbf{T}$ransition--Driven $\textbf{P}$olicy $\textbf{O}$ptimization $\textbf{(CSTPO)}$. Cog-Sim maintains the user's cognitive and affective states and generates responses through constrained state transitions across turns, so feedback depends on both the realized utterance and the user's current state. CSTPO organizes each action as a hierarchical strategy--utterance representation: a high-level strategy label constrains utterance sampling, and utterances are optimized within each label. Sparse complete-branch sampling reuses shared dialogue prefixes and estimates separate strategy-level and utterance-level advantages, enabling fine-grained optimization at both levels. Across three tasks, Cog-Sim exhibits monotonic dose--response relationships and is preferred over prompt-based simulators for naturalness. CSTPO improves Qwen3-14B's performance to a level comparable to that of GPT-5.5-based planning methods.
cs.AI / 95 / 2609.34949
VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang, Yifang Xu, Yuchuan Wu, Xiaolei Chen, Zhengtao Yao, Fan Shi, Yang Liu, Bin Li, Xiangyang Xue
cs.AI
Abstract
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.
cs.AI / 96 / 2609.34951
AX is the New AEO
Ido Finder, Assaf Elovic, Gad Shalev
cs.AI · cs.IR
Abstract
In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter breadcrumbs across forum threads, listicles, and off-site citations, so AI engines are likelier to surface and recommend them. But being surfaced is no longer enough: an agent opens the results and reads them before deciding, and one buyer question sends it through several rounds of search and fetch. What decides the outcome at this drill-down step is whether the agent can fetch and read the business's own site: agent experience (AX). We argue that AX is the new AEO. We run 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses, matched on fame, prior model knowledge, and two AEO proxies, then split based on their AX level. Only 7-10% of the finished answer comes from the model's training knowledge, whether or not the site is readable. Agent-ready businesses have answers built from their own pages 78% of the time against 56% and are clearly recommended 1.9x more often, while every grounded answer about a not-agent-ready business costs the agent 64% more. Holding business, harness, and question fixed, answers built from the site are 41% more accurate. The dominant failure is not fabrication but omission: web-built answers are 3.7x more likely to contain none of the facts the buyer asked for. Baselines differ sharply across the four harnesses, with clear-recommendation rates varying sevenfold from stack to stack, yet the effect holds in every one. In the agentic web era, being readable beats being talked about, and improving a site's AX is the strongest lever a business has.
cs.AI / 97 / 2609.34960
ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization
Feiming Wang, Daibo Li, Kun Yuan
cs.AI · math.OC
Abstract
Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system for Proof-Obligation-Driven Theory Construction. Given a published algorithm, target theorem, and source proof, ProofLoom autonomously constructs the Lean model and supporting theory. Open proof obligations drive the development of definitions, interfaces, lemmas, and proof plans. Signature contracts record evidence and obligations for model revisions; an independent Judge rejects unsupported assumptions and weakened conclusions. Planner expands the published argument into intermediate claims, and Audit checks whether the Lean proof follows it. Across tasks, SOptLib accumulates verified mathematics and construction experience: reusable results are extracted, generalized, and verified, while modeling decisions and failed proof routes are recorded. Later tasks retrieve these results and records and contribute new developments, forming a cycle of construction, accumulation, and reuse. On fifteen textbook and research-paper tasks, ProofLoom obtains mean human ratings of 6.3/7 and 6.4/7, compared with 4.9/7 and 5.0/7 for the strongest of six baselines. Across 33 developments, it produces 490,693 lines of algorithm-local Lean code with no sorry. The formalizations also expose 28 incorrect formulas, proof gaps, and algorithm-analysis mismatches in published sources across 22 developments, each with checked evidence.
cs.AI / 98 / 2609.34966
Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks
Hangzun Liu, Yuling Fan, Fang Tian, Zhilong Bie, Zaiwen Feng, Yongliang Qiao
cs.AI
Abstract
Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.
cs.AI / 99 / 2609.34973
APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
Puneet Mathur, Dinesh Manocha
cs.AI
Abstract
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.
cs.AI / 100 / 2609.34974
Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces
Ruozhao Yang, Mingfei Cheng, Xiaofei Xie
cs.AI · cs.CR
Abstract
LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArena-Single and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.
cs.AI / 101 / 2609.35013
Automated feature engineering, AutoML, and decision-focused learning for improved energy consumption forecasting
Nasser Alkhulaifi
cs.AI
Abstract
The rising cost and demand for energy, together with environmental sustainability goals, create major challenges for energy management. Energy Consumption Forecasting (ECF) supports planning by predicting future consumption, but Machine Learning (ML) models for ECF often depend on expert-driven Feature Engineering (FE). This thesis addresses that dependence through three contributions. First, it establishes and evaluates a comprehensive FE pipeline for ECF and investigates domain-specific features. Second, it introduces AutoEnergy, a domain-tailored automated FE algorithm that generates interpretable features from timestamps and lagged consumption and integrates with AutoML for end-to-end ECF modelling. Across eighteen real-world energy datasets spanning residential, commercial, industrial, renewable, and grid domains, AutoEnergy reduces forecasting error by 19.52%-84.72% relative to baseline AutoML and established automated FE methods, while running 1.31-4.41 times faster, with gains varying by dataset. Third, AutoEnergy is integrated with Decision-Focused Learning (DFL) for a Battery Energy Storage System problem, jointly forecasting electricity prices and demand while optimising charging and discharging decisions. On a real-world UK property dataset, this approach reduces operating costs by 22.9%-56.5% compared with the same DFL models without automated FE. Overall, the results show that domain-specific automated FE can reduce reliance on manual feature design, improve forecasting accuracy, and translate predictive gains into measurable operational benefits in energy management.
cs.AI / 102 / 2609.35017
TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation
Nicolas Dahan, Fran{\cc}ois Yvon, Rachel Bawden
cs.AI
Abstract
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assigns an interpretable verdict to every term occurrence: glossary-conforming occurrences are settled deterministically, while divergences are assessed under a two-step LLM-as-judge procedure using the full document context: the first detects and labels terminology errors; the second sorts valid document-level variations from inconsistencies. Validated against expert error annotations and document-level human MQM scores, TermJudge ranks first in both system- and segment-level meta-evaluation, ahead of glossary-conformity and quality-estimation baselines. When applied to eight systems translating academic documents, under two prompting conditions, we observe that glossary injection improves terminology translation in all paired comparisons, by removing genuine errors rather than valid variation. TermJudge is released as open-source code.
cs.AI / 103 / 2609.35018
Environmental requirements for the use of social information by artificial life agents using evolved plastic artificial neural networks
Hugh Charterton, James M. Borg, Aniko Ekart
cs.AI
Abstract
Evolved Plastic Artificial Neural Networks (EPANNs) consist of two principal processes, the first, evolution, and the second, development and in-life learning. In the context of the origins of social = learning, very few studies have been carried out using ALIFE models based on EPANN requirements. Studies in this field have usually involved an imitative teacher/pupil relationship. This, however, ignores the possibility that the observed behaviour is a consequence of social information cues rather than direct imitation or teaching. Starting with the first of the EPANN processes (evolution), a series of experiments was undertaken using artificial neural network (ANN) based agents in a variety of foraging environments to examine under what minimal environmental conditions the use of social information might have evolved, as measured by the number of generations taken to meet a specified fitness criterion. NEAT (Neuroevolution of Augmenting Topologies) was the ANN used as its evolutionary algorithm would evolve a network's topology as well its weights. Unintentionally, in the experiment there was a simple network topology based on the location of the nearest food item which enabled agents to swiftly meet the fitness criterion. With this topology, additional information, social or otherwise, was not required and could have proved to be a hindrance. However, this does indicate that for the use of social information to have evolved, it would require a greater degree of complexity in the environment to do so.
cs.AI / 104 / 2609.35025
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
Haotian Luo, Haoyu Wang, Zeyu Qin, Huanjin Yao, Yibo Wang, Zhuotao Tian, Shuai Wang, Jiaya Jia
cs.AI
Abstract
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collaboration. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent's ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where data is delivered sample by sample and each sample is accepted against a set of criteria rather than put straight into training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task for the same suite that meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. AutoDataBench provides a direct measure of an agent's capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run. Code and data are available at https://github.com/StarDewXXX/AutoDataBench.
cs.AI / 105 / 2609.35026
WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova
cs.AI · cs.CL
Abstract
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with parameters as a user or an agent acts. A task declares the events it requires, and success is decided by matching them, with no judge model and no scraping of rendered pages. The same instrumentation supports controlled UI variation: one configuration switch re-renders a task through a different implementation of a single control while the prompt and the success conditions stay completely identical, so sensitivity to interface form can be measured under a fixed task specification. The WebPageBench release consists of three components: 152 tasks, divided into 65 canonical scenarios and 87 control variants across light/dark UI-modes; a common runner evaluated with six browser/DOM harness configurations and five screenshot-only GUI-agent families; and a public leaderboard of 24 model-harness pairs. On the public 152-task leaderboard the gap between what agents declare finished and what the log confirms reaches 41 points (one configuration declares every task finished and satisfies the conditions on 59%).
cs.AI / 106 / 2609.35032
JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
Zhixi Cai, Fucai Ke, Sukai Huang, Maria Garcia de la Banda, Peter J. Stuckey, Gholamreza Haffari, Hamid Rezatofighi
cs.AI · cs.CV · cs.RO
Abstract
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at https://github.com/ControlNet/JRDB-AVR.
cs.AI / 107 / 2609.35088
When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents
Geonwoo Kim, Brent ByungHoon Kang
cs.AI · cs.CR · cs.SE
Abstract
Tool-enabled agents form calls from model-visible interfaces, while hosts later select their implementation. Standard dispatch omits the descriptor-handler relation. An unchanged and schema-valid call can therefore acquire a different security effect during rollout, reconnect, or delayed approval. We call this failure schema-epoch drift. We present formation-consistent dispatch (FCD), which connects implementation analysis to execution authority. Reviewed profiles produce provenance-bound over-approximations of declared in-scope effects from official source. Under a closed-target approval policy, a verifier applies each formed call to a summary and captures a successor only when its effects fit the call's security contract. Atomic admission and a final-hop fence preserve this decision to the effect. The exact source retains priority, and the captured successor becomes eligible only after source retirement. Stock releases and deployment changes reproduced the failure. Four profiles covered 32 official releases: 29 required no release-specific change and three escalated. A frozen 16-release expansion matched a separate source oracle. In a preregistered stock comparison, FCD completed all three pending calls whose effect remained private and blocked all three whose omission became public. Exact pinning and release-wide denial stopped all six calls, while release-wide approval completed all six but produced three public effects. A separate lifecycle experiment carried a formation-captured certificate across source retirement. The same safe certificate installed later governed new formations without expanding the pending call's authority.
cs.AI / 108 / 2609.35089
Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
Zehao Lu, Xingguo Xiong, Wopke van der Werf, Thijs L. van der Plas, Ioannis N. Athanasiadis
cs.AI
Abstract
Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches---direct zero-shot prompting, a staged workflow, and a multi-agent system---with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model--approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.
cs.AI / 109 / 2609.35107
DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery
Yulong Li, Rong Xia, Yuxuan Zhang, Jianxu Chen, Xiwei Liu, Haochen Xue, Maosheng Li, Yuhang Liu, Yibo Yuan, Yutong Xie, Chong Li, Jionglong Su, Hagai Rossman, Eran Segal, Imran Razzak
cs.AI · q-bio.QM
Abstract
We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, from longitudinal clinical phenotypes, medical imaging, and continuous physiological signals to eight molecular layers, together with an evidence network of approximately 4.7 million literature-derived records over 93,566 concepts and 149,383 candidate causal relations. DoAtlas-2 autonomously formulates research questions from evidence gaps and unresolved mechanisms, prespecifies their causal designs, and generates validated analyses. Supporting, challenging, and unresolved results continuously revise mechanistic interpretations, the causal evidence state, and the discovery frontier, so that DoAtlas-2 self-evolves within a closed loop of hypothesis generation, empirical testing, and renewed discovery. DoAtlas-2 has systematically evaluated 2,031 research questions. In the Human Phenotype Project (HPP), it formulated 4,014 candidate pathway questions across vascular, early-glycemic, and hepatic-metabolic systems, and screening of the first 1,079 yielded statistical support for 756. Representative studies identify blood pressure as a convergence node linking adiposity, hepatic, and lipid phenotypes to vascular outcomes, and show that an adiposity-inflammation-blood-pressure pathway is largely attenuated by joint adjustment for body mass index (BMI) and smoking. The discovered vascular network constitutes a completely interpretable predictive foundation, admitting exact attribution of every prediction and closed-form mediation effects. DoAtlas-2 thereby unifies causal mechanism discovery, population-evidence testing, and interpretable prediction within one continuously evolving foundation.
cs.AI / 110 / 2609.35115
DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines
Haixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie
cs.AI
Abstract
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime's speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-
cs.AI / 111 / 2609.35149
From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale
Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan, Jiaxing Song
cs.AI · cs.MA · cs.SE
Abstract
Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retention or target-contract satisfaction. We formulate agent calibration as constrained behavioral adaptation across three interacting layers: information preservation, harness adaptation, and user acceptance; the layers apply to every scenario, not one-to-one to the three. The basic objective is non-degradation on prespecified capability measures while satisfying target requirements; aggregate improvement is stronger. Information calibration preserves independently validated source content still applicable to the target task. Harness calibration aligns observable artifacts at semantic checkpoints and repairs them through iteration, tool substitution, or local replanning within explicit budgets. User calibration enforces recipient-specific output contracts: templates, schemas, and section-level preferences. A global e-commerce example shows how shared standards coexist with site- and market-specific adapters and validation. We distinguish trainable policies from frozen-backbone configuration or controller optimization, and evidence verification from relative judgment and DPO/GRPO optimization. Recent harness-transfer and judge-validity studies motivate target-native execution records, separate audits of task validity and near-tie ranking, and matched target-native optimization controls. We propose held-out evaluations for model changes, cross-border adaptation, and scale, including a factorial test of source evidence and checkpoint repair and group-level reporting to prevent aggregate gains from masking local failures. This is a methodological proposal; implementation and empirical validation remain future work.
cs.AI / 112 / 2609.35158
PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
Jiaan Zhu, Wei Gao, Youhui Bai, Zewen Jin, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Cheng Li
cs.AI
Abstract
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$--$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.
cs.AI / 113 / 2609.35160
FONDANT: Strong and Best-Effort Planning via Antichains
Benjamin Aminof, Tuan Khai Nguyen, Sasha Rubin
cs.AI
Abstract
A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available or there is no evidence that the environment is adversarial, one can resort to best-effort policies, which always exist, and which follow the classic decision-theoretic principle that an agent should not use a dominated strategy. A typical positional best-effort policy works as follows: from every state, it follows a strong policy if one exists from that state (such states are called ``strong-winning''), else a weak policy if one exists from that state (``weak-winning''), and else is unconstrained (``losing''). In this work, we introduce a sound and complete planner for both best-effort planning and strong planning. The algorithm that underpins the planner is quite simple: it represents certain sets of states, such as the winning regions, by their $\subseteq$-minimal elements. The algorithm returns uniform policies, i.e., it returns a policy $π_t$ that is a strong solution starting in every strong-winning state, and it returns a policy $π_w$ that is a weak solution starting in every weak-winning state, and it provides a certificate for the set of losing states. We implemented the algorithm with some simple optimizations (calling it FONDANT), and evaluated it on a benchmark set consisting of the instances that were used in the evaluation of leading strong planners PR2 and FOND-SAT, and the best-effort planner BeSyftP. On coverage, our implementation is at least as good on all domains, and outperforms on some domains; and on wall time, it is slower on small and medium-sized instances, and outperforms on larger instances.
cs.AI / 114 / 2609.35184
5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding
Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan, Jiaxing Song
cs.AI · cs.CL · cs.IR
Abstract
Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later tasks may be unable to use information that was omitted. If the index retains only open-ended text, however, rule-based reasoning lacks checkable premises. We propose 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding. The 5W1H questions organize source-grounded content units; Which points to versioned ontology elements and records mapping relations, scope, and validation status. Time, location, system environment, and participant roles are not merely retrieval labels: together, they constrain the contexts in which facts, bindings, and rules apply. Unbound content remains searchable, while bound content enters a formal reasoning path only after premise checks. The method further distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions. A worked example of migration from an on-premises server to a cloud environment illustrates the different treatment of world-state changes, ontology-version changes, and changes in rule applicability. We formulate three groups of falsifiable hypotheses concerning cross-task evidence coverage, control of contextual misuse, and incremental update cost. The planned evaluation includes a strong typed fact-graph baseline with the same evidence, temporal information, and budget, to test whether benefits arise from 5W1H organization, deferred binding, or additional information and engineering effort. The contribution is a testable indexing mechanism, not a claim to a new universal ontology or a demonstrated performance advantage.
cs.AI / 115 / 2609.35215
ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
Yang Li, Jinhan Yang, hai liu, Di Wan, Xiyu Chen, Zongsi Xu, Tuo Zhou, Sheng Zhong, Sergey Volkov, Ye Luo, Hao Sun
cs.AI
Abstract
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.
cs.AI / 116 / 2609.35261
Imprint Reader: From Weight-Update Readout to Behavioral Intervention
Guanxu Chen, Qihao Lin, Jing Shao
cs.AI
Abstract
As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter-level traces that can, in principle, be inspected directly. However, current models cannot decode these traces into an explicit account of what they have learned. To this end, we introduce the \textit{Imprint Reader}, a model trained with \textit{Semantic Mount-and-Read Tuning} (SaRT) to describe frozen weight updates. SMaRT mounts each update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, while no-change and random-perturbation controls discourage unsupported claims. On held-out updates, the joint Reader reaches judge-based Pass@100 of $2\%$ for knowledge and $16\%$ for behavior. These results demonstrate the feasibility of natural-language readout while pointing to reliability across updates as the next step. Beyond free-form generation, the Reader provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update. Its coordinate-aligned gradients support intervention through MetaEdit. At a $0.5\%$ pruning rate, Reader-guided selection raises measured harmful-prompt refusal from $57.9\%$ to $64.1\%$ under a safety-maintenance target. Using behavior descriptions without target-task training data, MetaEdit increases the frequency of backtracking and sub-goal expressions in mathematical reasoning traces and raises BFCL Overall from $41.69\%$ to $44.60\%$.
cs.AI / 117 / 2609.35286
The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
Michele Loi
cs.AI
Abstract
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.
cs.AI / 118 / 2609.35290
EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
Shihan Dou, Shaofan Liu, Zhonghang Lu, Jiahang Lin, Shichun Liu, Binghai Wang, Jiajie Jin, Guanting Dong, Tao Gui, Qi Zhang, Xuanjing Huang
cs.AI · cs.CL
Abstract
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model's own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document's length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.
cs.AI / 119 / 2609.35296
AbGaze: Attentive Geometric Representation Learning for End-to-End Antibody Design
Jiashuo Wang, Siqi Fan, Yizhen Luo, Zaiqing Nie
cs.AI
Abstract
Computational antibody design requires representations that capture the geometric patterns underlying antigen--antibody interactions, yet existing approaches often rely on scalar distances or surface-intrinsic features, leaving cross-molecular geometry largely implicit. We present AbGaze, an end-to-end antibody design framework based on attentive geometric representation learning, which encodes distance, spatial direction, and surface-normal orientation of antigen surfaces relative to antibody-residue local frames, and adaptively aggregates these geometric interactions according to their interfacial context. The learned interaction representation is shared across multi-CDR co-design, complex structure prediction, and affinity optimization, with local-frame geometric supervision further constraining the representation. AbGaze outperforms prior methods across all three tasks: relative to the second-best method, it improves amino-acid recovery by 7.1% and reduces structural error by 14.9% on average over the six CDRs, improves interface docking quality (DockQ) by 6.6%, and raises the affinity improvement rate (IMP) by 32.5%.
cs.AI / 120 / 2609.35298
Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework
Surajit Das
cs.AI
Abstract
Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Information Gate, knowledge-weighted evidence accumulation, disease similarity, and decisive clinical rules. Missing-aware normalization and coverage auditing distinguish absent from unavailable evidence. Candidate ranking is separate from outcome-label-independent K-means clustering, which uses four derived evidence coordinates (evidence strength, relative magnitude, directional similarity, and evidence completeness), not raw predictors or targets, to derive cohort-level assignments. Across six retrospective cohorts - four dengue (N = 1000, 1523, 989, 1018), malaria (N = 2190), and influenza (N = 4569) - a uniform, label-free, cohort-fitted K = 2 protocol yielded positive-class F1 scores of 0.996, 0.634, 0.936, 0.917, 0.695, and 0.842, and all-record accuracies of 0.996, 0.558, 0.914, 0.893, 0.707, and 0.906, respectively, with full partition-decision coverage using the frozen package and disease-specific knowledge representations. Neither scoring nor clustering uses outcome labels. Logistic regression provides a supervised baseline. Influenza incorporates confirmatory molecular PCR and is not independent pre-test prediction. Results characterize knowledge-grounded evidence separation, auditability, and sensitivity, not prospective clinical validity or comparative superiority. FOL/LLM-based clinical explanation remains unevaluated.
cs.AI / 121 / 2609.35316
Reliability Engineering for AI Systems: Challenges, Methods, and Directions
Rong Pan, Yili Hong, Min Xie
cs.AI · stat.AP
Abstract
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.
cs.AI / 122 / 2609.35328
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
cs.AI
Abstract
Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm's bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While MetaBBO helps advance the performance lower bound of the resulted optimization system, it is currently handcrafted and customized case by case to adapt different optimization problems, which inevitably introduces inherent subjectivity and hence restricts the performance upper bound and usability in practice. In this paper, we address this issue by regarding MetaBBO's design loop as coding task, where we could introduce openendedness into MetaBBO with recursive self-improvement capability of advanced coding agents. Specifically, we propose a dual-agent framework: i) a task agent continuously refines the codebase of a target MetaBBO approach through code evolution; ii) a hyper agent progressively modifies the task agent and itself to provide open-ended design behavior; iii) the evolved MetaBBO codebase is evaluated and all in-execution information is fed back to the agents for recursive self-referential improvement. As a result, given a naive MetaBBO template, our framework automates a design evolution and finds novel variants superior to up-to-date human-made MetaBBO baselines. Surprisingly, the experimental results also demonstrate that our framework supports fast adaption across different optimization domains. Solid interpretation analysis further reveals interesting design principles emerge in such open-ended process. This work serves as the first exploration on automating design of complex learning-assisted optimization algorithms.
cs.AI / 123 / 2609.35342
Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
Riccardo Porcedda
cs.AI · cs.LO · eess.SY
Abstract
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.
cs.AI / 124 / 2609.35350
Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
Lucas Biechy, Cédric Eichler, Adrien Boiret, Nicolas Anciaux
cs.AI · cs.CL · cs.LG
Abstract
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
cs.AI / 125 / 2609.35356
Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
Kajetan Dymkiewicz, Tim Farrelly, Adam Práda, Ishaan Panigrahi, Srishti Gureja, Helen Yannakoudakis, Robert Mullins, Victor Gillioz, Daniel Tan, Maxime Riché
cs.AI
Abstract
Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours co-occur in most training examples, so filtering out examples with undesired behaviour leaves only a small clean subset. We introduce stratified inoculation prompting (SIP). SIP leverages a small clean subset to demonstrate that desired behaviour should persist without the undesired one across different contexts. SIP oversamples these clean examples under diverse non-eliciting prompts while inoculating the rest. SIP substantially reduces expression of undesired behaviour while preserving more of the desired behaviour than IP. These gains persist even when we extend IP to oversample the same clean subset at the same rate as SIP. Moreover, SIP yields lower emergent misalignment rates in all harmful-advice setups we tested. SIP can be further extended to limit the undesired behaviour even under prompts that explicitly request it. We introduce backdoor dilution, which weakens expression under the inoculation prompt, and password-locked inoculation, which concentrates elicitation on a designated password. Taken together, our findings show that changing the training contexts for a small clean subset can significantly improve selective generalisation.
cs.AI / 126 / 2609.35370
A decision-support system applied to Law: Reasoning and explainability of the decision
Jeremy Bouche-Pillon, Pascale Zarat{é}, Yannick Chevalier, Nathalie Aussenac-Gilles
cs.AI
Abstract
The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ''Law Enforcement Directive (LED)'', appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal representation of these regulations can be part of decision systems that support LEAs in processing data in compliance with the regulations. Although many new formalisms have emerged to represent legal norms and rules, few are provided with a reasoning mechanism. Furthermore, systems used in decision-making processes in critical contexts such as medical diagnoses or legal decisions cannot be fully automated, and the explainability of their results is essential to ensure user confidence in decisions. This explainability aspect, while crucial, is lacking in most modern approaches that rely on machine learning. This paper describes a framework to operate formal rules from regulations, by focusing on explainability of the decision. After describing the general architecture of the proposed decision support framework, the paper showcases how symbolic AI and the SPARQL query language can support legal reasoning. It then describes an algorithm to generate a justification for the reasoning results, and outlines the procedure to be followed when the reasoning does not lead to a satisfactory conclusion. We notably focus on a method based on decision trees to determine what additional information to request from the user.
cs.AI / 127 / 2609.35400
Structural Alignment for Reliable Industrial AI: Bridging Physical Reality, Data, Models, and Human Intent
Lizhi Xiao, Sihong Wu, Victoria Xiao, Yiqiao Song, Chen Gu, Jianwei Ma, Xinming Wu, Aimé Fournier
cs.AI · physics.geo-ph
Abstract
Artificial intelligence is increasingly deployed in critical industrial domains, including healthcare, energy grids, subsurface exploration, where failures can have severe consequences for human safety, system stability, and economic outcomes. Yet AI is still evaluated primarily through benchmark accuracy, a model-centric metric that fails to capture the structural complexity and risks of real-world deployment. We propose a framework that views industrial AI reliability as a problem of structural alignment across four interacting worlds: physical, representational, machine, and human cognitive. These worlds are connected through two interfaces: digitalization, linking physical reality to computational representations, and goal encoding, translating human cognition to the machine objectives. Together, they define the space of admissible solutions. We characterize the solution space through four attributes: existence, non-uniqueness, robustness, and interpretability and show how mismatches arise at interfaces and propagate across worlds to produce reliability failures. Applications to healthcare, energy grids, and subsurface exploration illustrate that although dominant failure modes differ across domains, for example, interpretability in healthcare, robustness in energy grids, and non-uniqueness in subsurface exploration, all originate from a shared structural mechanism. By shifting the focus from model-centric evaluation to system-level alignment, this framework offers a principled foundation for assessing and governing reliability in industrial AI systems.
cs.AI / 128 / 2609.35412
Self-Adapting Group of Experts for Multi-Agent Reasoning
Mohammad Atif Quamar, Nurbek Tastan, Karthik Nandakumar, Junpei Komiyama
cs.AI · cs.CL · cs.MA
Abstract
Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often adapt communication by changing this context while leaving individual prompts fixed, even when a problem calls for different skills. We study whether agents' initial responses can identify a strategy better suited to the current problem and guide its transfer to other agents. To address this, we introduce SAGE (Self-Adapting Group of Experts), a training-free framework that uses answer agreement, prefix consistency, and reciprocal peer review to select a strategy donor. SAGE transfers the selected donor's reasoning strategy to the other agents while preserving their original roles. This transfer uses only the agents' original system prompts, without access to the problem or generated solutions. After strategy adaptation, agents exchange responses through a dynamic, sparse directed acyclic graph that routes information from higher-scoring agents to lower-scoring agents. Experiments across multiple agent backbones and reasoning benchmarks show that SAGE achieves higher average accuracy than the evaluated baselines. Our code is available at https://github.com/atifquamar07/sage.
cs.AI / 129 / 2609.35436
Building Transformation Layers for Riemannian Neural Networks
Ziheng Chen
cs.AI · cs.LG
Abstract
Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a framework for constructing FC and convolutional layers over computationally tractable Riemannian spaces. This framework incorporates several previous FC layers across different geometries as special cases and is instantiated on ten representative manifolds, including three hyperbolic models, five geometries of the symmetric positive definite (SPD) manifold, and two Grassmannian perspectives. Experiments on different manifolds demonstrate the effectiveness and applicability of our approach. Code can be found at https://github.com/GitZH-Chen/RieTrans.
cs.AI / 130 / 2609.35443
Just Initialize: A Training-Free Initialization Component for Large-Scale Routing Optimization
Jiale Zhao, Sirui Mao, Zimu Chen, Wentao Yang, Zihan Wang, Xuefeng Huang, Junji Cheng, Liyuanjun Lai
cs.AI
Abstract
Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational cost. We instead shift the focus to a useful initialization that can be refined into a high-quality solution with limited downstream refinement. We propose Just Initialize, a training-free and solver-agnostic initialization component for large-scale routing optimization. Just Initialize compresses a large routing instance into a compact surrogate space, optimizes its global routing structure, and recovers the resulting solution as an optimization-friendly starting point in the original space. Extensive experiments on Traveling Salesman Problems (TSPs), Capacitated Vehicle Routing Problems (CVRPs), Vehicle Routing Problems with Time Windows (VRPTWs), and Prize-Collecting Traveling Salesman Problems (PCTSPs) demonstrate that Just Initialize achieves high-quality solutions comparable to or better than state-of-the-art methods while substantially reducing computational cost across instances ranging from 1K to 100K nodes, including an average speedup of approximately 70$\times$, sub-second runtimes on 10K-node instances, and runtimes within tens of seconds on 100K-node instances.
cs.AI / 131 / 2609.35456
AutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain--Computer Interfaces
Muyun Jiang, Yi Ding, Wei Zhang, Jinbo Chen, Chenyu Liu, Zhenjie Yang, Yuxin Li, Jingyuan Chen, Yuhao Lu, Yong Li, Shuailei Zhang, Cuntai Guan
cs.AI
Abstract
EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection of EEG decoding architectures across tasks. The Designer Agent performs Pool-Guided Architecture Discovery (PGAD), generating and refining architectures through training and validation across multiple EEG tasks, such as emotion recognition, motor imagery, and sleep staging. The Forecaster Agent performs Performance Estimation from Early Knowledge (PEEK), using architecture code, the training protocol, and early learning curves to predict full-budget validation performance and select promising candidates for continued training. Across 14 EEG datasets spanning motor imagery, emotion recognition, and sleep staging, we evaluate AutoBCI with six LLMs, including Opus 5.5 and GPT 5.6 Sol, and compare the architectures selected by the search procedure against ten baselines: six conventional EEG models and four foundation models. The architecture discovered by AutoBCI with Claude Opus 5.5 achieves 64.16% average test balanced accuracy (bAcc), compared with 63.87% for REVE, the strongest baseline on this metric. Using ten observed epochs, PEEK reduces mean absolute error in predicting average validation bAcc from 2.20 to 1.36 percentage points, a 38.1% reduction relative to the best-observed-score baseline.
cs.AI / 132 / 2609.35463
A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans
Alessandro Emmanuel Pecora, Stefano Calzolari, Francesco Strada, Andrea Bottino
cs.AI · cs.GR
Abstract
Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify perception, reasoning, and action within a single control loop. A.D.A.M.O. maintains a dual visual-symbolic world model that combines egocentric visual input and synchronized symbolic state to support grounded task-oriented behavior from natural language prompts. To support diagnostic evaluation, we introduce a controlled task suite organized by a cd taxonomy that breaks down spatial tasks into procedural and linguistic complexity. Experiments in controlled scenes show that semantic labeling strongly influences task completion and failure modes, reducing perceptual ambiguity while shifting failures toward downstream execution, whereas reasoning errors remain comparatively rare.
cs.AI / 133 / 2609.35515
MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding, Yong Li
cs.AI · cs.LG · cs.SC
Abstract
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.
cs.AI / 134 / 2609.35532
ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
Kun Feng, Yuchen Fang, Yiyang Tan, Shuqi Gu, Yongxiang Zhao, Yu Liu, Xingyu Lu, Lintao Ma, Kan Ren
cs.AI
Abstract
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .
cs.AI / 135 / 2609.35549
RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong, Qingkai Zeng, Lijun Wang, Tien-Yin Wong, Peng Cui, Tianyu Liu
cs.AI
Abstract
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.
cs.AI / 136 / 2609.35551
BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
Peilin Feng, Zhengyang Huang, Soujanya Poria
cs.AI
Abstract
In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.
cs.AI / 137 / 2609.35561
RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement
Yaxin Du, Xiyuan Yang, Zhifan Zhou, Yujie Ge, Cheng Wang, Jiajun Wang, Sijie Chen, Zehui Liu, Yuxin Zhang, Weicheng Gu, Julian Zhang, Zixing Lei, Siheng Chen
cs.AI
Abstract
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0\% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.
cs.AI / 138 / 2609.35571
Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
Hyunwoo Yoo, Cassie Huang, Haebin Shin, Li Zhang, Gail L. Rosen
cs.AI · cs.CL
Abstract
While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL generation, we achieve high success rates where end-to-end approaches fail. This provides evidence that a primary bottleneck lies in representation alignment rather than raw model capacity. Our structural analysis demonstrates that intermediate representations are essential in retrosynthesis planning, highlighting the importance of representation-centric design in future systems.
cs.AI / 139 / 2609.35588
Source-preserving alignment for robust evidence localization in scientific PDFS
Zihao Liu, Wei Yang, Zixiao Dong, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
cs.AI
Abstract
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, and fragmented items alter text sequences and geometry. We present a source-preserving alignment framework: normalize text for robust matching while preserving provenance for accurate localization. It aligns evidence with normalized page text, maps matches back to source-character spans, and renders only their geometry. When exact alignment fails, line-break-aware token alignment recovers supported spans while excluding unmatched noise. Experiments on 1,020 chemistry papers show that the framework achieves a 92.6\% quote-level automatic localization rate, compared with 43.6\% for text search and 19.1\% for a precomputed bounding-box baseline. Component ablation confirms distinct contributions from normalization and approximate token alignment, while human verification assesses the visual correctness of returned highlights. Overall, these results demonstrate that reliable evidence verification requires robust matching and precise localization within a shared source-preserving alignment representation.
cs.AI / 140 / 2609.35618
From cacophony to hierarchy: a principled framework for assessing AI consciousness
Shamil Chandaria, Arvo Muñoz Morán, Fernando Rosas, Anil Seth, Henry Shevlin, Marcus Hutter, Thore Graepel, Adam Bales, Iulia Comsa, Murray Shanahan, Ruben Laukkonen, Morten Kringelbach, Chris Frith, Shane Legg
cs.AI · cs.CY
Abstract
The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a cacophony of competing theories that often talk past each other. Separating the hard problem from the mapping problem allows the deepest metaphysical disagreements to be set aside: granting that experience supervenes on a system's organisation, the tractable question becomes at which grain of description that supervenience base sits. We extend Marr's three levels of analysis into a five-level hierarchy of functional descriptions (behavioural, computational, intrinsic causal-structural, organismic, and organism-environment) grounded in supervenience, coarse-graining, and multiple realisability. The major theories of consciousness are positioned within this hierarchy according to which level they take to be critical, and for each level we develop operationalisable indicators and assess current AI systems against them. A Bayesian model then combines theoretical credences with indicator evidence into an overall credence in a system's capacity for consciousness. In illustrative assessments, the verdict for current LLMs is driven as much by where theoretical credence is placed as by how the evidence is read: under different stipulated readings and credence distributions, assessments range from below 0.01 to roughly 0.8, showing sensitivity to assumptions. Finally, the consciousness indicators at each level closely overlap with the architectural features needed for general intelligence, suggesting that increasingly capable AI may become a stronger candidate for consciousness. The framework supports a structured agnosticism, in which theoretical commitments are made explicit, credences are updated as evidence accumulates, and assessments take the form of aggregated probabilities rather than verdicts.
cs.AI / 141 / 2609.35671
PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
Yangqin Jiang, Lingrui Xu, Chao Huang
cs.AI
Abstract
Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld's official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app's navigation rather than one run, so it serves new tasks, not only repeated ones.
cs.AI / 142 / 2609.35677
Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
Christian Moya, Elliott Thornley, Guang Lin
cs.AI
Abstract
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.
cs.AI / 143 / 2609.35692
Report: Progressive Disclosure of Agent Skills
Guilin Zhang, Kai Zhao, Priyanka Mudgal, Waleed Ammar, Xiquan Cui, Xu Chu, Alet Blanken
cs.AI
Abstract
Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost. Progressive disclosure (lazy-loading) of skills as needed may reduce operational costs, but its impact on overall latency and skill-retrieval quality remains unclear. In this report, we investigate the impact empirically and find that progressive disclosure improves skill-retrieval quality but marginally degrades overall latency.
cs.AI / 144 / 2609.35732
Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
Junru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding, Chunxin Tang, Ruoyu Qi, Yulang Fei
cs.AI
Abstract
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.
cs.AI / 145 / 2609.35741
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim
cs.AI · cs.CL
Abstract
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
cs.AI / 146 / 2609.35744
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia Geng, Andrew Chin, Yin Luo, Edward Tong, Yu Yu, Zach Golkhou, Minkyu Kim, Igor Halperin, Young Cha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee
cs.AI · q-fin.CP
Abstract
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric's expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts' key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.
cs.AI / 147 / 2609.33818
Augmenting Visual Anomaly Detection with Automated Interpretability
Antonio De Santis, Arsenio Leo, Marco Brambilla
cs.CV · cs.AI
Abstract
Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different components of this signal. We decompose PatchCore nearest-normal residuals into sparse features using Sparse Autoencoders (SAEs), and provide high-activation and contrastive non-active examples to a Multimodal LLM, which describes each feature and labels it as anomaly, distractor, or uncertain. These labels guide interventions in the SAE hidden representation, where distractor features are suppressed and anomaly features amplified. The edited representation is then used to reconstruct patch embeddings, which are rescored with PatchCore. Across 40 categories from four benchmarks, applying both interventions jointly improves macro-average image-level AUROC from 0.8724 to 0.8857 on source data and from 0.8066 to 0.8210 under synthetic corruptions. On three additional RobustAD categories with real acquisition shifts, the same interventions improve AUROC from 0.8745 to 0.9056 on source data and from 0.6069 to 0.6599 under real acquisition shifts. Finally, individual feature interventions across all 43 categories show that the MLLM labels are aligned in aggregate with how features differently affect normal and anomalous images.
cs.AI / 148 / 2609.33855
Program-Verified Self-Evolution for Vision-Language Models
Ahmed Heakl, Sungik Choi, Moontae Lee, Salman Khan
cs.CV · cs.AI · cs.CL · cs.LG · cs.NE
Abstract
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94\% of VQS answers correct, against 76\% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS
cs.AI / 149 / 2609.33935
Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy
Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia
cs.CV · cs.AI
Abstract
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.
cs.AI / 150 / 2609.34047
ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
Kieran Sagar Parikh, Jose Luis Garcia del Castillo y Lopez
cs.CV · cs.AI
Abstract
Multimodal models increasingly interpret visual environments, but their ability to recognize the same building across photographs, floor plans, elevations, sections, and renderings remains poorly characterized. We introduce ARCH-B, a benchmark of 354 four-choice questions across 11 cross-representational archetypes, constructed from a building-linked corpus of 3.9 million architectural images using visually similar distractors, model-guided difficulty screening, and manual validation. We evaluate 25 multimodal models and collect 5,830 responses from non-expert human participants. Model accuracy ranges from 10.45% to 83.90%, compared with a human baseline of 35.35%. Models perform comparatively well on mixed-representation outlier detection and photograph matching, but remain weaker on floorplan-to-photograph correspondence. Human and model difficulty across archetypes is only weakly correlated (Spearman's (ρ=0.33)). Held-out evaluation confirms that the difficulty identified during screening generalizes beyond the curation models. ARCH-B provides a diagnostic evaluation of visual correspondence and representation transfer across architectural media.
cs.AI / 151 / 2609.34078
WhiteCon: Semi-Supervised Domain Adaptation Regression Through Whitening Transform and Dual Consistency
Se Jin Sim, Seoung Bum Kim
cs.CV · cs.AI
Abstract
Domain adaptation is crucial for addressing distributional shifts that degrade model performance across domains. While most existing research has centered on classification, semi-supervised domain adaptation regression (SSDAR) for continuous-output tasks remains largely unexplored, particularly in practical scenarios with limited labeled target data. To address this gap, we propose semi-supervised domain adaptation regression through whitening transform and dual consistency (WhiteCon), which combines domain-specific whitening transform (DWT) and dual consistency regularization to enhance training stability and domain adaptation. DWT reduces the variance of the model parameters by transforming the feature covariance matrix into an identity matrix, thus stabilizing training under ordinary least squares assumptions. In addition, variance consistency regularization, as part of dual consistency regularization, aligns the variances of weak, strong, and mixup-augmented features to improve resilience against augmentation-induced perturbations. Empirical evaluations on various benchmark datasets under SSDAR settings demonstrate that the proposed WhiteCon achieves state-of-the-art performance compared to existing methods, effectively addressing domain shifts in regression tasks. The code for WhiteCon is available at https://github.com/sejin-sim/WhiteCon.
cs.AI / 152 / 2609.34106
The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation
Kuniaki Saito, Yoshitaka Ushiku
cs.CV · cs.AI
Abstract
Large visual foundation models have demonstrated remarkable transferability across a wide range of downstream tasks. To deploy such models efficiently, feature matching has become a popular knowledge distillation approach that transfers teacher representations to smaller student models without requiring labeled data. However, we show that the conventional feature matching objective with L2-distance is inherently biased toward reconstructing dominant spectral directions of the teacher representation, while under-optimizing low-variance directions that often contain task-relevant information. To address this, we propose Spectrum-Balanced Feature Matching, SpecMatch, a simple objective that adaptively emphasizes under-optimized spectral directions while preserving the relative importance of dominant directions. SpecMatch is easy to implement and introduces negligible computational overhead. Extensive experiments on image recognition demonstrate that SpecMatch consistently improves downstream adaptation across diverse tasks, including image classification, anomaly detection, medical image analysis, and domain generalization. In particular, SpecMatch outperforms conventional feature matching in 40 of 42 teacher--student and training-setting combinations, while consistently improving over the original student model in all settings. We further demonstrate that the proposed objective generalizes beyond vision, improving downstream performance across six protein understanding tasks.
cs.AI / 153 / 2609.34133
PrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise Preferences
Chuanzhi Xu, Langyi Chen, Chengkun Yue, Xuanhua Yin, Boyu Wei, Qingwen Zeng, Zihan Deng, Weidong Cai
cs.CV · cs.AI
Abstract
Photographic color editing is inherently personal: the same image can appear too warm, too muted, or already satisfactory to different users. Most lookup table (LUT) and reference-guided methods target a specified appearance rather than model persistent preferences from repeated user choices. To address this gap, we introduce PrefLUT, a reusable and refinable user-preference modeling framework for deployable 3D LUTs, encoding ordered preferred/non-preferred image pairs into a lightweight Reusable User Profile that is reused across queries and refined using additional user preference pairs, without per-user optimization. A Query-Conditioned LUT Predictor combines this profile with each image to predict a LUT latent vector and edit strength. An Identity-Residual LUT Decoder and Edit-Strength Controller then produce an exportable 3D LUT. Experiments on three datasets demonstrate effective personalized editing and general-purpose enhancement. Each quantized profile requires only 260 bytes, and editing takes 1.365 ms/image on an RTX 5090 GPU. We also introduce the Preference-Conditioning Verification Protocol (PCVP), an evaluation protocol to verify whether personalized image edits depend on user preferences and the query image through controlled changes to user profiles, preference orders, pair correspondences, and query images.
cs.AI / 154 / 2609.34277
See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology
Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li, Xinyu Liu, Jiaming Yang, Jie Chen, Zhang Zhang, Yuhao Yi, Hong Bu, Jiancheng Lv
cs.CV · cs.AI
Abstract
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.
cs.AI / 155 / 2609.34314
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
Shayekh Bin Islam, Hwanjun Song
cs.CV · cs.AI · cs.CL
Abstract
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.
cs.AI / 156 / 2609.34335
SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering
Yanwei Huang, Mingxuan Zhu, Shujie Li, Shiyuan Liu, Yuanxing Zhang, Arpit Narechania
cs.CV · cs.AI
Abstract
Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents shot logic, composition, lighting, sound design, and other filmmaking cues in a fine-grained format, and retrieves movie references categorized as resonators (good matches), dissonants (weak matches), and divergents (creatively useful near-misses). The first two refine when and how a skill should be applied, while divergents inspire alternative cinematic realizations at different degrees of modification while preserving the user intent. Candidate skills are assessed through generated videos along prompt fidelity, cinematic quality, narrative appeal, and creativity to construct the final skill libraries. Experiments on StoryEval and VBench show improvements of up to 1.40 points over the strongest external baseline and 0.51 points over seed skills on 7-point four-dimensional evaluation, while remaining competitive on benchmark-native metrics. Overall, SkillPE offers a practical approach to balancing fidelity and creativity in cinematic text-to-video generation. Code is available at https://github.com/Ais0n/SkillPE .
cs.AI / 157 / 2609.34363
SyncRA: Learning Temporal Correspondence in Omni-Modal Models
Zelong Xu, Yan Li, Wenhe Hu, Xiyang Hu
cs.CV · cs.AI · cs.MM
Abstract
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.
cs.AI / 158 / 2609.34470
Precise Editing and Flexible Referencing for Interactable Worlds
Xinyao Liao, Xianfang Zeng, Zhu Liang, Zhoujie Fu, Qianxun Xu, Jiachi Liu, Gang Yu, Guosheng Lin
cs.CV · cs.AI
Abstract
We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends world modeling from exploration to precise modification by streaming editing instructions and reference images during autoregressive generation. To support these capabilities, EditWorld introduces Gated Causal Attention for temporally varying editing conditions and reference images, together with a Sparse Context mechanism that maintains a bounded historical context for long-horizon inference. We further adopt joint autoregressive and bidirectional training with annealed self-resampling, and construct a dedicated data synthesis and annotation pipeline that provides supervision for world editing. We also present WBench-Editing to systematically evaluate streaming world editing capabilities. EditWorld achieves the best overall performance on WBench-Editing with an overall score of 73.8 and an editing score of 80.0, substantially outperforming existing methods on editing-related metrics. https://github.com/leoisufa/EditWorld
cs.AI / 159 / 2609.34472
VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation
Sukju Oh, Soo Ick Cho, Dae Hun Suh, Sukkyu Sun
cs.CV · cs.AI
Abstract
Acne assessment is crucial for clinical decision-making, yet traditional grading and counting are subjective and fail to account for lesion size. While area-based assessment has emerged as a promising alternative, acne segmentation has continued to rely on general-purpose architectures. To address this gap, we propose VL-AcneSeg, a multimodal framework for acne lesion segmentation that leverages CLIP and region-level text prompts to incorporate spatial priors, enabling lesions to be localized across the whole face. Because region-level prompts indicate which facial areas contain lesions, we report a single global prompt, which requires no such information, as our primary setting. On our internal clinical dataset, VL-AcneSeg achieves a Dice score of 0.5082 and an IoU of 0.3407 under this protocol, the highest among all compared methods, including recent vision-language segmentation methods that are themselves given region-level prompts; region-level prompting raises these to 0.5296 and 0.3602. Moreover, lesion area measurements derived from our segmentation correlate with IGA scores at a level comparable to expert annotations (Pearson r = 0.719 versus 0.658). Notably, our framework maintains consistent performance across external validation datasets, performing reliably even on uncontrolled smartphone images without requiring additional training or fine-tuning. By pairing a protocol that requires no lesion-location information with area-based severity estimation, this work provides a foundation for objective acne assessment outside the clinic. Our implementation is publicly available at: https://github.com/sukjuoh/VL-AcneSeg
cs.AI / 160 / 2609.34547
ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy, Hung-Ting Su, Winston H. Hsu
cs.CV · cs.AI · cs.CL
Abstract
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276
cs.AI / 161 / 2609.34587
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
cs.CV · cs.AI
Abstract
Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that is poorly aligned with the contribution of individual tokens. A generated program may contain operations that accurately reproduce some parts of the target image alongside others that introduce errors, yet all tokens are trained from the same final outcome. We observe that many intermediate code prefixes are not only executable, but already produce meaningful partial renders that reflect progress toward the target. This property provides a natural source of denser supervision during generation. Based on this observation, we introduce IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. We evaluate our approach on Image-to-SVG and Image-to-TikZ generation. Across both tasks, our method improves over supervised fine-tuning and standard GRPO, yielding new state-of-the-art open-source models. This shows that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code models.
cs.AI / 162 / 2609.34598
Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding
Nanxing Hu, Xiaoyue Duan, Qiwei Yan, Kailin Lyu, Jinchao Zhang, Guoliang Kang
cs.CV · cs.AI
Abstract
Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a ``summarize before grounding'' framework (named ``SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.
cs.AI / 163 / 2609.34672
Evidence-Aligned Multimodal On-Policy Self-Distillation for Fine-Grained Visual Understanding
Nanxing Hu, Qiwei Yan, Jinchao Zhang, Guoliang Kang
cs.CV · cs.AI
Abstract
Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories. Ideally, teacher corrections, the distributional changes from the student toward the privileged teacher, should be driven by task-relevant visual evidence. However, the designs that make the teacher effective also introduce other interference. Using a lagged or frozen teacher improves training stability but introduces a model-state gap from the evolving student, while cropping enhances task-relevant evidence but also loses the visual context. These two sources of interference make the teacher corrections not purely rely on the visual evidence. We introduce Evidence-Aligned multimodal on-policy self-Distillation (EAD), which retains the crop-conditioned teacher as the target but constructs a separate evidence reference for weighting the corrections. To exclude the effect of lagged model-state from this reference, EAD measures prediction changes using the current student. To avoid crop-induced context changes, EAD masks the evidence region in the original image while preserving the other visual context. The change from the student's masked-image prediction to its original-image prediction provides a controlled reference for the direction in which the visual evidence shifts the student's prediction. EAD weights each teacher correction by its cosine alignment with the reference, i.e., retaining aligned corrections and downweighting the rest. Retaining only 6\% of the supervision mass of dense OPSD, EAD consistently outperforms previous state-of-the-art methods.
cs.AI / 164 / 2609.34749
CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving
Yu Meng, Baining Zhao, Junta Wu, Tengfei Wang, Rongze Tang, Haiyu Zhang, Wenqiang Sun, Chen Gao, Zhibo Chen, Xinlei Chen, Yong Li, Xiao-Ping Zhang, Chunchao Guo
cs.CV · cs.AI
Abstract
Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video generation framework that jointly generates observations of vehicles sharing the same dynamic scene with precise camera-trajectory control. CoDrive interleaves local self-attention, which models spatiotemporal dependencies among the views of each vehicle, with global self-attention, which enables information exchange and consistency modeling across vehicles. To explicitly encode their spatial relationships, all camera trajectories are represented in a shared world coordinate system and injected into the attention layers through projective relative positional encoding. We further adopt a progressive mixed-task training strategy that combines large-scale real-world single-agent data with synthetic cross-agent interaction data, allowing the model to benefit from real-world appearance distributions while learning cross-agent consistency from simulation. For systematic evaluation, we introduce CoDrive-Bench, a benchmark covering real and synthetic multi-vehicle scenarios and evaluating trajectory controllability, scene geometry consistency, and instance-level consistency. Experiments show that CoDrive improves trajectory controllability and cross-agent geometric and instance consistency while maintaining competitive visual quality.
cs.AI / 165 / 2609.34781
When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
Yuxing Cheng, Yuan Wu, Yi Chang
cs.CV · cs.AI
Abstract
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45\% to 58.51\%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
cs.AI / 166 / 2609.34817
ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild
Hongyu Ma, Hairong Qu, Shiqi Zhao, Yongsong Yang, Peng Yin
cs.CV · cs.AI
Abstract
Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-end model nor an in-the-wild benchmark. We propose ESTHER, a model whose stereo geometry, temporal reasoning, and output representation are designed for wearable egocentric stereo. It is trained on pseudo-labels from a calibrated labeling pipeline and in turn assembles our benchmark ESTHER3D, an egocentric stereo hand dataset pairing a large in-the-wild training set of model-generated labels with a motion capture test set of true metric ground truth. Experiments show state-of-the-art accu?racy, superior external generalization, and robustness to the missing views, dropped frames, and lighting and motion blur extremes of real egocentric capture that break existing meth?ods. This robustness runs deeper than graceful degradation: stereo guidance teaches the model to bind apparent hand scale to metric depth, so it not only adapts to different stereo rigs and modalities with minimal fine-tuning, but more strikingly preserves true metric scale even after collapsing to a single monocular view.
cs.AI / 167 / 2609.34853
EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation
Sungho Moon, Kota Shimomura, Junwoo Park, Wonhyeok Choi, Seunghun Lee, Takayoshi Yamashita, Sunghoon Im
cs.CV · cs.AI
Abstract
Open-vocabulary 3D scene understanding enables object localization and segmentation from free-form text queries without a fixed category vocabulary. Many recent methods build on 3D Gaussian Splatting and consolidate multi-view observations, such as masked crops from individual views, into language features or compact object descriptors before the query is known. However, observations of the same object vary across viewpoints and are not equally informative: some reveal cues relevant to a particular query, whereas others provide incomplete or misleading evidence. Pre-query consolidation can therefore suppress cues on which a later query depends. We introduce EviSplat, which preserves individual observation features as evidence for later text queries. EviSplat retains individual observation features within class-agnostic 3D instances that represent objects, object parts, or background regions. It also learns, for each Gaussian, a distribution describing which visual appearances its observations support. Given a text query, EviSplat scores each instance using its most relevant observations. It then computes a score for each Gaussian by combining instance-level relevance with locally supported evidence, weighted by how often and how unambiguously that Gaussian was observed. Different queries can thus draw on different visual cues from the same preserved evidence. Experiments across diverse datasets and evaluation protocols demonstrate state-of-the-art performance, supporting the benefit of preserving multi-view evidence until query time and aggregating it according to the query.
cs.AI / 168 / 2609.35002
Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models
Qiankun Li, Yuechen Zhang, Bowen Chen, Shilinlu Yan, Zhenhong Zhou, Kun Wang, Li Sun
cs.CV · cs.AI · cs.CR
Abstract
Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial input that remains correct under full-token inference but fails after compression, casting compression-induced risk as a paired failure attribution problem. Within a controlled diagnostic cohort, counterfactuals show that retained-set allocation causally changes compressed correctness and reveal a negative association between recovery and representation drift in displaced evidence. Motivated by these findings, we propose CIRA, a Compression-Induced Risk Attack for Large Vision-Language Models. Under a vision-encoder white-box setting, CIRA optimizes image perturbations through encoder-side objectives that manipulate token priorities across candidate compression budgets while preserving displaced evidence. CIRA uses no downstream questions or labels and requires no access to the language model, deployed compressor, or exact compression budget. Across 12 dataset-compressor settings evaluated at four budgets, CIRA achieves a mean CSFR of 20.35% while limiting full-token attack success to 6.92%, with similar behavior on additional LVLM families. A cross-view selection-stabilization defense substantially suppresses CIRA, although Adaptive CIRA partially restores its effectiveness. These results show that compression-specific failures persist under restricted access and support paired evaluation of full-token and compressed inference for attributing risk to visual-token compression.
cs.AI / 169 / 2609.35052
OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models
Hao Wang, Tao Yu, Liuzhou Zhang, HeXin Wang, Haopeng Jin, Yuxuan Zhou, Xinming Wang, Hongzhu Yi, Xinye Li, Yuanlei Wang, Ping Nie, Yan Huang, Yuxuan Zhang, Pengfei Zhou, Yanyan Zou, Wei Yang
cs.CV · cs.AI
Abstract
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded benchmark that strictly anchors the assessment to a fixed set of object instances from the initial observation for evaluating multi-object memory in video world models. The OPIS dataset comprises 500 cases across real-world, embodied-robotic, and game-world domains, providing dense object-level annotations for 12,672 rigid, articulated, and deformable instances. Our object-centric evaluator combines association and explicit visibility reasoning to hierarchically measure Object (O) Presence (P), Identity (I), and Structure (S), utilizing static or dynamic evaluation tracks based on object kinematics. Across eight image-to-video or camera-conditioned world models, our proposed OPIS scores range from 48.65 to 56.01. As the reference inventory grows from less than 20 to more than 40 objects, the Presence, Identity, and Structure scores show an overall decline, with the average Identity score falling from 40.22 to 23.11. The results demonstrate that preserving the particular object instances in the input is considerably harder than generating plausible visual elements.
cs.AI / 170 / 2609.35143
Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
Gunin Gupta, Nirmit Arora, Pavan Kalyan Tankala
cs.CV · cs.AI · cs.MM
Abstract
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.
cs.AI / 171 / 2609.35195
CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation
Md Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans, Walia Farzana, Asfaqur Rahman, Ahmed Temtam, Khan M. Iftekharuddin
cs.CV · cs.AI
Abstract
Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically distinct tumor subregions: non-enhancing tumor core (NETC), surrounding non-enhancing FLAIR hyperintensity (SNFH), enhancing tumor (ET), and the resection cavity (RC). Among these, RC segmentation is particularly challenging because of its low prevalence, heterogeneous postoperative appearance, and lesion-wise evaluation protocol, leading conventional segmentation networks to prioritize dominant tumor classes during optimization. The proposed nnU-Net-based framework explicitly addresses RC segmentation through four complementary components: (i) RC-weighted Dice and Cross-Entropy optimization to alleviate class imbalance, (ii) anatomically consistent cavity augmentation to increase the diversity of postoperative cavity appearances, (iii) a residual encoder architecture for enhanced multi-scale feature learning, and (iv) lesion-aware morphological post-processing to suppress false-positive cavity predictions while preserving anatomically plausible structures. The framework is evaluated on the BraTS-MET 2026 Task 1 online validation benchmark. Among the evaluated configurations, the ensemble model (Residual Encoder nnU-Net + nnU-Net + RC-aware CarveMix) achieves the best performance, with lesion-wise Dice scores of 0.732, 0.752, 0.708, and 0.575 and corresponding NSD scores of 0.794, 0.798, 0.727, and 0.474 for ET, TC, WT, and RC, respectively. These experimental results show that integrating RC-aware optimization, anatomically consistent augmentation, and lesion-aware post-processing provides an effective strategy for improving rare resection cavity segmentation in post-treatment brain metastases.
cs.AI / 172 / 2609.35410
Spectral Super-Resolution using Spatial-Spectral Residual Operator Networks
Seokhyun Chin
cs.CV · cs.AI
Abstract
Spectral super-resolution of multispectral satellite images can enable high temporal- and spatial-resolution hyperspectral satellite imagery at a modest cost, significantly increasing the applicability of hyperspectral remote sensing. This task is inherently ill-posed, making it well-suited for deep learning-based methods. In this study, the spectral super-resolution task is framed as an operator learning problem, and SSRON is proposed as a Deep Operator Network that effectively learns function-to-function mappings from downsampled spectra to continuous spectra. The model is trained to super-resolve Sentinel-2A-like multispectral imagery to EMIT images. Compared to baseline models, SSRON achieves superior performance across all metrics. The model also demonstrates zero-shot spectral super-resolution capability by predicting bands unseen during training. Furthermore, its continuous-output formulation suggests the potential to estimate spectra at finer wavelength intervals than the native sensor. These results suggest the potential of SSRON and establishes operator learning as a promising direction for spectral super-resolution.
cs.AI / 173 / 2609.35504
SolveEdit: Benchmarking Visual Problem Solving in Generative Models
Wenjie Shu, Yexin Liu, Harold Haodong Chen, Xuerui Qiu, Zehan Wang, Yidi Zhang, Yizhan Chen, Zunwei Wang, Minghao Liu, Qi Chen, Harry Yang, Xiaogang Xu
cs.CV · cs.AI
Abstract
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
cs.AI / 174 / 2609.35530
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
Yuta Oshima, Ku Onoda, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta
cs.CV · cs.AI
Abstract
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.
cs.AI / 175 / 2609.35745
Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, Wu Liu, Xi Peng, Chun Jian Ho, Hongyuan Zhu
cs.CV · cs.AI · cs.LG
Abstract
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.
cs.AI / 176 / 2609.35767
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu
cs.CV · cs.AI
Abstract
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
cs.AI / 177 / 2609.35770
FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets
Srinjay Sarkar, Prakhar Kaushik, Soumava Paul, Alan Yuille
cs.CV · cs.AI · cs.GR
Abstract
Realistic and editable animal fur reconstruction from multi-view images is challenging due to fine-scale detail, self-occlusion and obfuscation, and, unlike human hair, the lack of animal-fur datasets. Fur usually covers most of an animal's body, with large inter-species and intra-species variability. We present FurE, an efficient strand-based animal fur reconstruction method that recovers a per-strand, editable groom by optimizing a root-conditioned latent field, decoded into strand geometry via a PCA-based decoder. We reconstruct a defurred animal body using local fur-thickness cues from a surface-constrained Gaussian Frosting representation together with part-based priors. We further show that a PCA-based decoder learned from human-hair strand data can alleviate animal-data scarcity while enabling substantially faster optimization. FurE achieves a 10x speedup in strand training over current SOTA dense per-strand optimization while retaining strand fidelity and generalizing across synthetic and real-world sequences, with quantitative and qualitative validation despite the reduction in training time.
cs.AI / 178 / 2609.34258
Investigating Human--AI Discrepancies via Multiple-Solution Problems
Zihao Wang, Francesco Insulla, Andrea Montanari
cs.CY · cs.AI
Abstract
Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and model reasoning lead to different distributions over valid solutions. Our testbed comprises 270 reasoning puzzles across five puzzle families. These multiple-solution puzzles each have 3 to 8 valid solutions and are simple enough that humans and models can solve them reliably. The resulting distributions differ markedly: models differ from one another, yet resemble each other far more than they resemble humans. Model distributions are, moreover, within every puzzle family, less diverse than human ones. We compare these discrepancies across puzzle categories, and trace how they respond to reasoning-effort settings, to prompting, and to perturbations of the puzzle that leave its solutions unchanged. Together, these results point at significant differences between human and AI problem-solving processes, and their choice among equally defensible solutions. As progressive deployment of AI systems in society comes into focus, evaluating such differences (beyond one-dimensional accuracy metrics) is increasingly important. Data and code are available at https://hai-discrepancies.github.io/
cs.AI / 179 / 2609.33916
Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters
Lenore M. Mullin, Gaetan Hains
cs.DC · cs.AI
Abstract
We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forward+backward, decode, and the complete block (RMSNorm, gated MLP) as a hardware-independent specification (DNF) transformed to a machine-specific realization (ONF) via gamma, with verification to machine precision against PyTorch. This paper checks those predictions against measured performance on two HPC clusters (Purdue Anvil, NCSA Delta) across CPU and GPU. Three results stand out. (1) We identify and fix a GPU regression: fusing forward+backward, proven to avoid materializing an O(n^2) intermediate, initially ran slower than naive on GPU due to atomic contention. Profiling confirmed 2.00x more atomic instructions; a targeted ONF rewrite reversed it, yielding up to 2.5x speedup. (2) Identical derivations produce markedly different real costs by topology: 535x NUMA-locality penalty on one cluster vs <3x oversubscription on another, showing optimal deployment is a function of the machine's array structure. (3) We report a partially resolved anomaly: identical denotational computations run faster in C than Fortran on CPU but faster in Fortran than C on GPU, narrowed to one dominant kernel and one memory-latency stall mechanism (3.17x time gap matches 3.35x stall gap). We treat hardware-specific optimization as a routine ONF rewrite with fixed, verified DNF, a candidate methodology for scaling AI onto evolving hardware without re-deriving correctness.
cs.AI / 180 / 2609.34376
Before Agents Act: Assurance-Aware Semantic Scheduling for Evidence Acquisition in Distributed Systems
Jun He, Deying Yu
cs.DC · cs.AI · cs.CR
Abstract
Tool-using agents can initiate consequential infrastructure changes, yet evidence required for admission may expire while other checks run or depend on a shared fault domain. We formulate evidence acquisition as joint witness selection and scheduling under quorum, diversity, freshness, deadline, and resource constraints. Assurance-Aware Semantic Scheduling (AAS) combines integer-program selection, dispatch-aware temporal scheduling, bounded diagnostic expansion, and receipt-aware repair. Formal results state the assumptions needed for dispatch-time freshness and finite diagnostic expansion. In three generated infrastructure workloads, AAS produces 1,075/1,200 valid candidates versus 647/1,200 for constraint-aware forward scheduling; stale candidates fall from 440 to 12. Paired sensitivity studies reuse the same instances and operation latency draws across parameter settings. A corrected timeout intervention finds 18/20 admissions with repair or full resynthesis versus 0/20 for a static plan, with lower committed cost when receipts are reused. On 20 constructed cases requiring a certified decomposition cut, refinement recovers an oracle-matching feasible plan every time. These are controlled simulation results; the bounded oracle shares a temporal search component, and transfer to deployed systems remains untested.
cs.AI / 181 / 2609.34380
DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory
Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung, Thanh Tuan Dao
cs.DC · cs.AI · cs.ET · cs.PF · cs.SE
Abstract
Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this design by allowing a single stored model to support both full-accuracy and lower-precision execution, making the effective weight footprint runtime-dependent. This creates an opportunity under bursty workloads, where temporary spikes in KV-cache demand often determine throughput and SLO compliance. We present DPS, a dual-precision LLM serving system that turns weight memory into an elastic resource: under normal load, DPS serves the full-accuracy model; under KV pressure, it switches to a nested, lower-precision variant and repurposes unused weight memory for KV cache blocks. DPS is built on Semi-Unified Memory (SUM), which partitions the weight region into a persistent lower-precision sub-region and a shared region that alternates between residual weight tensors and KV-cache blocks, preserving compatibility with paged KV-cache management. We implement DPS on top of vLLM and evaluate it across both dense and MoE models and various production workload traces. Our results show that \sysname improves sustained throughput by $2.1$--$3.3\times$ and effective pass@1 by up to $+41$\,pp over Static FP16, while preserving FP16-class accuracy.
cs.AI / 182 / 2609.35263
WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse
Aaryam Sharma
cs.DC · cs.AI
Abstract
Pipeline parallelism can improve prefill throughput by processing multiple request chunks concurrently across different stages of the model. However, keeping the pipeline fully utilized requires efficient scheduling and request preparation. In systems where stages retain and evict cache state independently, a local cache hit does not guarantee that the same prefix can be reused across the pipeline. Here, coordination overhead can impede request admission cadence and thus reduce overall throughput. In this paper, we present WavePP, a prefill runtime built on top of TensorRT-LLM that addresses these challenges by overlapping request admission with pipeline execution. WavePP asynchronously finds a prefix that can be reused across all stages, protects the cached state, and reserves space for the remaining input while earlier requests continue to execute. It subsequently plans the chunk sizes of each request dynamically to maximize pipeline fill. Each stage then completes the local preparation before executing the request. In the same system and pipeline topology, WavePP improves TensorRT-LLM's prefill throughput in 37 of 40 tested settings on GLM 5.2 and MiniMax M2.7. At concurrency 128 with high cache reuse, these changes increase throughput by factors of 2.91 and 2.02, respectively. Across 28 Kimi K3 settings, WavePP also has the highest measured throughput in all 18 settings at concurrency eight or higher, compared with tensor/expert-parallel and pipeline-parallel baselines from TRT-LLM, SGLang, and vLLM.
cs.AI / 183 / 2609.35639
GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation
Yuchen Sun, Jinjin He, Sinan Wang, Bo Zhu
cs.DC · cs.AI
Abstract
Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks cover fluids, deformable solids, and granular materials, from individual simulation operators to complete simulators. Agents write, compile, test, and optimize GPU code with access to a NVIDIA GPU under fixed time budgets. We report pass rates and runtime performance relative to expert-optimized reference implementations. In a single-attempt evaluation of six frontier model-harness pairs, the two strongest pass all 50 tasks, but even the fastest reaches at least 0.9 the reference speed on only 22% of them, and no submission is more than 5% faster than the reference. The largest gaps arise in collision detection, constraint solving, and iterative solvers. GPUPhysBench brings physical simulation workloads to coding-agent evaluation, testing both the ability to implement numerical methods correctly and the ability to make them run efficiently.
cs.AI / 184 / 2609.34391
P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction
Haojie Yang, Ran Su
cs.ET · cs.AI · cs.LG
Abstract
AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never observed as a pair, and cells under one condition remain heterogeneous and noisy. Regression on individual cells absorbs sampling variation into the estimated effect, whereas interpretation requires the reproducible population effect. P2P (Perturbation-to-Perturbation) takes a stochastic cell-set view as its supervision unit. Two views drawn from the same condition share a reproducible population effect and differ by view-specific variation. A permutation-invariant set encoder summarizes the control population, a structured encoder represents perturbation tokens, cellular context, dose, and combination interactions, and a gate blends empirical condition-effect memory with a neural residual. A heteroscedastic head predicts the population mean and gene-wise response variance. Under one protocol and five seeds, P2P attains the lowest expression RMSE and the highest Effect Pearson, DEG F1, and DEG average precision on each of Adamson, Norman, Replogle K562, and Replogle RPE1 relative to GenePert, LinearPert, SLIM, Scouter, and scPILOT. On Replogle K562, Effect Pearson rises from 0.643 to 0.702 and DEG F1 rises from 0.067 to 0.178 relative to Scouter, the strongest baseline on both metrics.
cs.AI / 185 / 2609.35331
Reverse Sequential Proportional Approval Voting Rule: Proportionality and Approximation Guarantees
Georgios Papasotiropoulos
cs.GT · cs.AI
Abstract
We study the Reverse Sequential Proportional Approval Voting Rule (RevSeqPAV) in approval-based committee elections. Despite its historical prominence and practical use, its properties and guarantees are much less understood than those of Sequential PAV. We analyze it along two dimensions: proportional representation (measured by Extended Justified Representation, its approximations, and proportionality degree) and approximation of the maximum PAV score of instances. We first establish strong negative results for general, unrestricted election instances and then identify settings in which the rule provides meaningful fairness and optimization guarantees.
cs.AI / 186 / 2609.33962
EEG-Fusion: Failure-Informed Source-Free Expert Routing for Robust Motor Imagery EEG Decoding
Abdul Basit, Saim Rehman, Muhammad Shafique
cs.HC · cs.AI · cs.LG · eess.SP
Abstract
Subject-independent motor-imagery (MI) EEG decoding can exhibit subject-level failures even when average performance appears acceptable: under subject shift, a decoder can become an overconfident near-one-class predictor. This is especially problematic in source-free deployment, where target-user labels are unavailable during adaptation and expert selection. We present \textit{EEG-Fusion}, a failure-informed decision-level fusion framework that treats source-free MI decoding as label-free reliability estimation over heterogeneous experts. EEG-Fusion applies subject-wise Euclidean alignment and normalization-only test-time adaptation, then routes each target subject to a neural, covariance-based, or physiological-feature expert using a reliability gate trained on source-held-out folds to predict expert performance and collapse risk from label-free stream diagnostics. The gate uses confidence, entropy, prediction diversity, expert agreement, and predicted class balance; collapse is measured as the maximum predicted class fraction. In 9-fold leave-one-subject-out (LOSO) evaluation with three seeds, relative to a no-alignment raw EEGNet source-free anchor, EEG-Fusion improves subject macro-F1 from 0.417 to 0.529 on BCI IV-2a local protocol, from 0.314 to 0.482 on BNCI2014-001, and from 0.607 to 0.708 on BNCI2014-004; corresponding collapse-index reductions are 0.199, 0.227, and 0.169. In a 9-subject Cho2017 external subset, EEG-Fusion improves macro-F1 from 0.516 to 0.630. These results suggest that label-free reliability estimation can reduce subject-level failure modes in source-free MI-EEG deployment.
cs.AI / 187 / 2609.33967
ThinkNet: Compact Architecture Selection and Validation-Gated Ensembles for Subject-Independent MI-EEG Decoding
Abdul Basit, Saim Rehman, Muhammad Shafique
cs.HC · cs.AI · cs.LG · eess.SP
Abstract
Practical assistive and rehabilitative brain--computer interfaces require subject-independent motor-imagery EEG (MI-EEG) decoders that generalize to new users under limited target-user data and constrained compute. However, held-out-subject performance can be overstated when test-subject information influences preprocessing, model selection, or ensemble selection. We present \textit{ThinkNet}, a validation-controlled framework that combines train-only normalization, validation-guided evolutionary search, and validation-gated inference to identify compact decoders and inference policies for held-out subjects. We evaluate four-class BCI Competition IV-2a (session T) decoding with nine Leave-One-Subject-Out (LOSO) folds, three seeds, seven fixed decoder entries, and a broader search over ten representative decoder families; the held-out subject is never used for normalization, hyperparameter, architecture, or ensemble-policy selection. In the fixed benchmark, the validation-selected compact decoder achieved 44.35$\pm$15.41\% accuracy with 4.9K parameters, 19 KB FP32 weights, and 0.99 ms batch-1 Orin CUDA inference. Across the broader search, compact models ($\leq$25K parameters) achieved higher mean held-out accuracy than mid-size and large alternatives after selected retraining (40.10\% vs. 35.09\% and 34.78\%). Validation-gated ensembling improved over validation-selected single-model inference, reaching 43.98$\pm$16.25\% in the fixed benchmark and 43.31$\pm$15.88\% for the compact six-family ensemble. A non-deployable oracle analysis revealed a 6.1-point family-selection gap and near-zero validation--test correlation, showing that validation reliability remains a key bottleneck under subject shift. Thus, ThinkNet is a validation-controlled framework for compact MI-EEG model and inference-policy selection, rather than a single-architecture benchmark.
cs.AI / 188 / 2609.35197
Alignment Games: A Framework for Conceptual Repair in Human-AI Collaboration
Hari Subramonyam, Maneesh Agrawala, Sean Follmer
cs.HC · cs.AI
Abstract
The meaning of a concept in use is shaped by the situation, task, goals, and prior knowledge. For example, a request to make a poster "visually appealing for a five-year-old" might evoke bright colors and cartoon imagery for one collaborator, but less text, bold shapes, and visual simplicity for another. We call such task-relevant differences conceptual misalignment. We introduce Alignment Games, a framework for making these differences visible and repairable during human-AI interaction. Drawing on theories of situated conceptualization, we characterize task-specific conceptual frames in terms of relevant attributes, values, relations, constraints, and priorities. We then define alignment moves that intervene on the situation, the reasoning used to interpret it, or the resulting frame. Through examples from educational content generation, creative coding, and argumentative writing, we show how these moves can be composed into repair sequences and derive design principles for supporting task-sufficient conceptual alignment at runtime.
cs.AI / 189 / 2609.34127
STITCH-RAG: Spatio-Temporal Influence Tracing over Topic Hypergraphs for Multi-Hop Retrieval-Augmented Generation
Haodong Yang, Mengzhu Chen, Jia Cai
cs.IR · cs.AI
Abstract
Multi-hop retrieval-augmented generation requires a retriever to connect evidence distributed across documents while preserving a concise, faithful generation context. Existing indexes leave two complementary gaps: chunk-based RAG can break cross-passage evidence chains, whereas an unlabeled pairwise projection without generating-topic provenance cannot jointly preserve topic-level co-participation and per-occurrence entity descriptions. We propose STITCH-RAG, a hypergraph-based framework with three coupled components. First, a semi-merged topic hypergraph encodes multi-entity co-participation as topic-summary hyperedges while retaining per-chunk entity states linked by canonical-name equivalence. Second, spatio-temporal influence bridging propagation (STIBP) combines topic-space propagation with deterministic chunk-index linkage across name-equivalent states under frequency-adaptive decay. Third, continuous STIBP scores replace binary entity-match seeds in localized Personalized PageRank (PPR). We characterize the condition under which this prior assigns more PPR mass to ground-truth evidence than a binary prior. Under the reported protocol, STITCH-RAG attains the highest reported Contain-Acc and LLM-Acc point estimates among the compared methods on HotpotQA and 2WikiMultiHopQA, and higher Recall@8 than the methods included in the standardized retrieval comparison. Results on the mixed-domain benchmark remain auxiliary preference-based evidence because only LLM-judged accuracy is available.
cs.AI / 190 / 2609.35031
PEAR: Progressive Evidence-Based AutoResearch for Industrial Search Systems
Yifan Wang, Shipeng Zhu, Fei Xiong, Yuqin Yang, Yonghui Huang, Kunyao Wu, Yue Wang, Weichao Meng, Yu Gong
cs.IR · cs.AI
Abstract
AutoResearch improves systems through iterative experimentation: agents propose candidate modifications, evaluate them, and use the results to guide subsequent exploration. Applying this paradigm to industrial search presents two challenges. (1) Common AutoResearch approaches follow a keep-if-better rule, retaining the highest-scoring candidate for subsequent experiments. Under non-stationary traffic, transient gains may be mistaken for persistent improvements, impairing reliable accumulation of search knowledge. (2) Candidate modifications can be evaluated at multiple fidelity levels, from low-cost proxies to online validation, differing in cost, objective alignment, and statistical reliability. Existing methods rely on individual signals or task-specific procedures, lacking a unified basis for using evidence across levels to guide search. We introduce Progressive Evidence-Based AutoResearch (PEAR) with two complementary components. Evidence-driven AutoResearch maintains an independent, hypothesis-guided research state for each strategy task within a predefined objective and intervention scope. Each state evolves through a Plan-Execute-Evaluate-Update transition that links experimentation to context-aware evidence interpretation and hypothesis revision. Confidence-Gated Verifier Ladder organizes evaluation into four levels of increasing fidelity: Offline Replay, Shadow-Traffic Evaluation, Rapid Online Evaluation, and Decision-Grade Online Evaluation. A unified confidence-based gate promotes candidates only when evidence supports a statistically significant positive effect, enabling broad low-cost exploration while reserving costly online experiments for promoted candidates. In a real-world industrial search system, strategies optimized with PEAR significantly increased Main Order/DAU by 2.7336% and 3.2957% relative to their respective baselines in two A/B experiments.
cs.AI / 191 / 2609.35425
Semantic Prefix Oracles for LLM Decoding: Contracts and Differential Validation
Paul Kronlund-Drouault
cs.PL · cs.AI · cs.CL
Abstract
Constrained decoding can enforce regular or context-free output formats, but many program-generation failures are semantic: scope, typing, and declaration effects depend on context. We present semantic grammar specifications, a declarative formalism that attaches such constraints to a context-free surface and executes them during Earley descent. Our implementation enforces \emph{safe pruning}: it rejects only prefixes whose semantic contradictions cannot be repaired by any continuation. A separate, grammar-dependent, \emph{dead-end freedom} property guarantees the existence of a realizable witness for each remaining branch. We give simple sufficient conditions based on surface productivity, type coverage, and left-to-right constraint flow. Our finite-lambda, core ML, and C-like fragments satisfy them, while the STLC instance used in our experiments does not: plain STLC can violate type coverage, and we show how restricting its type universe recovers it. A tokenizer-lifting lemma carries character-level witnesses to token sequences under an explicit vocabulary-coverage hypothesis. We validate the implementation differentially against production compilers (\texttt{ocamlc}, \texttt{cc}). Across every prefix of 65 compiler-valid programs we observe zero false prunes. The semantic oracle localizes 25/30 invalid programs mid-stream, against 0/30 for a syntax-only oracle, and agrees on 42/42 recursion probes. A twelve-model generation study, including a matched semantic-versus-syntactic ablation for nine models, finds nonnegative observed semantic-minus-syntactic point estimates for every model-language pair, with maxima of $+15.2$ points on STLC task correctness and $+14.3$ points on ML validity.
cs.AI / 192 / 2609.33832
Achieve What You Imagined: Learning to Align Actions with Visual Plans
Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo, Yichen Luo, Zhibo Pang, Peng Zhou, Sichao Liu
cs.RO · cs.AI · cs.LG
Abstract
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $π_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/
cs.AI / 193 / 2609.33872
Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
Sichao Liu, Zekun Wang, Lixuan Tang, Yiming Li, Xiaohan Wang, Hanzhi Zhang, Daqiang Guo, Peng Zhou, Lihui Wang
cs.RO · cs.AI · cs.LG
Abstract
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. Our approach constructs a high-fidelity robotic environment from RGB-D observations using 3D Gaussian Splatting and SAM3D, enabling ``simulation and evaluation before acting''. It integrates visual observations and language instructions with spatio-temporal reasoning for long-horizon task planning using large vision-language models. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Before execution, candidate action sequences are simulated and evaluated in the Gaussian-SAM environment to filter infeasible behaviours. We validate our approach on representative manipulation tasks involving rigid, soft, and deformable objects, including cube placing, toy packing, and duck rearrangement, demonstrating that geometry-aware spatio-temporal reasoning and state-aware execution improve manipulation reliability across different object categories. Our results suggest that combining geometry-aware reconstruction with high-quality rendering and simulation provides a scalable approach for evaluating robotic manipulation behaviours. Website: https://robot-gst.github.io
cs.AI / 194 / 2609.34085
AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska
cs.RO · cs.AI · cs.CV · cs.LG
Abstract
Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA
cs.AI / 195 / 2609.34182
Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
Wenqiao Li, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu, Tengyu Liu, Kaifeng Zhang, Chuan Wen, Siyuan Huang
cs.RO · cs.AI
Abstract
Dexterous manipulation requires tactile feedback.However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
cs.AI / 196 / 2609.34250
WAM-OPD: Sharpening World Action Models via On-Policy Distillation
Panjun Liu, Xiaohan Lei, Shiqi Zhang, Yikun Wang, Yongxin Zhang, Mingyi Hu, Shida Sun, Jiateng Shou, Wengang Zhou, Jiajun Deng, Zhiwei Xiong
cs.RO · cs.AI
Abstract
Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. WAM-OPD inherits the advantage of OPD methods that transfer task-specific teacher knowledge under the student's own induced distribution, rather than directly fitting the student to a narrow task-specific data distribution. However, in closed-loop manipulation, the observation histories change as the student policy evolves, requiring fresh environment rollouts to remain on-policy. Applying OPD to WAMs entails repeated data collection, which is costly even in simulation and often impractical on real robots. To avoid repeated environment rollouts during distillation, we introduce prefix-weighted trajectory replay (PWTR). PWTR uses a fixed trajectory pool composed primarily of initial-student rollouts, supplemented with task-specific teacher rollouts to broaden trajectory coverage. For each trajectory replayed from this pool, PWTR conditions the current policy on successive stored histories to generate fresh denoising paths, along which the task-specific teacher provides supervision. Although these denoising paths are refreshed as the policy evolves, the replayed environment trajectories remain fixed. PWTR therefore reweights per-decision distillation losses using proxy importance weights derived from path scores accumulated over the trajectory prefix preceding each decision to mitigate the resulting shift in the history distribution. Simulated and real-world experiments demonstrate task adaptation without additional environment interaction during distillation. In both settings, WAM-OPD improves target-task performance while retaining near-initial performance on tasks excluded from adaptation.
cs.AI / 197 / 2609.34300
When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
John Cao, Somil Bansal
cs.RO · cs.AI · eess.SY
Abstract
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.
cs.AI / 198 / 2609.34356
Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control
Taekyung Kim, Salem Fradi, Yanning Dai, Mateusz Ostaszewski, Jürgen Schmidhuber
cs.RO · cs.AI · cs.LG
Abstract
Physical interactions can create future hazards that are not apparent from the robot's current geometric surroundings. We present a framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering. A vision-language model (VLM) predicts physical events and their timing or directly predicts object displacements. An explicit motion model converts event hypotheses into object trajectories. Split conformal prediction calibrates position errors jointly across specified objects, observation times, and future times; geometric shape bounds convert the resulting position regions into predicted object occupancy. PSS evaluates a prescribed backup maneuver against this occupancy and derives input-affine constraints for minimally modifying the nominal input while preserving backup feasibility under the robot dynamics and input limits. MuJoCo experiments with a Unitree Go1 consider falling fixtures, impact-driven support loss, and contact propagation. PSS achieves a safe episode rate of 99.3%, compared with 43.3% for a Backup Control Barrier Function baseline that only uses current obstacle geometry.
cs.AI / 199 / 2609.34695
Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control
Yisheng Zhang, Tao Wang, Sicun Gao
cs.RO · cs.AI
Abstract
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.
cs.AI / 200 / 2609.34782
CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration
Hyunjin Park, Jebeom Chae, Minwoo Park, Sunghyun Park, Hanjun Yoo, Seoyeon Choi, Soochul Yoo, Joohwan Seo, Sarmad Idrees, Jae-Sang Hyun, Jongmin Lee, Roberto Horowitz, Youngwoon Lee, Jongeun Choi
cs.RO · cs.AI · cs.CV · cs.LG
Abstract
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations. CoHuB provides 10 tasks, eight with two humanoids and two with three humanoids, spanning diverse collaboration patterns. We also provide synchronized demonstrations collected through a multi-operator VR teleoperation pipeline, in which each operator controls one humanoid from its egocentric view. Experiments with representative visuomotor policies reveal substantial challenges across different forms of coordinated perception and control. CoHuB provides a foundation for developing and evaluating multi-humanoid collaboration policies.
cs.AI / 201 / 2609.34911
Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye
cs.RO · cs.AI · cs.CV · cs.LG
Abstract
Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose *Action Upcycling*, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2--1.7$\times$ with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.
cs.AI / 202 / 2609.34968
RoboFL: Federated Expert Assembly for World Action Models
Rongyu Zhang, Ruizhi Fan, Yunfan Lou, Hengyu Fang, Shenli Zheng, Chenrui Wu, Yili Jin, Li Du, Dan Wang, Yuan Du, Shanghang Zhang
cs.RO · cs.AI
Abstract
Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.
cs.AI / 203 / 2609.35003
Learning to Act under Visual Interruptions with Vision-Language-Action Models
Mingle Jiang, Rui Xu, Yunke Wang, Chang Xu
cs.RO · cs.AI
Abstract
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on $π_{0.5}$ and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/
cs.AI / 204 / 2609.35047
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
cs.RO · cs.AI · cs.LG
Abstract
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
cs.AI / 205 / 2609.35200
ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation
Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li, Moritz Reuss, Rudolf Lioutikov
cs.RO · cs.AI · cs.CV · cs.LG
Abstract
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3\% average success on LIBERO and 62.4\% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7\% average success, against 8.3\% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT
cs.AI / 206 / 2609.35249
Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies
Dingsheng Liu, Yangzheng Wu, Mahboubeh Asadi, Zhiyuan Li, Jinbang Huang, Yixin Xiao, Tongtong Cao, Yingxue Zhang
cs.RO · cs.AI · cs.CV
Abstract
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $π_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.
cs.AI / 207 / 2609.35375
From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations
Bangjun Wang, Longyan Wu, Yukun Wei, Shenghe Shao, Chaoyi Huang, Wenze Cui, Zetong Xu, Hanlin Wu, Long Chen, Yi Ma, Hongyang Li
cs.RO · cs.AI
Abstract
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.
cs.AI / 208 / 2609.35575
F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
cs.RO · cs.AI
Abstract
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
cs.AI / 209 / 2609.34648
SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows
Tianxin Xie, Pengfei Zhang, Kai Jiang, Zelin Zhao, Li Liu
cs.SD · cs.AI
Abstract
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.
cs.AI / 210 / 2609.34907
SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living
Debolina Chowdhury, Suman Samui, Sujoy Saha
cs.SD · cs.AI · cs.LG
Abstract
Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filter bank followed by a depthwise-separable convolutional body. Each sinc filter is controlled by two frequency parameters, allowing the learned passbands to be inspected directly in hertz while keeping the front end small. To reduce room-specific leakage, recording sessions and environments are separated before overlapping windows are assigned to the training, validation, and test partitions. We further use multi-objective Bayesian optimization as a design tool to examine the validation performance--model-size trade-off across 24 configurations. The selected designs span different operating points: the best-performing model achieves 80.2\% accuracy and 0.760 macro-F1 with 14{,}040 parameters, while the compact $N_f=25$ configuration uses only 2{,}848 parameters and achieves 75.7\% accuracy, 0.661 macro-F1, and 0.716 MCC on the held-out environment. Analysis of the learned filters and confusion patterns shows that spectral overlap contributes to confusion among water-related events, while the \textit{Door}/\textit{Walker/Crutch} errors also reflect similarities in their transient temporal structure.
cs.AI / 211 / 2609.34931
JazzSAMBA: A Synchronous and Asynchronous Multi-take Band Audio Dataset of Jazz Standards for Live Music Models
Phillip Long, Jacob Nguyen, Jace Hosto, Gage Hosto, Jett Takazawa, Fares Nofal, Sebastian Stade, Nithya Shikarpur, Julian McAuley, Cheng-Zhi Anna Huang, Stephen Brade, Aleksandra Teng Ma
cs.SD · cs.AI · cs.LG · cs.MM · eess.AS
Abstract
Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-annotated corpus of clean per-stem combo recordings on standards. We introduce JazzSAMBA (Jazz Synchronous and Asynchronous Multi-take Band Audio) to fill this gap: the first originally recorded jazz-combo multitrack dataset of standards with asynchronous (overdubbed) and synchronous (live ensemble) protocols, preferred and alternate takes chosen by the musicians, and timed annotations for bars, chords, sections, and soloists. JazzSAMBA covers 76 standards by eight musicians on drums, bass, piano, trumpet, and saxophone, with per-stem audio, mixtures, and MIDI. It can support chart-conditioned accompaniment, combo source separation, and form-aware music information retrieval. We demonstrate the dataset on two tasks: a jazz combo source-separation baseline and a chart-conditioned accompaniment ablation. The dataset, code, and samples are linked from the project demo page.
cs.AI / 212 / 2609.35411
GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection
Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu
cs.SD · cs.AI
Abstract
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.
cs.AI / 213 / 2609.35196
GAC-PINN: Geometry-Adaptive and Constraint-Enhanced Physics-Informed Neural Networks
Yanxin Zhang, Yong Zhang, Houbiao Li
math.NA · cs.AI
Abstract
For systems with steep gradients, sharp interfaces, or severe spatio-temporal coupling, Physics-informed neural networks (PINNs) suffer from spectral bias, geometric inflexibility, and boundary constraint conflicts, which undermine accuracy and convergence. To overcome these issues, we propose a geometry-adaptive and constraint-enhanced PINN (GAC-PINN). The framework comprises four components: a gradient-driven adaptive grid mapping (AGM) for diffeomorphic point concentration with Jacobian regularization, an adaptive bandwidth hard-constraint ansatz with spatially-varying boundary transition widths, a Gaussian Fourier feature mapping as a spectral preconditioner to further enhance high-wavenumber representation, and an operator-aware router that automatically selects the appropriate hard-constraint construction based on whether the governing PDE contains temporal derivatives. An AGM callback mechanism and a three-stage training strategy ensure stable coordination. Benchmarks including the viscous Burgers equation, a sharp-peaked 2D Poisson problem, and the Allen-Cahn phase-transition equation show that GAC-PINN attains relative (L^2) errors of ((1.747\pm 0.450)\times 10^{-4}), ((2.868\pm 0.947)\times 10^{-5}), and ((1.756 \pm 0.712)\times 10^{-3}), respectively, consistently outperforming the baselines. Ablation studies further reveal that AGM alone yields a substantially lower error than residual-based adaptive refinement (RAR), while RAR becomes beneficial only when combined with FFM, demonstrating a context-dependent module interaction. Convergence analysis verifies rapid error reduction and saturation with increasing resolution, establishing a practical adaptive framework for high-fidelity simulation of problems with localized sharp features in applied mechanics and computational physics.
cs.AI / 214 / 2609.33994
A packet-level digital hardware twin for commissioning megahertz diagnostic edge AI and plasma control system integration in tokamaks
Semin Joung, Abhilasha Dave, Luca Scomparin, Filipp Khabanov, Zheng Yan, Benedikt Geiger, George McKee, Ryan N. Coffee, David R. Smith
physics.plasm-ph · cs.AI
Abstract
High-bandwidth plasma diagnostics increasingly provide inputs to machine learning and signal-processing algorithms intended for real-time tokamak control, but the complete path from diagnostic sampling to control-system handoff is difficult to commission because of their sampling rates. We develop a packet-level digital hardware twin for the megahertz diagnostic edge-AI architecture. The simulator represents a 64-channel, 1 MHz beam emission spectroscopy (BES) diagnostic embedded in a 96-channel dual-carrier acquisition system, two 48-channel streams with SPAD0 sample counters, 10 GbE Hardware UDP transport, packet loss and network jitter, FPGA parsing and dual-carrier alignment, causal preprocessing and edge inference, a compact Ethernet result packet, receiver-side shared state, and a 1 kHz PCS-like control cycle. Binary UDP payloads and PCAP files are generated rather than emulating transport only at the array level. With a baseline of 20 samples per packet, each carrier generates 50,000 packets s$^{-1}$ and 100 MB s$^{-1}$ of user payload. A 110 ms reference run produces 11,000 HUDP packets; an intentionally dropped 20-sample packet is detected by the sample-counter continuity logic and invalidates the two overlapping 128-sample inference windows without silent interpolation. For valid windows, the configured engineering latency model gives a median last-input-to-shared-memory latency of 91.6 us and a 99th percentile of 108.1 us. A separate operating-system loopback test sends binary FPGA-result datagrams through a UDP receiver into POSIX shared memory and preserves packet sequence and CRC for 20/20 packets. Interactive GUI interfaces expose timing, packetization, network faults, inference thresholds, and control-state inspection. The framework provides a reproducible environment for testing diagnostic-to-accelerator interfaces and fail-safe behavior for deployment on fusion devices.
cs.AI / 215 / 2609.33835
An Active-Bottleneck Mechanism for Weak-to-Strong Generalization
Mohammad Zeinalpour, Amir Najafi
stat.ML · cs.AI · cs.LG
Abstract
Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no assumption that the student is more expressive than the teacher. In two-stage linear regression, a teacher is fit from $n$ labeled examples and a student is trained solely on the teacher's predictions on $m$ fresh, unlabeled inputs. Although both stages share the same hypothesis class and the same training rule, we show that the student outperforms the teacher exactly when $m$ lies in an explicit intermediate range: too few pseudo-labels leave the student without enough signal, too many let it inherit the teacher's noise. Under power-law covariance, we derive this range in closed form as a function of the spectral decay and noise level, including regimes where the improving region splits into two disjoint intervals of $m$. We then study a random-feature model in which the student has strictly more features than the teacher, and identify two regimes, again given by explicit thresholds: one where improvement occurs only for $m$ in a bounded interval, and one where it occurs only once the student width $N_S$ exceeds an explicit threshold. Both regimes are governed by a single "active-bottleneck" principle: whichever of $m$ or $N_S$ is scarcer controls how much teacher error is filtered out, while increasing the other resource only reduces estimation noise. Together, these results show that finite data and finite width can themselves regularize a two-stage learner, with no explicit mechanism doing so.
cs.AI / 216 / 2609.33972
HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases
Quan D. Bui, Nguyen Do, An Nguyen Dang, Huyen Nguyen, Nhu Duc Minh Nguyen, My T. Thai
stat.ML · cs.AI · cs.LG
Abstract
Existing interpretable graph additive models still face limitations in either computational scalability or modeling flexibility. In terms of structural modeling, previous approaches either face quadratic scaling costs or sacrifice explicit source-to-target contribution decomposition. In terms of feature components, they rely either on per-feature neural networks or on single shared bases with limited feature specialization. We address both problems by introducing HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases, an interpretable-by-design framework. For feature modeling, HARMONIA introduces a Mixture of Neural Bases (MoNB), which routes features to specialized basis experts, enabling parameter sharing without sacrificing feature-specific specialization. For structural modeling, HARMONIA uses Relative Random Walk Probabilities (RRWP) to capture multi-hop and multi-path relationships, and proposes Sparse RRWP Aggregation (SRA) to compute these interactions through sparse graph propagation without quadratic pairwise complexity. HARMONIA retains a simple additive form in which predictions decompose into feature responses modulated by structural influence. Empirically, HARMONIA achieves stronger explanation recovery than existing interpretable graph baselines while maintaining competitive predictive performance and scaling to graphs with millions of nodes. These results show that interpretable graph learning can remain both faithful and scalable without sacrificing predictive effectiveness.
cs.AI / 217 / 2609.34458
Understanding Generalization Requires Universal Induction
Aram Ebtekar, Marcus Hutter, Danica J. Sutherland
stat.ML · cs.AI · cs.IT · cs.LG · math.ST
Abstract
Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments to underperform on others. We might hope that past experience informs which environments to expect, but NFL applies equally to meta-learning. Thus, any method that makes meaningful predictions necessarily begins with an inductive bias external to the data. Choosing to bias toward short programs yields Solomonoff induction (SI), whose performance is competitive against all computable learners - albeit up to "constants" that become large when comparing against specialized methods that exploit background information. We therefore relativize SI to an information vantage point, biasing toward short programs with access to all preexisting information. This reframes the inductive bias: instead of seeking some absolute notion of simplicity, we favor accessibility with respect to our vantage point. An algorithm can only outpredict the relativized SI to the extent that its code contains additional information about the data, and no algorithm can generate such information. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute, and there is evidence to suggest that frontier AI systems roughly approximate it. Thus, the only known answer to meta-NFL is rooted in algorithmic information theory, which we should expect to play a fundamental role in explaining the generalization behavior of modern (and future) AI systems.
机器学习 (cs.LG)
296
cs.LG / 1 / 2609.34292
Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale
Jun-qiang Lu
cond-mat.dis-nn · cond-mat.mtrl-sci · cond-mat.other · cs.LG
Abstract
We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.
cs.LG / 2 / 2609.33978
Adapting neural operators for mechanics decisions under changing operating conditions
Prashant K. Jha, Koffi Enakoutsa, Ian Galloway, Henry Anderson
cs.CE · cs.LG · math.NA
Abstract
Neural operators can accelerate repeated nonlinear mechanics calculations, but their accuracy can deteriorate as operating conditions move beyond the training range. This work studies whether high-fidelity solutions acquired during use can be reused to adapt a neural operator and improve subsequent mechanics-based command selection. Two hard-magnetic soft-material systems are simulated using high-fidelity finite-element (FE) models, providing reference solutions for evaluating surrogate predictions and selected commands. A neural operator predicts deformation from known material, loading, and magnetic-field inputs, while an empirical error estimator determines which predictions may be used for command selection. Selected FE evaluations supplement these predictions, and their complete loading paths are retained for periodic updates of the neural operator and estimator. In both examples, the fixed operator loses substantial accuracy when stiffness and loading move outside the training range. Updates using 16 acquired paths recover much of the lost accuracy while preserving accuracy in the nominal regime. Under the same FE evaluation budget, the updated operators also improve command selection, although the benefit varies with the operating condition. Error estimation is less consistent, with inaccurate predictions sometimes accepted and accurate predictions rejected. These results demonstrate that reusing high-fidelity loading paths can extend the useful operating range of a neural operator. However, improved forward accuracy alone does not guarantee reliable prediction acceptance, highlighting prediction-specific error assessment as a separate requirement for trustworthy decision making.
cs.LG / 3 / 2609.34035
3D Point Tracking with State Space Models
Masahiro Ogawa, Qi An, Atsushi Yamashita
cs.CV · cs.LG · cs.RO
Abstract
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.
cs.LG / 4 / 2609.34765
Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models
Minchan Kang, Kyeonghye Park, Seungyeon Sa, Seoyoung Cho, Daeshik Kim, Yucheol Cho
cs.CV · cs.LG
Abstract
Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss with respect to the full-precision model, potentially over-preserving FP behavior and calibration-specific bias. Rather than treating quantization solely as an error to be minimized, we observe that it can also provide beneficial regularization for certain layers and modalities. Motivated by this observation, we propose Balanced Fitting, a quantization effect-based framework that balances precision and regularization beyond reconstruction-based optimization. By measuring layer- and component-wise quantization effects for weights, vision activations, and text activations, Balanced Fitting combines fine-grained fitting for sensitive components with coarser fitting to exploit potential regularization benefits. Experiments on multiple LVLMs show that our method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance. The source code is publicly available at https://github.com/kmc3661/BFQ
cs.LG / 5 / 2609.34809
From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models
Rong Yu Xu, Prayag Tiwari, Shaolei Zhang
cs.CV · cs.LG
Abstract
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.
cs.LG / 6 / 2609.34861
When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model
Minchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim, Yucheol Cho
cs.CV · cs.LG
Abstract
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT
cs.LG / 7 / 2609.35096
DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection
Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou, Panos K. Papadopoulos, Yannis Spyridis, Dimitrios Zarpalas, Vasileios Argyriou
cs.CV · cs.LG
Abstract
Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this need since localization-based approaches lack semantic descriptions while language-based explanation methods are only weakly grounded in visual evidence. In this work, we propose DF-CBM, a region-aware concept bottleneck model for explainable deepfake detection. DF-CBM builds a compact vocabulary of manipulation-related concepts from textual artifact annotations and links each concept to plausible facial and boundary regions. It then predicts these concepts from visual features using a concept-specific masked attention mechanism guided by parsed facial masks and the final real/fake decision is made from the predicted concept bottleneck. Our experiments show that DF-CBM outperforms concept-based baselines in concept prediction and deepfake classification while remaining competitive with state-of-the-art black-box detectors. Finally, qualitative results and intervention analyses demonstrate that DF-CBM provides spatially grounded concept evidence and enables counterfactual explanations of how individual manipulation concepts influence the final prediction. Our code is available at: https://github.com/GeorgeTsoumplekas/DF-CBM.
cs.LG / 8 / 2609.35189
G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA
Jia Song, Wenhow Li, Lichen Bai, Bada Ye, Zeke Xie
cs.CV · cs.LG
Abstract
Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G$^3$-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.
cs.LG / 9 / 2609.35232
Beyond Selection: Token Parameterization for Extreme Visual Token Compression
Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo
cs.CV · cs.LG
Abstract
Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under $23\times$--$64\times$ compression and remains competitive at $144\times$, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using $16.6\times$/$78.8\times$ lower compressor latency/FLOPs.
cs.LG / 10 / 2609.35247
AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
Mohamed Eltahir, Fardows Adam, Duaa M. Tahir, Lama Alamoudi, Sana Ammar, Atheer A. Alboloshi, Jory Albluey, Tanveer Hussain, Naeemullah Khan
cs.CV · cs.LG
Abstract
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes'' posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.
cs.LG / 11 / 2609.35449
From internal representations to model improvement through prediction errors
Yushi Nakaya, Kenichi Higuchi, Shuichi Ishida
cs.CV · cs.LG
Abstract
With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but those signals do not directly capture changes in the model being improved. The target model's own internal features reflect what it has learned so far and change with retraining, making them a natural cue for choosing the next training data. However, feature rarity alone does not reveal the errors that matter for performance. Here we link internal features to prediction errors and their expected impact on performance and select images for labeling and retraining without using labels for candidate images. We evaluated the method with an object detector on two datasets and two pairs of random seeds. Adding internal features improved the identification of prediction errors in 15 of 16 conditions. When performance was averaged over successive labeling rounds, the method outperformed selection based only on feature rarity in all four evaluation settings and ranked among the top two of six methods. With other conditions held fixed, performance after retraining was again higher than with rarity-based selection, even though the latter collected more errors. With longer retraining, the proposed method ranked first among six methods. These results suggest that linking a model's internal features to its errors and their effects on performance may help select training images that improve performance, thereby allowing the model's current state to guide which images are labeled next.
cs.LG / 12 / 2609.35473
Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
Carlos Garrido-Munoz, Jorge Calvo-Zaragoza
cs.CV · cs.LG
Abstract
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.
cs.LG / 13 / 2609.35611
On-Policy Self-Distillation for Multi-Turn Image Editing
Liangbing Zhao, Le Zhuo, Mohamed Elhoseiny
cs.CV · cs.LG
Abstract
Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure to a train-test mismatch in the conditioning distribution: models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference time. To address this, we propose MT-OPSD, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations. We further introduce LME-Bench, a benchmark of 100 ten-turn editing sessions for evaluating long-horizon robustness. Experiments across three editing backbones show that MT-OPSD substantially improves long-horizon editing success and reduces multi-turn collapse while largely preserving single-turn editing quality.
cs.LG / 14 / 2609.34586
AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum
Michalis Kasioulis, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
cs.DC · cs.LG
Abstract
Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent development, observability, or benchmarking separately, offering limited support for the full lifecycle of distributed agentic applications. This paper presents AgentWare, an AgenticOps framework that automates the provisioning, deployment, observability, and evaluation of agentic applications across Edge-to-Cloud infrastructures. AgentWare introduces an end-to-end lifecycle pipeline that automatically prepares heterogeneous execution environments, transforms user-defined agent implementations into distributed applications, deploys agent components across the continuum, and performs unified collection of execution traces, infrastructure telemetry, and evaluation metrics. The framework further supports automated semantic evaluation through LLM-as-a-Judge workflows and generates reproducible reports covering correctness, performance, resource utilization, and energy consumption. We demonstrate the applicability of AgentWare through a distributed book assistant agent deployed across real Edge-to-Cloud infrastructure under multiple deployment and model configurations. The results show that AgentWare enables systematic experimentation and evaluation of distributed agentic applications while significantly reducing the manual effort required for deployment, instrumentation, and analysis.
cs.LG / 15 / 2609.34818
Accelerator Choice Is Not Enough: AlphaFold2 Inference on Cloud TPUs
Lorenzo Pazienza, Ihab El Bani
cs.DC · cs.LG · cs.PF
Abstract
AlphaFold2 is written in JAX, so the same inference code compiles and runs unchanged on CPUs, GPUs and Google Cloud TPUs. That portability makes the accelerator look like the main decision a user has to make. We show that it is not. Running one AlphaFold2 inference workload across a Colab CPU runtime, an NVIDIA T4 GPU and a dedicated eight-chip Cloud TPU v5e slice, we find a large hardware advantage for the TPU, 0.47 s per call in steady state on a single chip against 13.1 s on the T4 in the same measurement campaign, and three ways in which the software layer decides how much of it a user actually gets. The default execution path uses one chip of the eight, and at list prices the idle capacity makes the slice cost about as much per prediction as the GPU. Batching with jax.vmap never exceeds single-query throughput, while mapping queries across chips with jax.pmap gives eight chips 6.5-7.9x the throughput of one on a matched grid; automatic sharding leaves the per-chip footprint unchanged, consistent with replication, most plausibly because AlphaFold2 carries no sharding annotations. Our retained trace analysis of a first call at a new input shape reports about three quarters of the traced span in JAX tracing and compilation rather than execution. Reruns five weeks later reproduced neither cloud baseline, the GPU one off by roughly a factor of two, so the hardware ratio above is specific to one campaign.
cs.LG / 16 / 2609.35481
TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training
Jiacheng Zhu, Xie Zhao, Gongming Zhao, Hongli Xu, Yao Fei, Jin Fang
cs.DC · cs.LG
Abstract
Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2\%--11.4\% across three representative MoE models.
cs.LG / 17 / 2609.34074
The Double-Edged Sword of Information: Revealed versus Hidden Lotteries in School Choice
Parinaz Naghizadeh, Jingyan Wang
cs.GT · cs.LG · econ.TH
Abstract
In school choice, a lottery number is often used by the matching mechanism to break ties when there are more students who prefer the same school than the number of seats available. There has been growing theoretical and empirical interest in understanding the impact of revealing the lottery number to students. In practice, in recent years, the NYC Public Schools started revealing the lottery number to students to improve transparency. Theoretical findings from prior literature also suggest that revealing the lottery number strictly improves the number of matches under the deferred acceptance algorithm. However, these theoretical results are based on the over-simplifying assumption that all students share the same preference ranking for schools. In this work, we relax this assumption and allow students to have heterogeneous preference rankings. Under a game-theoretic model where student strategies form a Bayesian Nash equilibrium, we characterize scenarios where revealing the lottery number can either improve or worsen the matching outcome, measured by two metrics of match rate and social welfare. We further consider revealing partial information about the lottery, and demonstrate non-monotonic effects in the amount of information available to students. These results together illustrate complex tradeoffs induced by the lottery revealing policy.
cs.LG / 18 / 2609.35034
Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition
Riki Okamura, Toshiharu Sugawara
cs.IR · cs.LG
Abstract
Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We propose two estimators based on the decomposition of clicks into examination and relevance. First, the latent-examination independent inverse propensity score (LE-IIPS) estimator corrects the IIPS bias using policy examination probability ratios. Second, the examination-decomposed doubly robust (ED-DR) estimator extends LE-IIPS to a doubly robust framework. ED-DR is unbiased if the examination probabilities are correct regardless of relevance accuracy, or under ranking-independent examination, even if both model estimates are inaccurate. Experiments show that ED-DR achieves a lower MSE than existing methods with large sample sizes, especially when the examination depends on ranking. We also highlight its limitations under small samples or cascade user behavior conditions.
cs.LG / 19 / 2609.35684
The Hidden Perception Constraint in Task-Aware Compression
Sahan Liyanaarachchi, Semih Akkoc, Sennur Ulukus, Aylin Yener
cs.IT · cs.LG
Abstract
With the recent advancements of neural compressors, explicitly incorporating perception constraints into the design of compression schemes has gained significant attention. Traditionally, these perception constraints ensure that the distribution of the reconstruction does not significantly deviate from the distribution of the source, thus attesting to the perceptual quality of the reconstruction. In this work, we uncover several perception constraints that are naturally present in task-aware compression. In particular, we consider a problem where the primary task is reconstruction and the secondary task is classification (i.e., a statistical test). We study this problem at varying levels of domain information available to us and discuss how to utilize the naturally emerging perception constraints to design rate-minimal compression schemes that also maximize the utility of our secondary task. We show that in this setting, if the decision boundaries of the classifier are ill-defined (mismatch) for our source distribution, then matching onto a target distribution enhances our classification accuracy.
cs.LG / 20 / 2609.33828
dOPT: Differentiating Conic Optimization via Geometric Reduction
Fengyu Yang, Connor W. Magoon, Tyler Watts, Shahar Z. Kovalsky
cs.LG · math.OC
Abstract
Optimization layers enable the incorporation of structured constraints and decision problems into learning systems. Training such systems requires differentiating through the embedded optimization problem, which can be challenging for general conic programs. We introduce dOPT, a solver-agnostic framework that, rather than differentiating the full conic formulation, reduces it at a computed primal-dual solution to an equality-constrained quadratic program that preserves the reference solution and its first-order sensitivity. The reduction captures the local first- and second-order conic geometry relevant to differentiation and remains well defined at singular configurations. Computing solution derivatives then requires a single symmetric linear solve, independently of the forward solver. We derive explicit reductions for convex NLPs, QPs, SOCPs, and SDPs. Numerical experiments validate the computed gradients and show favorable backward-pass scalability, with substantial speedups over existing differentiable conic optimization methods as problem size increases.
cs.LG / 21 / 2609.33851
Rethinking Contextualization by Reinterpreting Attention Head Channels
Hakaze Cho, Haolin Yang, Zhun Sun, Naoya Inoue, Benjamin Heinzerling, Kentaro Inui
cs.LG · cs.CL
Abstract
Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete dictionary, lacking a global view of their general behavior. Therefore, we propose a general principle: Globally, we find and estimate that different words carry different amounts of information, and less-informative words tend to absorb more contextual information. Specifically, these low-information words do not absorb contextual words uniformly, and finer-grained selectivity enables more precise routing to promote information transmission between matched words. Moreover, to find what mechanism causes such processing, we reinterpret attention heads as channels gated by their singular vectors and find that: (1) these singular vectors point to the hidden states of more informative words, allowing such words to write their information to others more strongly to act as information sources, and vice versa; and (2) these singular vectors can be viewed equally as hidden state features, enabling automated interpretation of attention heads beyond prior heuristic head discovery, also embedding heads into a continuous space rather than treating them as discrete, independent dictionary entries.
cs.LG / 22 / 2609.33874
JET: Justification Evaluation in Transformer
Shenghao Ding
cs.LG · cs.PF
Abstract
JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.
cs.LG / 23 / 2609.33888
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits
Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer
cs.LG
Abstract
Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance weighting. For $A$ actions, $T$ rounds, and a finite prediction class $F$, vanilla PO achieves $\widetilde O(\sqrt{AT\log(|F|)})$ regret with high probability. Our analysis reveals an implicit exploration mechanism of independent interest: gradual policy updates prevent actions from losing probability too quickly, allowing the regression oracle to learn their expected losses. We further develop a batched version using only $O(\log T)$ regression calls and policy switches, and show how private regression oracles yield differentially private contextual bandit algorithms without composition across batches. For a finite class, this gives pure $\varepsilon_{\rm priv}$-DP and regret $\widetilde O\left( \sqrt{AT \log(|F|/δ)}(1+\varepsilon_{\rm priv}^{-1/2}) \right)$. Finally, experiments across oracle-based contextual bandit algorithms, with and without privacy, demonstrate the practical effectiveness of policy optimization and the value of explicit exploration under stronger privacy constraints.
cs.LG / 24 / 2609.33891
Augmented Feature Boosting for Multicalibration
Ira Globus-Harris, Inbal Livni Navon
cs.LG
Abstract
Multicalibration requires a predictor's residuals to be unbiased not only globally, but also after conditioning on the predictor's own level sets and reweighting by a rich class of test functions. Standard boosting approaches in the distributional setting achieve this by repeatedly discretizing the predictor's range then auditing and repairing the resulting level sets. One consequence is that in practice, the algorithm's guarantees are sensitive to this parametrization of the rounding parameter. A natural theoretical question, then, is how to do discretization-free boosting which avoids this rounding within the boosting process itself. Here, we analyze an alternative feature-augmentation boosting paradigm inspired by Tax et al. (2026): at each round, a squared-loss oracle is called on hypotheses that receive the previous predictor's output as an additional feature, and only the final predictor is rounded to have a finite set of level sets to provide the multicalibration guarantee with respect to. We give a theoretical analysis of this procedure through the expressivity of the augmented hypothesis class, and show how the expressivity of this class yields a hierarchy of guarantees, including multiaccuracy, multicalibration, and the stronger notion of level-set multicalibration.
cs.LG / 25 / 2609.33893
MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
Zhi Wen Soi, Giulio Segalini, Jian-Jia Chen, Lydia Chen
cs.LG
Abstract
Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the audio. We challenge correctness-based evaluation by defining two hallucination categories: (i) context, where claims are not grounded in the audio; and (ii) knowledge, where claims about audio-related topics lack support from externally verifiable facts. We introduce MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering both categories. To evaluate open-ended responses, we propose a groundedness judge that uses reference rubrics and judge prompts guided by human annotations. We extensively evaluate ten state-of-the-art LALMs and show that hallucination remains substantial. Even a frontier model such as Gemini 3.7 Flash reaches a hallucination rate of 36.5%. We further adapt and benchmark four mitigation methods from multiple domains for LALMs. Despite some improvements, effective hallucination mitigation remains an open challenge. Finally, we call on the community to evaluate hallucination and benchmark mitigation methods with MISHAP-Bench.
cs.LG / 26 / 2609.33898
No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability
Yiyong Liu, Jun Sakuma, Michael Backes, Rui Wen
cs.LG · cs.CR
Abstract
Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient pre-training, and simplified reinforcement learning pipelines. While these strategies drastically reduce overhead, they prompt a critical, yet neglected question: Is efficiency achieved at the expense of model robustness and security? To our knowledge, we present the first systematic cross-domain investigation of the efficiency-vulnerability trade-off. Across vision and language models, we show that efficiency-oriented training increases susceptibility to adversarial and privacy attacks. We characterize this vulnerability by analyzing the models' internal geometry and functional representations, demonstrating that the evaluated efficient variants consistently exhibit sharper loss geometry together with systematic changes in representational structure. We further extend our analysis to "zero RL training", finding that models trained using simplified RL recipes exhibit substantially greater susceptibility to catastrophic forgetting and more pronounced overconfidence than those trained through conventional alignment pipelines. Our findings suggest that training efficiency is rarely a "free lunch"; rather, the mechanisms that minimize computation can inadvertently compromise safety. We conclude by calling for a paradigm shift toward multi-objective training that jointly optimizes for performance, cost, and security.
cs.LG / 27 / 2609.33901
Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym
cs.LG · cs.AI
Abstract
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical foundations remain limited. In this work, we study when finite probe-based representations are sufficient for learning neural functionals. We establish general identification and universality results for probing, and show that using intermediate hidden representations can provide significantly more informative representations than relying only on final outputs. Motivated by these results, we introduce HIDDENPROBE, a simple architecture for learning from hidden probe responses. Across a range of neural functional benchmarks, including both MLPs and Transformers, HIDDENPROBE consistently improves over existing probing methods and achieves state-of-the-art performance. Our code is publicly available on GitHub.
cs.LG / 28 / 2609.33904
On the Two Faces of Adam in Separable Linear Classification
Chen Fan, Csaba Szepesvári
cs.LG
Abstract
We consider the behavior of deterministic, full-batch, bias-corrected Adam in separable linear classification with softmax parametrization under log-loss. In this setting, under a wide range of conditions Adam is known to approach max-norm-margin optimality when its stability constant $ε$ is zero, while with a positive $ε$, it is known to approach Euclidean-margin optimality. Our main contribution is the quantitative description of Adam's behavior for small fixed positive $ε$. We give sufficient conditions under which an Adam-trained classifier nearly maximizes the max-norm margin before the updates become gradient-like. We also show that the classifier reaches a fixed target Euclidean margin only much later. Specifically, we show that for polynomially decreasing stepsizes with exponent \(a\), where \(1/3<a<1\), the updates become approximately proportional to the negative gradient after $Θ(\log(1/ε)^{1/(1-a)})$ iterations. At that time, the classifier still nearly maximizes the max-norm margin. Reaching a fixed target Euclidean margin above that of every max-norm-optimal classifier, but below the optimum, is shown to require $ε^{-Θ(1)/(1-a)}$ iterations. Under inverse-linear stepsize decay (\(a=1\)), the update transition takes polynomially many iterations, whereas reaching the target margin takes exponentially many. Experiments support these predictions. The later change in the classifier can improve or worsen generalization after training error reaches zero, connecting the analysis to grokking and its reverse.
cs.LG / 29 / 2609.33906
JIVE: Jacobian-Informed Volume Expansion for Diverse Generative Sampling
Guangxun Zhang, Brian Cai, Boxuan Zhang, Chao Chen, Ruixiang Tang
cs.LG · cs.AI · cs.CV
Abstract
Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of the resulting endpoints. We introduce JIVE, a training-free framework that enhances generative diversity by injecting velocity perturbations aligned with the leading right singular subspace of the generator's endpoint Jacobian. By leveraging this local geometric structure, JIVE provably maximizes endpoint diversity while preserving sample quality. To maintain practical efficiency, we compute these perturbation directions via matrix-free iterations rooted in classical numerical linear algebra, requiring only a small computational overhead. Across different benchmarks, JIVE boosts both pixel and feature-level diversity in few-step and one-step generation.
cs.LG / 30 / 2609.33915
Training Witnesses: Trusting the Training without Trusting the Trainer
Houjun Liu, Pratyusha Sharma
cs.LG · cs.CR
Abstract
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
cs.LG / 31 / 2609.33927
Optimizing the Phi-2 Small Language Model for Real-time Chatbot Applications Using Parameter-Efficient Fine-Tuning (PEFT) with QLoRA Quantization
PhanTan Khanh Nguyen, Ashfaq Ali Shafin, Khandaker Mamun Ahmed
cs.LG
Abstract
This study explores the optimization of the Phi-2 Small Language Models (SLMs) for real-time chatbot applications through Parameter-Efficient Fine-Tuning (PEFT) and Quantized Low-Rank Adaptation (QLoRA). QLoRA specifically refers to the integration of PEFT with LoRA alongside a 4-bit quantization process, aimed at enhancing computational efficiency. These models, initially designed for high performance with minimal computational overhead, are further refined to address the constraints of mobile and edge computing environments. By integrating PEFT with QLoRA, the research aims to reduce memory usage significantly while maintaining, or potentially improving, the accuracy of model responses in real-time interactions. The effectiveness of these techniques was evaluated using the ROUGE metric system, which showed notable improvements in the summarization tasks performed by the models. This approach not only confirms the feasibility of using SLMs in resource-restricted environments but also opens up new avenues for deploying advanced AI-driven applications in real-time settings. The study's findings have significant implications for the development of efficient, scalable, and accessible AI technologies, paving the way for broader adoption in various industries.
cs.LG / 32 / 2609.33936
How Strong Is the Evidence for the Artificial Hivemind? Reevaluating Evidence for the Open-Ended Homogeneity of Language Models
Rylan Schaeffer, Brando Miranda, Joshua Kazdan, Jessica Chudnovsky, Sanmi Koyejo
cs.LG
Abstract
Recent research argues that language models exhibit pronounced homogeneity in open-ended generation, framing such behavior as an Artificial Hivemind that poses a long-term threat to human creativity. We examine three of its central results. First, the flagship example is that model responses to "Write a metaphor involving time" collapse into two clusters. Visualization, spectral analysis, clustering, and language model labels all contradict this description. The labels record each response's vehicle, what it compares time to. Our responses and the original authors' own show one dominant vehicle plus a heavy tail of distinct minority vehicles. "Time" is one of our least diverse topics, so the example is a favorable case, not a representative one. Second, the paper measures homogeneity against an undemanding null: responses to unrelated prompts. Under a more demanding null (same-prompt responses expressing genuinely different ideas), 20%-32% of such pairs already exceed the paper's 0.8 convergence threshold. A residual effect survives this null. The paper's same-prompt pairs exceed 0.8 roughly two to three times as often as our different-idea pairs. Much of what the paper calls homogeneity is the shared geometry of answering the same prompt. The remaining measurements lack any null: no human baseline is collected, and the model-indistinguishability statistic has no null. Third, the paper concludes that inference-time interventions are inadequate for combating the Artificial Hivemind, writing that "more generalizable solutions are needed at the model training level." We show that this conclusion is unsupported in three ways, and that an inference-time intervention (prompting) reliably raises measured response diversity. We do not resolve whether the Artificial Hivemind is real. We show that the published evidence does not establish it.
cs.LG / 33 / 2609.33940
Behavioral Monitoring of JEPA World Models with Jacobian Centroids
Thomas Walker, Randall Balestriero, Richard Baraniuk
cs.LG
Abstract
Detecting failures in World Model (WM)-based planning requires monitoring whether the model is behaviorally aligned with the current task, which in turn requires studying its internal representations. Here, we show that centroids---sub-component Jacobian row-sums---effectively identify the behavioral properties of WMs, complementing traditional activation-based knowledge signals. The centroids of a model are easily computed through Jacobian vector products and characterize how the model organizes the geometry of its input space, yielding an efficient perspective on internal representations, including the generation of task-relevant saliency maps. Evaluated on continuous control tasks using JEPA WMs, this behavioral view reveals a structural dissociation, where the encoder correctly represents the goal while the predictor remains behaviorally unresponsive. This failure mode directly predicts planning failure before any action is taken, allowing for goal resampling to recapture out-of-distribution success. Moreover, centroid-based methods outperform baseline methods as distribution-shift detectors. Together, these tools yield a behavioral monitoring stack that is operational and consequential under distribution shifts.
cs.LG / 34 / 2609.33980
DynGraphAgentBench: A Benchmark for Agentic Lifecycle Control in Dynamic Graph Anomaly Detection
Yuwei Han, Lingwei Wei, Wooseong Yang, Liangjie Huang, Liancheng Fang, Huanhuan Ma, Philip S. Yu
cs.LG
Abstract
Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedback. It comprises seven temporal graph datasets with node- and edge-level anomaly tasks, eleven selectable detectors, and eight chronological deployment windows per dataset. In each window, a controller sees only time-causal aggregate context, registered model cards, and its own matured history. It must choose a detector before current-window training or candidate scores exist. A sandboxed executor trains the chosen architecture on mature data, scores a hidden deployment window, and releases the outcome after a one-window delay. A deterministic verifier checks decision timing, leakage guards, legal actions, training scope, and persisted artifacts. We measure detection utility with average precision and capture at fixed review depth, and characterize adaptation through model switches and compute. Complete eight-window trajectories from two primary controllers and a no-memory reference on four datasets, together with three additional controllers on three datasets, expose useful, costly, and ineffective reactions to delayed evidence without granting an exhaustive current-window oracle.
cs.LG / 35 / 2609.33981
Future Information-Directed Sampling for Bayesian Nonstationary Bandits
Yichen Song, Alessio Russo, Aldo Pacchiano
cs.LG
Abstract
Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal arm, which can be insufficient in nonstationary settings where future optimal arms may differ substantially from current ones. In this paper, we propose Future Information-Directed Sampling (FIDS), a new algorithm for Bayesian nonstationary bandits that explicitly explores to gather information about future optimal arms. We show that FIDS achieves regret comparable to Thompson Sampling up to a small constant factor, while being able to exploit predictive information structures that conventional exploration objectives fail to capture. To address the practical difficulty of posterior inference, we further propose a supervised-learning-based approximation framework that learns the FIDS policy from offline data, and demonstrate its effectiveness on synthetic benchmarks.
cs.LG / 36 / 2609.33984
From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting
Chaoqi Zhang, Yu Wang, Haixu Tang
cs.LG
Abstract
Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingular history covariance, the minimum-MSE finite-window linear predictor factors into a Hankel cross-covariance matrix and an inverse Toeplitz covariance matrix. Shared lags and scale cancellation specify this predictor using $H+L-1$ autocorrelations for lookback $L$ and horizon $H$. Building on the innovations representation, our Hankel-Toeplitz Forecaster (HTF) learns one impulse response that defines both an inverse filter and a forecast map. We characterize the finite-history correction and, under summability assumptions, bound the excess risk of truncating the true filters. HTF uses $H+L-1$ trainable coefficients while allowing a full-rank forecasting matrix. Across seven benchmarks at $L=336$, its horizon-averaged MSE is within 1.2% of Dense Linear on each dataset with 75-229 times fewer trainable parameters.
cs.LG / 37 / 2609.33986
ICMAPE: In-Context Multiagent Pure Exploration
Xinyi Hu, Alessio Russo, Aldo Pacchiano
cs.LG
Abstract
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-specified models, while there is currently a gap for practical multi-agent methods that can perform active sequential testing. We fill this gap with ICMAPE, a Bayesian learning-based framework for decentralized multi-agent pure-exploration driven by inference objectives. ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration. It jointly learns a centralized neural inference network that estimates a posterior distribution over hypotheses from global trajectory data, and decentralized policies that select actions from local observation histories and learn when to stop collecting data once the target confidence is reached. On two synthetic benchmarks and a Maryland nitrate concentration monitoring task based on real-world data, ICMAPE-TD3 achieves target accuracy with fewer exploration steps.
cs.LG / 38 / 2609.33993
ASTRA: ADMM-Accelerated Topology Reconfiguration for Dynamic Satellite Constellations
João Norberto, Ricardo Ferreira, Cláudia Soares
cs.LG
Abstract
Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Satellite Topology via Regret-Aware learning (ASTRA), a theoretically-grounded framework for dynamic satellite topology reconfiguration that builds on an online learning formulation and makes it computationally practical. ASTRA combines an ADMM-based offline solver with efficient online updates for both online gradient descent and online conditional gradient, yielding markedly cheaper constrained updates than generic optimization pipelines. On the theory side, we show that for a relevant class of entry-wise nonzero utility matrices, the objective is strongly convex, which yields logarithmic static regret for online gradient descent, and we further instantiate known dynamic-regret guarantees under inexact ADMM inner loops. Empirically, ASTRA matches or improves topology quality, presenting a good trade-off with computational time on synthetic constellations, and it remains effective on real Starlink data under partial deployment and non-uniform spacing, where idealized structural assumptions break down. These results position ASTRA as an efficient and theoretically grounded approach to topology reconfiguration in realistic Low Earth Orbit networks.
cs.LG / 39 / 2609.34002
T-SNN: Temporal Simplicial Neural Network for EEG Decoding
Nikita Malik, Shubhajit Roy, Mohit Kataria, Isuru Herath, Suraj Yadav, Inés García-Redondo, Dhananjay Bhaskar
cs.LG · q-bio.NC
Abstract
Decoding brain states requires models that capture both the evolution of neural activity and interactions among groups of brain regions. Existing EEG methods often treat recordings as multivariate time series or represent functional connectivity with pairwise graphs, leaving dynamic higher-order interactions largely unmodeled. We introduce the Temporal Simplicial Neural Network (T-SNN), which represents EEG recordings as sequences of evolving simplicial complexes. By combining simplicial convolutions with recurrent updates, T-SNN jointly learns higher-order interactions and their temporal evolution. On the seven-class SEED-VII emotion recognition task, T-SNN outperforms convolutional, recurrent, graph-based, and Transformer methods in both trial-wise and cross-subject evaluations. Incorporating eye-movement features further improves performance, demonstrating the framework's potential for multimodal brain-state decoding.
cs.LG / 40 / 2609.34004
RICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock Forecasting
Tong Liu, Lanmiao Liu, Xiang Hu
cs.LG
Abstract
Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, information availability, and transition reliability are modeled. Existing LLM-based financial agents incorporate historical evidence, yet they provide limited support for preserving issuer-specific chronology under point-in-time constraints and for identifying when historical transitions contribute information beyond the current forecast. We present RICE-Alpha (Reliability-Informed Correction with Event Graphs), a point-in-time stock-scoring framework that separates a history-aware multi-view Base Alpha from a reliability-calibrated residual correction derived from historical event continuation. A Multi-Tier Memory Layer grounds news interpretation in temporally eligible issuer-specific history, while a Typed Event Agent constructs event states whose successor relations are formed within issuers and pooled across firms only after valid local pairing. Matured transitions are calibrated by their empirical reliability, and the resulting graph signal is residualized against the Base Alpha and technical view to obtain the RICE Delta. On daily Nasdaq-100 and Hang Seng Index panels from 2024 to 2026, RICE-Alpha achieves the strongest results among the evaluated LLM-based agents and momentum across four predictive and four portfolio-level metrics. Its ICIR more than doubles that of the strongest baseline, while net Sharpe ratios reach 1.656 and 1.725 in the U.S. and Hong Kong, respectively. U.S. ablations further show significant reductions in IC and RankIC after Holm adjustment when major components are removed. These results indicate that historical event continuation adds incremental information when it is temporally grounded, reliability-calibrated, and introduced as a residual correction to a multi-view forecast.
cs.LG / 41 / 2609.34014
LTV-CTDNet: Compositional Turning Decomposition for Short-Term Turning-Movement Forecasting
Md Atiqur Rahman Mallick, Kamrul Hasan, Robert T. White
cs.LG
Abstract
Short-term turning-movement forecasts can support signal control and corridor operations, but unconstrained neural networks may produce physically impossible negative counts or outputs that are not explicitly tied to an approach-demand total. This study introduces the Linear Temporal-Variable Compositional Turning Decomposition Network (LTV-CTDNet), a forecasting framework designed to combine competitive accuracy with structurally admissible outputs. LTV-CTDNet was evaluated using seven months of 15-minute LiDAR observations from eight monitored corridor locations in Nashville, Tennessee. Its lightweight encoder combines recent turning-movement history, weekly time-slot embeddings, and location embeddings. The Compositional Turning Decomposition framework separately predicts nonnegative approach totals and within-approach turning proportions, then reconstructs movement forecasts from these components. Among the evaluated predefined configurations, LTV-CTDNet achieved a movement-level MAE of 1.8189 and RMSE of 3.8072. Its accuracy gains over the strongest sequence models were modest, but it produced no negative forecasts, while unconstrained learned models generated negative values in approximately 10.6% to 29.2% of raw forecast cells. The framework enforces nonnegative outputs and exact agreement between each model-predicted approach total and the sum of its component movements by construction, providing directly interpretable forecasts without clipping or coherence correction.
cs.LG / 42 / 2609.34019
SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making
Shyam Sundar Murali Krishnan, Dean Frederick Hougen
cs.LG · cs.AI
Abstract
In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alternative is to use models that are interpretable by design. However, existing rule-based approaches, such as RuleFit and decision trees, while transparent, often lack stability and predictive strength, reinforcing a perceived trade-off between traditional performance measures and model understandability. To address this, we propose Sparse Relaxed Regularized Regression Rule-Fit (SR4-Fit), an intrinsically interpretable algorithm for both classification and regression that produces compact and stable rule sets without sacrificing performance. Using demographic data from the U.S. Census Bureau's American Community Survey, SR4-Fit predicts U.S. House election outcomes with high accuracy and interpretability while uncovering demographic interactions missed by black-box models. We further validate SR4-Fit across fourteen benchmark datasets (six classification and eight regression), where it outperforms existing rule-based methods, including RuleFit and decision trees in terms of accuracy, stability, and compactness while remaining competitive with black-box models in predictivity. These results demonstrate that interpretability and predictive reliability need not be mutually exclusive, offering a practical and transparent alternative for high-stakes decision-making.
cs.LG / 43 / 2609.34025
Structure-Adaptive Tree Field Integrators
Millend Roy, Soham Samal, Ivan Zelich, Krzysztof Marcin Choromanski
cs.LG · cs.CV · cs.DS · math.NA · math.OC
Abstract
We present a new class of near-linear algorithms for efficiently integrating general tensor fields defined on trees with distance dependent kernels, the Structure-Adaptive Tree Field Integrators (STAD-TFIs). STAD-TFIs exploit the tree's underlying structure through decompositions built around path backbones and single vertex separators, and use two-dimensional fast Fourier transforms to compute interactions jointly. By exploiting this structural information, STAD-TFIs achieve more computationally efficient integration than their regular efficient tree field integrators (TFI) counterparts. We provide a detailed theoretical analysis of our proposed approach and complement it with an exhaustive empirical evaluation, ranging from speed tests on synthetic trees, through accelerated Sinkhorn-based relaxations of the Optimal Transport algorithms on real meshes, to Topological Attention Transformers for vision tasks. To the best of our knowledge, we provide some of the first results showing that efficient to compute and accurate relaxations of the geodesic Sinkhorn-based solutions of the Optimal Transport problem can be derived by applying fast TFI methods.
cs.LG / 44 / 2609.34031
Posterior Regimes and Latent Deception: Variational Bayesian Inference in Hidden Markov Models for Sequential Fraud Detection in Financial Transactions
Joseph Uririoghene Obukofe, Anthony O'Hare, Chioma Sandra Dike
cs.LG · stat.ML
Abstract
We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number of latent behavioural regimes, one of which is empirically identified as fraud-associated. The Neural VBEM HMM replaces the fixed Gaussian-multinomial emission family with a learned encoder, compressing a 741-dimensional transaction representation into a 64-dimensional latent space in which the VBEM HMM's posterior operates; a UMAP projection of this space reveals that the discovered regimes are not discrete clusters but ordered segments of a single continuous behavioural manifold, with confirmed fraud concentrated at its extreme. We show that the model's natural output, that is, the posterior probability of regime membership, is routinely mistaken for a fraud probability, and quantify the resulting miscalibration (the regime-membership interpretation error, MRIE); a corrected posterior-predictive score, closes most of this gap. We further distinguish batch (smoothed) inference, which uses look-ahead unavailable at deployment time, from filtered (forward-only) inference, and report both. On IEEE-CIS transaction data, the neural tier achieves a 14.4$\times$ fraud enrichment in its identified regime; while its AUPRC trails a discriminative XGBoost baseline, we show this gap is structural and not incidental, and argue the model is best positioned as a calibrated triage and interpretability layer rather than a drop-in ranking replacement.
cs.LG / 45 / 2609.34036
UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
Wenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang, Hengrui Cai
cs.LG · cs.AI
Abstract
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8\%$ relative to standard OPD.
cs.LG / 46 / 2609.34041
Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
Samuel Tetteh, Cody Fleming
cs.LG
Abstract
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6\% to 19.4\%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.
cs.LG / 47 / 2609.34054
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction
Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim
cs.LG
Abstract
Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.
cs.LG / 48 / 2609.34056
Steering Language Model Goals with Value Transplant
Pengcheng Jiang, Fabien Roger
cs.LG · cs.CL
Abstract
Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study whether changing such a signal can retarget the model's search toward a different goal. We test value transplant: at each token, we shift the host model's activation along a candidate value axis by the donor-host difference in value coordinates (multiplied by a large scalar), aiming to redirect the host toward the donor's goal. We study this intervention in Qwen3-8B and GPT-OSS-20B models fine-tuned into honest and cheating variants. We test several candidate value axes, including a self-rating axis constructed from activations preceding high versus low elicited self-ratings of progress. The intervention works in both directions, with an honest donor reducing test-gaming in a cheating host and a cheating donor increasing test-gaming in an honest host, showing that this signal can influence which strategy the model follows. On solvable coding tasks, transplant from an honest donor also improves the cheating host's hidden-test performance. Value transplant also works across model families, providing preliminary evidence for the intervention in a setting relevant to model control.
cs.LG / 49 / 2609.34060
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim
cs.LG
Abstract
Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and construct a separate cache with high computation and memory overhead. Selective recomputation reduces this redundancy but still retains substantial model execution, while existing delta correction methods either support only recurring context relations or maintain memory-intensive online correction states for dynamically changing context. For first seen shared context, these methods also construct a reference cache outside the agent workflow, and an approximate correction at the first agent affects the outputs passed to subsequent agents. We present KVCMAS, an online KV cache correction framework that represents cross-agent cache deviations using compact low-rank states and seamlessly chains corrections along the agent workflow without an additional reference prefill. This design supports dynamically changing shared context while preserving an exact first-agent cache. Across multiple language and vision-language workloads, KVCMAS matches or improves the accuracy of prior KV cache sharing methods while achieving the lowest TTFT under highly concurrent serving. Under controlled serving traces, it provides a 2.0x TTFT speedup over inference without KV cache sharing and reduces peak GPU memory by up to 3.7x relative to a prior KV cache correction method. These results establish KVCMAS as an accurate and scalable KV cache sharing approach for prompt-specialized multi-agent serving.
cs.LG / 50 / 2609.34063
Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
Yasuto Hoshi, Daisuke Miyashita, Jun Deguchi
cs.LG · cs.CL
Abstract
Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improves the fidelity of the final model output. We study training-free RESA and learned Top-K+$φ$ with frozen backbone language models. A prespecified single-layer screen yields two Qwen3-0.6B/Multi-LexSum interventions for which direct-runtime measurements show positive prespecified request-aggregate local reconstruction gain but worse final KL fidelity than the corresponding all-abstain Exact Top-K baseline on both discovery and prompt-token-disjoint holdout requests. Exact restoration at the same layer instead improves final fidelity, showing that the reversal is specific to approximate completion in these cases. In complementary multi-layer experiments, a task-independent local diagnostic often repairs the tested completion estimators, although the repaired models do not consistently outperform Exact Top-K. Together, these results show that better local reconstruction need not translate into better final-model fidelity.
cs.LG / 51 / 2609.34070
FLARE: Flow Matching with Local Axis-Angle Representations for Stochastic Micromagnetic Evolution
Pengyu Li, Renjie Tong, Xuanlue Jiang, Jianmin Li, Yuanyuan Zhou
cs.LG · cond-mat.mes-hall · cs.AI
Abstract
Long-horizon micromagnetic simulation remains expensive because conventional and learned solvers typically propagate Landau--Lifshitz--Gilbert (LLG) dynamics step by step. Existing learned approaches generally retain stepwise integration or model deterministic evolution, leaving full-field, direct-horizon stochastic prediction largely unexplored. We propose FLARE, a flow-matching framework that recasts stochastic finite-time magnetization prediction as conditional transport over anchor-relative local axis-angle rotations. This rotation-space formulation respects the intrinsic geometry of magnetization dynamics and preserves pointwise unit norm by construction. By explicitly conditioning on the physical prediction horizon, FLARE directly generates full-field stochastic endpoints across multiple target times without stepwise integration. Against the strongest single-checkpoint external baseline on each metric, FLARE achieves 29.9% lower angular energy distance ($15.30^\circ$), and a 37.3% lower fair energy score (0.393). On a representative composed 5-ns two-segment protocol, FLARE achieves a $3{,}062\times$ best-batch speedup over the widely used GPU micromagnetic solver MuMax$^3$ on a single GPU.
cs.LG / 52 / 2609.34077
MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
cs.LG · cs.AI
Abstract
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
cs.LG / 53 / 2609.34083
Beyond One Epoch: Uncertainty-Weighted Sensitivity Regularization for Recommendation Models
Richard Lettich, Shagun Gupta
cs.LG
Abstract
Recommendation models with sparse embeddings and a shared consumer often exhibit the one-epoch phenomenon: a second epoch lowers training loss while sharply degrading generalization. We present a view based on the violation of the prequential principle. On the first epoch, an example's label has not affected the embedding rows used to score it. On later epochs, those rows contain a displacement induced by the labels earlier update. This creates an incentive for the shared consumer to exploit this displacement in subsequent epochs, which fails to generalize. We call this self-influence asymmetry. Using an exact scalar model and local influence analysis, we connect this mismatch to the uncertainty in the embeddings and the consumers incentive to exploit it in subsequent epochs. We verify this hypothesis using an embedding-consumer-update interventions in deep recommendation models and propose uncertainty-weighted sensitivity regularization (UWSR) which counteracts this mismatch by augmenting the loss function to penalize the consumer for relying on uncertain embeddings. Unlike existing remedies, UWSR preserves the learned embeddings and across three benchmarks, four-epoch UWSR reduces test cross-entropy by 1.38%-6.78% and improves AUC by 0.0058-0.0231 relative to one-epoch training.
cs.LG / 54 / 2609.34088
TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
Lovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold, Minxiao Wang, Runze Yan, Patricia Dykes, Brian J. Gow, Tom J. Pollard, Jessica K. Zègre-Hemsey, Dillon J. Dzikowicz, Lekshmi Kumar, Xiao Hu, Ran Xiao
cs.LG · cs.AI
Abstract
TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-style training, which often struggles with noisy clinical text and fails to leverage the complementary strengths of unimodal (from ECG) and cross-modal (between ECG and matched cardiologist reports) learning. To bridge this gap, we propose a hybrid architecture that jointly learns unimodal and cross-modal representations via uncertainty-weighted multi-task learning while utilizing an LLM-based pipeline to extract high-fidelity findings from cardiologist reports. We evaluate TRACE across a spectrum of clinical urgency, establishing robust performance on public benchmarks for arrhythmia classification and structural abnormalities relative to existing unimodal and multimodal ECG models. To demonstrate real-world utility, we further validate the model on acute coronary occlusion (ACO), where the prevailing ST-elevation criteria miss 25-34% of true occlusions. Utilizing a large private ACO dataset with expert-annotated ground truth, TRACE significantly outperforms real-world clinical practice, yielding a 19.0% increase in sensitivity or a 62.6% reduction in false positive rates at the clinical baseline. This extensive evaluation confirms that TRACE delivers both strong performance on benchmark tasks and tangible clinical impact in the most acute, high-risk cardiac scenarios.
cs.LG / 55 / 2609.34091
When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
Yucong Cao, Chenqi Li, Tingting Zhu
cs.LG
Abstract
Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly changes a latent's activation, one might conclude that the latent represents alpha activity. Across 27 settings spanning three backbones, three EEG datasets, and three network depths, this interpretation initially appears compelling: alpha removal changes latent firing 7.3 times more than an equal-width sham notch (95% CI [6.2, 8.7], bootstrapped over settings). However, the alpha filter also deletes far more signal than the sham. After normalizing by removed spectral energy, the ratio falls to 0.28 (95% CI [0.22, 0.36]) and exceeds one in none of the 27 settings. Latents selected for their response to alpha removal are, on clean EEG, slightly anti-correlated with relative alpha power (mean r = -0.073), giving no support for a simple alpha-detector reading. Motivated by this failure case, we propose a validation ladder for semantic interpretations of SAE latents: it asks in turn whether a latent responds, whether that response survives controlling for how much signal the intervention removes, whether it is specific rather than broadly fragile, and whether the proposed property is visible on unperturbed data--while separately testing the stronger claim that the latent matters to a task classifier. Perturbation sensitivity alone does not establish what an SAE latent represents.
cs.LG / 56 / 2609.34098
Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim
cs.LG
Abstract
Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.
cs.LG / 57 / 2609.34114
Evolution of fairness in multi-objective reinforcement learning framework
Jingyi Zhang, Xin Ou, Guozhong Zheng, Shengfeng Deng, Jiqiang Zhang, Li Chen
cs.LG · cond-mat.dis-nn
Abstract
Fairness, as a fundamental social norm, continues to pose a longstanding puzzle regarding its emergence. Traditional game-theoretic models largely rely on the assumption of \emph{Homo economicus}, wherein individuals are purely rational and self-interested, acting solely to maximize material payoffs. Such accounts, however, overlook the multidimensional nature of human decision-making, which is often shaped also by other considerations beyond economic incentives. To address this gap, we propose a multi-objective reinforcement learning framework that models the evolution of fairness as a dynamic trade-off between material payoff maximization and fairness-driven moral behavior, regulated by a fairness pressure coefficient. Using simulations of a two-objective Q-learning ultimatum game, we find that increased fairness pressure promotes fair outcomes, as expected. Strikingly, however, under moderate pressure, responder behavior reverses: responders become ``forgiving" by accepting low offers -- a pattern in line with our daily experience. Microscopic analyses reveal that this strategy reversal stems from competition between payoff-maximizing and fairness-oriented preferences. We further extend our framework to an asymmetric setting, where proposers and responders assign different weights to the two objectives. Overall, our work expands the reinforcement learning paradigm from a single-objective to a multi-objective formulation, offering a versatile tool for elucidating a broader range of human social behaviors.
cs.LG / 58 / 2609.34117
SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
Gunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin, Baeseong Park, Minsoo Rhu
cs.LG · cs.AI
Abstract
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
cs.LG / 59 / 2609.34120
Probabilistic electrical power demand forecasting with uncertainty quantification
Mahesh Neupane, Pragya Dhungana, Pradip Khatri, Swechhya Baskota, Hariom Dhungana
cs.LG · cs.AI
Abstract
The majority of research on electricity consumption forecasting has focused on deterministic approaches, which generate a single point estimate for each time step in the forecasting horizon. However, the increasing penetration of renewable energy sources and the growing complexity of modern smart grids have introduced greater variability and uncertainty into power-system demand and operation. Consequently, probabilistic forecasting, which quantifies the uncertainty and variability associated with future electricity demand, is becoming increasingly important for reliable power-system planning and operation. This study presents an empirical comparison of four contemporary probabilistic forecasting models for electricity consumption, highlighting their respective strengths and limitations. We have performed comparision on real-world power systems related datasets. Across all power-consumption zones, NGBoost demonstrates superior probabilistic forecasting performance, achieving the lowest MAE and RMSE while providing well-calibrated uncertainty estimates with high prediction-interval coverage and reasonably narrow intervals. These results indicate that NGBoost offers a more accurate and reliable forecasting framework than Bayesian, Monte Carlo (MC) Dropout, and Gaussian Process Regression (GPR) models for the considered electricity consumption data.
cs.LG / 60 / 2609.34123
What Does a Stream Model Buy You in Flow Matching?
Jian Xu
cs.LG
Abstract
Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We ask what such a stream model actually contributes. Three results answer the question. (i)~\emph{Reduction.} The stream-level CFM objective depends on the stream law only through the per-time joint law of $(s_t,\sdot_t)$, so the conditional paths a Gaussian stream can reach are exactly the Gaussian conditional paths CFM already parametrises; in the coordinate-wise, shared-scalar-kernel construction gpcfm actually uses, the entire design space collapses to two scalar curves $(m_t,v_t)$, and cross-time covariance affects only estimator variance. (ii)~\emph{The GP is a constrained chart of that space.} One kernel sets both $m_t$ and $v_t$, so the paper's own recipe for widening coverage-shrinking the SE length-scale---destroys the interpolant (the midpoint mean weight falls from $1.03$ to $0.00$). On the 2-Gaussian benchmark this makes the GP chart diverge on $15/200$ runs at high coverage against $0/200$ for a decoupled $(m_t,v_t)$ chart ($p=6.6\times10^{-5}$), and crossing the two curves shows the divergence tracks the mean, not the variance. On MNIST the same sweep does not diverge and the ordering reverses, so whether the coupling is harmful is benchmark-dependent; what holds on both is that the recipe buys nothing---no coverage level beats the paper's own, and past $\max_t\sqrt{v_t}\approx0.6$ both charts degrade. (iii)~\emph{Audit.} The released code does not implement the mechanism it describes: state and velocity are drawn independently ($\mathrm{corr}=0.00\pm0.01$ against an intended $\pm0.83$--$0.99$).
cs.LG / 61 / 2609.34130
Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
Burc Gokden
cs.LG · cs.CL
Abstract
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs). Exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects. Positive affine blocking retains restarts at the row-constant face, while the augmented AdamW state supplies the complete dynamical description. Predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks, retaining optimizer memory, remaining data, schedule, and numerical policy. Autonomous reductions require closure; approximate reductions carry successor and emission errors. Finite-population covariance, matched physical clocks, matrix fluxes, and signed temporal energy connect row dynamics to model-wide observations. Absolute row collapse, relative row concentration, operator stabilization, and predictive accuracy are distinguished. Experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction. Independent single-pass families exhibit moving finite fluctuation regions without establishing a thermodynamic critical class. Conditional symmetry, head limits, covariance flows, and readout error budgets specify assumptions needed to transfer scaling laws to inference. The theory separates exact identities, conditional dynamical claims, and finite empirical findings, with proofs, selected formal checks, and compact numerical evidence.
cs.LG / 62 / 2609.34146
ExpertoRhythm: Morphology-Aware Learning for Waveform Reconstruction and Cuffless Blood Pressure Estimation from Single-Channel PPG
Amir Arjomand, Kenneth B. Kent, Georgiy Krylov
cs.LG
Abstract
Continuous cuffless blood pressure (BP) monitoring from photoplethysmography (PPG) has strong potential for wearable health and telemonitoring, but accurate estimation remains difficult because PPG-to-BP mapping must preserve subtle waveform morphology and pressure-range-dependent dynamics. We introduce ExpertoRhythm, an attention-enhanced 1D U-Net that reconstructs the arterial blood pressure (ABP) waveform from a single-channel PPG signal and derives systolic and diastolic BP directly from the reconstructed waveform. The central contribution is a composite morphology-aware learning objective that integrates range-weighted SmoothL1 reconstruction with a window-range regularizer to emphasize high-dynamic BP segments and reduce amplitude under/over-shoot. On the UCI cuff-less BP dataset with 942 subjects, ExpertoRhythm achieves 2.46/1.46 mmHg MAE for systolic/diastolic BP (SBP/DBP), while obtaining a 30.4% average relative error reduction over pure MSE across waveform reconstruction and BP estimation metrics. Clinical-style evaluation further demonstrates low bias and strong agreement across the BP range, including high-pressure windows up to 200 mmHg, satisfying AAMI criteria and achieving BHS Grade A. These results suggest that morphology-aware waveform reconstruction from a single PPG channel can provide an accurate and practical pathway toward continuous cuffless BP monitoring in wearable and remote-care settings.
cs.LG / 63 / 2609.34153
SPINET: Sheaf Protein Inverse Folding Network
Jens Lundsgaard, Colin Mikulski, Zhixuan Yan, Dhananjay Bhaskar
cs.LG · q-bio.BM
Abstract
Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.
cs.LG / 64 / 2609.34156
Transfer Calibrated Prediction Powered Inference
Aditya T. Vadlamani, Jae Ho Chang, Srinivasan Parthasarathy, Subhadeep Paul
cs.LG
Abstract
Prediction-powered inference (PPI) and its power-tuned extension (PPI++) improve confidence intervals by combining a small gold-standard labeled sample with a large AI model's predictions. Its efficiency gain relies on low residual variance, which may not hold if the predictor is pre-trained on a different source domain. We propose Transfer Calibrated Prediction-Powered Inference (TC-PPI), adapting the source-domain predictor to the target domain using gold-standard samples through cross-fitting. This approach supports various adaptation methods, such as sparse linear calibration, LoRA, and fine-tuning. Our jointly tuned cross-fit estimator, Joint-TC-Cross-PPI++, maintains unbiasedness and is simultaneously at least as efficient as classical inference, PPI, and PPI++, thereby protecting against negative transfer. We provide high-dimensional MSE bounds for calibration and show empirical improvements over baseline methods across various real-world applications.
cs.LG / 65 / 2609.34159
WorldGraph: Graph-Native World Modeling
Zezhong Ding, Yipeng Li, Xike Xie
cs.LG · cs.AI · cs.CV
Abstract
World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their properties change over time. Prior graph-related world models use graph structures to organize internal states or support task-specific reasoning, rather than treating an evolving graph itself as the modeled world. We instead study graph world modeling (GWM), where graph evolution itself constitutes the world dynamics. We formulate graph world modeling over observed graph evolution, latent graph states, and heterogeneous graph-transition predictions. Based on this formulation, we construct GWM-Zero, a benchmark covering node-, edge-, and graph-level transitions over eight temporal graph datasets. We propose WorldGraph, which combines a state-aware graph transformer for multi-granularity structural and transition-conditioned evolution modeling with transition-aware GRPO using dynamic grouping and structure-aware verifiable rewards. Extensive experiments on GWM-Zero show that WorldGraph consistently outperforms representative graph representation, temporal graph learning, graph pretraining, and graph world-model baselines across all three transition granularities.
cs.LG / 66 / 2609.34166
Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations
Marco Armenta
cs.LG · cs.NE
Abstract
We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair $(W,f)$, a thin representation $W$ of its quiver and an activation $f$; its function factorizes through the space of quiver representations, each input $x$ inducing a representation, and the knowledge matrix $M(x)\in\mathbb{R}^{C\times(d+1)}$ is the contraction of that representation to one matrix whose rows sum exactly to the logits. At one trained network we ask what determines it, what it is invariant to, what it determines, and what its geometry measures. Under (LCS), a locally constant slope diagonal, as for ReLU, the matrix at a regular input is a function of the realized germ; its stabilizer among encodings regular there is exactly the germ stabilizer at inputs with no vanishing coordinate, neuron permutation a special case; and it recovers the germ, whereas hidden activations, gauge-covariant and germ-incomplete, are not enough. Under (LCS) it equals per-class gradient$\times$input plus an exact aggregate bias attribution, grounding it in attribution theory and computing it by $C$ vector-Jacobian products instead of probing. The fixed shape gives an alignment-free per-sample distance between ResNet-152, DenseNet-121 and GoogLeNet; the row-sum identity gives an exact visible/invisible displacement decomposition whose unit-free coherence $A=(d_Ψ/d_M)^2$ puts adversarial germ motion at median $A\le 0.23$, with an attack-family ordering concordant across six architectures (Kendall $W=0.921$; $0.97$ on the three networks at full scale). Two honest negatives: on AlexNet/CIFAR-10 penultimate features win 5 of 6 detectors and all 16 attacks, and a matrix-direction counterfactual fails 0/54.
cs.LG / 67 / 2609.34174
GradLev: Token-Parallel Test-Time Training Via Costate Prediction
Bo Liu, Qiang Liu
cs.LG · cs.AI
Abstract
Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.
cs.LG / 68 / 2609.34191
Cardinality-Stratified Interaction Decomposition for Interpretable Pairwise and Higher-Order Structure in Transactional Basket Data
Hidetoshi Kawase, Toshihiro Ota
cs.LG
Abstract
Transactional basket data can reveal associations among items, but observed co-occurrence conflates item-specific relations with basket-size structure and unmodeled higher-order dependence. We introduce Cardinality-Stratified Interaction Decomposition (CSID), an interpretable framework that decomposes log-odds contrasts stratified by the number of remaining items into item-set-specific and cardinality-common components, without fitting a global joint distribution. CSID uses an information-weighted, gauge-constrained ridge projection to estimate pair and triple components and to diagnose higher-order contributions to pairwise structure. CSID is designed primarily for interpretable decomposition of association structure rather than for full-distribution prediction. In a simulation with zero pair effects, increasingly strong small-basket cardinality potentials drive ordinary Ising couplings spuriously negative, whereas CSID pair estimates remain centered near zero. Detection power rises with the magnitude of planted triple effects, and local deprojection reduces pair-coefficient RMSE from 0.244 to 0.073. Across three grocery datasets, high-information triple components are reproducible over time. In the matched cross-period partial-transfer evaluation, transferred CSID triple components show closer agreement with later-period stratified contrasts than the nodewise-symmetrized cardinality-aware higher-order pseudolikelihood comparator, with gains in weighted Lin's concordance correlation of 0.038--0.122. These results support CSID as an exploratory and interpretable decomposition framework for pairwise and higher-order association structure in transactional data.
cs.LG / 69 / 2609.34198
Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
Jiapeng Li
cs.LG
Abstract
Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned agent-task cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare improvements that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge confidently reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and reference outcome, while a task-solvability prediction from AgentRewardBench reverses sign in SWE-bench. Transporting old-version calibration raises mean absolute comparison error on SWE-bench from 3.8 to 19.5 percentage points; 24.6% of ratio-bootstrap draws are undefined near the correction boundary. A tuned paired audit narrows a classical interval by only about 5% at 80 labeled tasks. A randomized three-arm test does not support the predicted increase in false acceptance from showing the agent's final report (all three Holm-adjusted p-values = 1.0). These results favor explicit reference standards and paired audits of current outputs over judge-only release decisions or transported old-version calibration.
cs.LG / 70 / 2609.34212
X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths
Bowen Dong, Yilong Fan, Tengyu Pan, Yike Zhang, Zhenyu Li, Zijian Zhang, Xuewei Li, Mei Yu, Jianyong Wang
cs.LG · cs.CL
Abstract
Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.
cs.LG / 71 / 2609.34218
Loop Dropout: Regularizing Shared Updates in Looped Language Models
Zirui Zhu, Hailun Xu, Xuanlei Zhao, Yong Liu, Yingxuan Ren, Kanchan Sarkar, Kun Xu, Yang You
cs.LG · cs.CL
Abstract
Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update is more effective at later loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation.
cs.LG / 72 / 2609.34264
Rotated Manifold Optimization for Low-Rank Adaptation
Yuhui Ding, Javier Zazo, James Hensman
cs.LG
Abstract
We propose a novel optimizer for low-rank adaptation (LoRA) that explicitly incorporates the gauge symmetry of low-rank factorization. Our optimizer extends recent matrix optimizers for full-parameter training to the manifold of fixed-rank matrices by interpreting them as normalization under a rotated basis. We show how rotation and normalization can be integrated with the fixed-rank manifold efficiently. Our optimizer converges faster to lower held-out loss and achieves better or comparable downstream performance on both supervised finetuning and reinforcement learning tasks.
cs.LG / 73 / 2609.34265
ZeroCode: On-demand Error-Correcting Code Construction from the Zero Matrix via Reinforcement Learning
Ju-Hyeong Lee, Yongjune Kim, Sang-Hyo Kim, Dae-Young Yun, Hee-Youl Kwak
cs.LG
Abstract
Error-correcting codes (ECCs) are essential across diverse applications, from wireless communications and storage to quantum computing, yet each application imposes distinct design requirements on the parity-check matrix (PCM). To address these on-demand requirements in a unified framework, we propose ZeroCode, a reinforcement learning (RL)-based approach that constructs PCMs sequentially from the all-zero matrix. ZeroCode formulates construction as a discrete sequential decision-making problem and uses proximal policy optimization with action masking to select valid edges. ZeroCode achieves a gain of approximately 1 dB over the prior RL-based construction method at a bit error rate (BER) of $10^{-4}$ for the (32,16) code and outperforms existing genetic, differentiable, and classical code-design methods in our experiments. Beyond optimizing decoding performance, the masking mechanism allows on-demand structural constraints, such as a maximum degree, 4-cycle-free structure, and quasi-cyclic structure, to be flexibly incorporated. Moreover, a single policy rollout yields a library of PCMs with varying edge counts, offering trade-offs between decoding performance and complexity without retraining. Overall, ZeroCode addresses diverse code-design requirements within a unified framework, providing solutions with optimized decoding performance under given constraints.
cs.LG / 74 / 2609.34272
Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
Junlin Chen, Daize Dong, Huanwei Di, Haolong Jia, Jiawei Wu, Haotian Xie, Mingkai Zheng, Yang Li, Leshang Chen, Huishu Wang, Eric P. Xing, Hongyi Wang
cs.LG · cs.DC · math.NA
Abstract
BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a problem: training was healthy for 25B tokens, then the gradient norm grew a thousandfold and the loss ended 0.2 nats above FP32 attention, without a single NaN. Recomputing the attention backward of just two layers in FP32 removes almost all of the excess gradient. Part of the cause is known: a fused multiply-add in the forward softmax, so far treated as an extreme-input NaN case and never fixed in FlashAttention-3. Repairing it stops the blow-up, but the query gradient is still wrong by more than its own size, and training still drives attention logits to thousands of times their size under accurate gradients. The remaining error comes from a broken conservation law. The softmax score gradient sums to zero along every row, which makes the query gradient blind to where the keys sit as a group; rounding it to BF16 leaves a small nonzero sum that leaks the mean key into the gradient, and the leak grows exactly as late training makes keys large and attention sharp. We introduce GProj (gauge projection), which restores the zero sum after the cast with two rank-one corrections per row. It cuts the remaining median query/key gradient errors from 219%/13% to 0.34%/0.37%, on par with FP32 attention, for 4.7% more time per training step. In matched from-scratch runs it trains to the same loss as FP32 attention, while FlashAttention-3 and key smoothing both destabilize.
cs.LG / 75 / 2609.34279
Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training
Yuyang Deng, Yu Wang, Jiayun Wang
cs.LG · cs.AI
Abstract
Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.
cs.LG / 76 / 2609.34281
Agentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided Search
Zhixuan Gao, Ke Xue, Rongxi Tan, Ming Chen, Chao Qian
cs.LG
Abstract
High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on low- and moderate-dimensional problems. We ask whether this paradigm can transfer to the higher-dimensional regime. Our experiments show that these methods do not remain reliable in the high-dimensional regime, where the challenge is not only where to evaluate, but also which modeling assumption and search geometry to use when the objective's useful structure is unknown. We therefore introduce HERA, a Hypothesis- and Evidence-guided Research Agent that uses task context, optimization feedback, and structural diagnostics to revise search hypotheses, select and configure HDBO strategies, and determine their execution length. PRISM, its numerical optimization engine, generates and evaluates candidates sequentially within each search block, updating numerical models after each observation. HERA remains competitive with strong numerical HDBO baselines and outperforms the evaluated LLM-based and agentic methods on four metadata-free synthetic functions. Across eight real-world tasks, HERA achieves the best mean final objective among all evaluated systems on most benchmarks. Further analyses show that structural diagnostics change strategy use, metadata effects vary across tasks, and adaptive search blocks reduce inference cost.
cs.LG / 77 / 2609.34285
Epistemic Learning from Imprecise Annotation
Kaizheng Wang, Siu Lun Chau
cs.LG
Abstract
Imprecise annotations may support several plausible labelling distributions, yet learning methods often resolve this ambiguity into a single predictive distribution. This can obscure what the annotation evidence leaves unresolved. We introduce epistemic learning from credal supervision, a framework that uses convex sets of plausible labelling distributions, called credal sets, as supervision and learns sets of predictive distributions. We instantiate the framework with the pessimistic--optimistic credal classifier (POCC), which combines a shared backbone with two classification heads trained to minimise worst-case and best-case losses over the supervision sets. Their outputs define a predictive credal set whose spread provides an uncertainty score. We also show how credal labels can be obtained through a simple relaxation of existing probabilistic labels, reducing commitment to their precise probability assignments. This construction admits closed-form inner optimisation under cross-entropy loss, enabling efficient training. Assuming the supervision sets contain the true conditional label distributions, and other regularity assumptions, we establish a finite-sample generalisation bound for the averaged predictor with an explicit penalty for supervision imprecision. We evaluate POCC using human annotator disagreement and teacher predictions, alongside label smoothing as a controlled proxy for annotation imprecision. Across these settings, POCC achieves a favourable balance of predictive accuracy, calibration, and uncertainty-based selective classification versus competitive baselines.
cs.LG / 78 / 2609.34310
Riemannian Difference-of-Convex Optimization for K-Means Clustering
Meng Xu, Bo Jiang, Hanfu Zhang, Ya-Feng Liu, Anthony Man-Cho So
cs.LG · math.OC
Abstract
K-means is a widely adopted clustering approach in signal processing and machine learning. In this paper, we study K-means clustering through a cardinality-constrained formulation on a compact embedded submanifold. We replace the cardinality constraint with a difference-of-convex (DC) penalty and establish a global error bound to prove that the penalized and constrained formulations share the same global minimizers whenever the penalty parameter exceeds a finite threshold. To solve the resulting nonsmooth Riemannian DC problem, we reformulate it as a minimax problem and propose RADA-DC, a Riemannian alternating descent ascent method combining dual regularization with DC linearization. Under standard assumptions and suitable parameter choices, RADA-DC finds an $ε$-Riemannian critical point within $O(ε^{-3})$ iterations. We conduct experiments on synthetic and real-world datasets to demonstrate that the proposed method outperforms the tested baselines, including K-means++, in solution quality at competitive computational cost when the number of clusters is large.
cs.LG / 79 / 2609.34321
One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
Ziqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang, Zihuan Jiang, Linqiang Guo, Siobhan Reid, Zhi Liu, Yang Wang
cs.LG
Abstract
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.
cs.LG / 80 / 2609.34343
The Composition Gap in Dataset Distillation
Guang Li, Takahiro Ogawa, Miki Haseyama
cs.LG
Abstract
Dataset distillation compresses a training set into a small synthetic set, usually evaluated one at a time. In federated and data-governance settings, several parties distill their own data and a user trains on their union. We ask whether the union of separately distilled sets reproduces training on the union of the real data composability and show that it can fail even when every source is distilled exactly and the total budget admits an exact joint distillate. Compressing a training trajectory into fewer steps transforms the source statistics nonlinearly, so averaging compressed sources differs from compressing their average. For quadratic objectives we derive the exact composition error for two-to-one step compression in terms of the source-Hessian variance and the linear terms of the losses, and on a smooth network at small step sizes this prediction captures the local endpoint discrepancy in magnitude and direction. For learned synthetic sets, however, the composed error decomposes exactly into this local discrepancy and an aggregate source residual. Under endpoint matching the residual exceeds the structural term by more than an order of magnitude, and under distribution matching the two terms partly cancel. Joint distillation also retains an accuracy advantage when both sets are distilled from the same dataset, where the local discrepancy is exactly zero. Training fidelity and downstream accuracy are therefore distinct requirements, neither established by evaluating each set on its own.
cs.LG / 81 / 2609.34348
Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
John Sweeney
cs.LG · cs.AI · cs.CL
Abstract
Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $η$ on each of sources $A$ and $B$, the weight difference $θ_{AB}-θ_{BA}$ is, to leading order, $η^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $θ_{AB}-θ_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.
cs.LG / 82 / 2609.34354
FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model
Shucheng Liu, Changchun Shi, Kai Zhang, Hongtu Zhu
cs.LG · q-bio.NC
Abstract
Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.
cs.LG / 83 / 2609.34358
FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents
Xi Xiao, Yunbei Zhang, Chen Liu, Lin Zhao, Jialin Chen, Tianchen Zhao, Xiang Xu, Youngeun Kim, Tianyang Wang, Min Xu
cs.LG · cs.CL
Abstract
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.
cs.LG / 84 / 2609.34370
Spexis: Speculative Lookahead Scheduling for LLM Inference
Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim, Jinho Lee, Jiwon Seo
cs.LG · cs.DC
Abstract
Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.
cs.LG / 85 / 2609.34373
Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
Mihir Chauhan, Aniket Bera
cs.LG · cs.MA · cs.RO
Abstract
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.
cs.LG / 86 / 2609.34374
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff
cs.LG · cs.AI · stat.ML
Abstract
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at https://github.com/SewoongLab/abc-align .
cs.LG / 87 / 2609.34375
LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao
cs.LG · cs.AI · cs.RO
Abstract
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.
cs.LG / 88 / 2609.34398
GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
Moshe Eliasof, Eldad Haber
cs.LG
Abstract
Critical mineral discovery is a positive-only problem: deposits are observed as sparse locations, while unlabeled regions are not reliable negatives, and similar geophysical signatures can arise from different subsurface states. We therefore model mineral targeting as learning a conditional spatial distribution over occurrence locations, $π(p\mid d)$, given geo-images $d$, rather than predicting a deterministic per-pixel score map. We introduce GeoCFM, a conditional flow-matching model that generates mineral occurrence point sets conditioned on multi-channel geo-images; GeoCFM learns a point-wise transport field in $\mathbb{R}^2$, using UNet features with point-conditioned velocity prediction to bridge dense rasters and sparse supervision without pseudo-negatives. On a synthetic magnetics--geochemistry benchmark with latent activation and on USGS Earth MRI data with a spatially disjoint tile split, GeoCFM improves geometric agreement with observed occurrences over score-map and non-conditional baselines, while representing epistemic uncertainty through conditional sampling.
cs.LG / 89 / 2609.34409
MASCIT: A Mask-Aware State Space Classifier for Naturally Irregular Time Series
Yoo-Min Jung, Hyeon-Gi Kim, Jonghun Park
cs.LG · cs.AI
Abstract
Naturally irregular time series combine asynchronous observations, missing values, unequal lengths, and nonuniform sampling, while dense adapters can discard temporal structure. We propose a mask-aware state space classifier for irregular time series (MASCIT), which supplies observation masks to the encoder and excludes invalid steps from gated temporal aggregation. Across 34 irregular time series datasets, MASCIT yielded the strongest aggregate point estimate and was the only evaluated neural model with three-seed results on every dataset. MASCIT retained the lowest point rank across six overlapping irregularity indicators, while factorial ablations favored partial over full selectivity. These results support selective state space models as effective, executable backbones for naturally irregular time series classification.
cs.LG / 90 / 2609.34420
MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism
Tong Qiao, Ao Zhou, Yingjie Qi, Chunming Hu, Jianlei Yang
cs.LG
Abstract
Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8\% and achieves up to 4.51$\times$ training speedup while maintaining model accuracy.
cs.LG / 91 / 2609.34422
Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
Lirui Luo, Kelong Mao, Heming Xia, Rongqing Li, Xinwei Yang, Luyu Chen, Kieran Wong, Yudong Guo, Xinrui Wang, Jiayin Zhu, Simiu Gu, Sulong Xu, Cong Fang
cs.LG · cs.AI · cs.CL
Abstract
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
cs.LG / 92 / 2609.34423
On the Relation Between Interval Regret and Dynamic Regret
Yi-Han Wang, Peng Zhao, Zhi-Hua Zhou
cs.LG · math.OC · stat.ML
Abstract
Non-stationary online learning has attracted much attention in recent years, as static regret is insufficient to guide algorithm design in changing environments. To address this limitation, interval regret and dynamic regret have been introduced as two representative performance metrics that strengthen static regret in complementary directions. Interval regret requires an online algorithm to achieve competitive static regret over every local time interval, whereas dynamic regret evaluates performance against an arbitrary sequence of time-varying comparators. Despite their importance, the relation between these metrics has long remained unclear. Prior work has often regarded interval regret as the stronger notion, based on the intuition that local guarantees should naturally induce global guarantees. Consequently, it is widely conjectured that an algorithm with optimal interval regret should automatically attain optimal dynamic regret. In this paper, we first establish a negative result that refutes this intuition of a metric-level implication. Specifically, for both convex and curved functions (including exp-concave and strongly convex functions), we show that there exist instances in which an algorithm with optimal interval regret nevertheless fails to achieve optimal dynamic regret. We then show how to leverage local adaptivity to obtain optimal dynamic regret. In particular, optimal dynamic regret can be attained by invoking an interval regret minimization process over an enlarged Euclidean ball containing the original convex feasible domain and using a suitable domain-converted surrogate loss. This reduction applies to both convex and curved functions. As a byproduct, we obtain the first proper and efficient algorithm with optimal dynamic regret for exp-concave functions, improving prior results while significantly simplifying the analysis.
cs.LG / 93 / 2609.34426
Q-learning Penalized Transformer for Safe Offline Reinforcement Learning
Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, Dacheng Tao
cs.LG
Abstract
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds.
cs.LG / 94 / 2609.34427
LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization
Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng, Dongdong Ge, Yinyu Ye
cs.LG · cs.CL
Abstract
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.
cs.LG / 95 / 2609.34442
Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths
Jialu Wang, Peizhi Niu, Haoteng Yin, Hans Hao-Hsun Hsu, Pan Li, Rongzhe Wei
cs.LG
Abstract
While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact. To achieve true forgetting, we propose a general deep unlearning framework compatible with existing unlearning algorithms. Our approach adaptively explores both explicit responses and latent internal representations to discover valid reasoning paths, compiles them into a confidence-aware supporting subgraph, and we apply a graph minimum cut to sever all recovery paths while preserving unrelated knowledge. To rigorously evaluate deep unlearning, we introduce a model-specific pipeline that extracts and completes knowledge graphs from raw text, filtering them by calibrated model confidence to reflect what the model genuinely retains. Comprehensive experiments demonstrate that selectively unlearning supporting knowledge yields substantially deeper forgetting than superficial methods while preserving model utility, highlighting that genuine unlearning requires breaking the relational structures that enable factual reconstruction.
cs.LG / 96 / 2609.34446
Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
Rostislav Makarov, Tal Peer, Danilo de Oliveira, Timo Gerkmann
cs.LG
Abstract
Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly from paired data. We target a setting that spans both regimes: unknown degradations can be learned from few paired examples, while known degradations can be learned from self-generated samples. We introduce Likelihood Score Approximation (LSA), a generative framework that keeps a pretrained unconditional model fixed and learns an observation-conditioned model that approximates the likelihood score from paired samples. Within a conditional stochastic-interpolant framework, LSA can be trained in either score or velocity coordinates, independently of the unconditional model's native parameterization, and supports both deterministic and stochastic sampling. We further show empirically that the prior model can be swapped post-training while keeping the same LSA model. Across speech and image inverse problems, LSA operates effectively even at roughly 0.01% of the full training dataset. On the ImageNet-256 benchmark it achieves competitive or better restoration quality than strong posterior-sampling baselines while requiring up to several orders of magnitude fewer network evaluations.
cs.LG / 97 / 2609.34453
SPACE-LoRA: Allocating Activation-Subspace Protection for Continual Learning
Seunghyun Yoo, Kiseok Kim, Hyeontae Joo, Junyeop Bang, Hwangnam Kim
cs.LG · cs.AI
Abstract
This study addresses the catastrophic forgetting problem that occurs when sequentially learning successive tasks using Low-Rank Adaptation (LoRA) from a lifelong learning perspective. While existing approaches have primarily constrained parameter updates or learning subspaces to reduce interference with past knowledge, they have not fully considered additive interference. This occurs when a newly added residual adapter on top of a fixed past model generates non-zero responses along input directions important for old tasks, thereby altering previous predictions. To this end, we propose Subspace Protection with Allocated Capacity for Efficient Continual Adaptation (SPACE-LoRA). SPACE-LoRA directly suppresses the responses of the new residual branch along input activation directions that are important for old tasks and adaptively determines the protection coverage for each module based on past-task sensitivity estimated via a common Fisher sensitivity-based coverage target. Under a fixed LoRA rank, this approach adaptively adjusts module-specific protection coverage while suppressing interference along input directions sensitive to old tasks. We assess the effectiveness of activation-subspace protection in mitigating catastrophic forgetting and examine the role of sensitivity-guided protection in continual learning across diverse tasks. Code is available at https://anonymous.4open.science/r/SPACE-LoRA-7864.
cs.LG / 98 / 2609.34457
ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models
Hai Duong, Thanh Le, ThanhVu Nguyen
cs.LG · cs.SE
Abstract
Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and provide auditable guarantees before deployment. However, prior work remains limited to small or restricted Transformers, and maintaining precision across deep models remains challenging. In this work, we introduce ZonoGPT, an abstract domain for verifying large transformers that maintains a space complexity independent of network depth. ZonoGPT uses a structured zonotope and a generator reduction mechanism to efficiently preserve correlations. To maintain precision, it introduces block-specific fused transformations for Attention and LayerNorm that retain feature relations, along with an affine transform for GELU that preserves generator relations. These mechanisms enable \tool{} to be the first approach to verify standard architectures, scaling to official HuggingFace models up to GPT-2 Medium (24 blocks, 300M+ parameters) and successfully verifying 1,339 instances across text and vision tasks.
cs.LG / 99 / 2609.34466
PhysioTRACE: Provenance-Aware Stress Tests for Physiological Foundation Models
Ayana Mussabayeva, Anuar Aimoldin, Olivier Oullier, Xue Liu, Kun Zhang
cs.LG
Abstract
Physiological foundation models encode how a signal was recorded alongside the physiology it reflects. When recording conditions are associated with diagnosis, this acquisition provenance can become a shortcut, yet the usual evidence, shifted transfer and provenance decodability, does not show whether a predictor uses it. We introduce PhysioTRACE, a four-axis behavioral audit for frozen encoders that separates what a probe can decode from what a fixed task head relies on. Recover scores how decodable provenance is; Stress reverses only the provenance-target association on the same held-out records; Intervene removes a train-localized provenance component; and Verify certifies that removal only if it beats matched random projections within a declared utility margin. Each audit thus ends in one of three verdicts: no reliance, or reliance with the remedy certified or refused. Across EEG and ECG, five training objectives, and five frozen foundation models, the relation between Recover's calibrated score and out-of-distribution utility changes sign between datasets, so neither can stand in for a reliance test. On paired EEG views where the shortcut is known by construction, the audit detects it (the exposed head loses about 0.2 AUROC when the association is reversed, while a control head is unaffected) and certifies removal of a rank-two component that restores control-level behavior without measurable utility loss, for both encoder objectives tested. On real ECG device metadata it returns all three verdicts: it certifies a remedy that removes 91% of one model's excess vulnerability, finds no reliance where device and diagnosis are barely associated, and refuses the remedy for a second model whose localized direction also carries task signal. Robustness to how inputs were recorded therefore needs a behavioral test, and PhysioTRACE provides one that can pass, fail, or refuse a remedy.
cs.LG / 100 / 2609.34474
Deep kernel hedging
Jean-Loup Dupret, Donatien Hainaut, Edouard Motte
cs.LG · math.FA · q-fin.CP · q-fin.RM
Abstract
We introduce a deep kernel hedging framework that combines the flexibility of deep learning with the structural inductive bias of kernel methods. The hedging functional is restricted to a reproducing kernel Hilbert space whose kernel is parameterized through a neural network embedding of the input features. The framework minimizes a regularized empirical risk under convex loss functions and can accommodate path-dependent information through truncated time-augmented signature features. We derive a generalized representer theorem for the joint hedging problem, reducing the empirical optimization to a finite-dimensional problem. To further reduce the computational cost associated with large kernel matrices, we develop a scalable random Fourier feature approximation and establish convergence guarantees. The random Fourier parameters are sampled once and remain fixed throughout training, while the deep kernel adapts to market data through the learned neural representation. We evaluate the performance of the proposed deep kernel approach on both synthetic and real data and compare it with standard kernel methods and classical deep hedging architectures. Numerical results indicate competitive and robust hedging performance, particularly in low-data regimes, which highlights the benefits of combining expressive neural representations with the inductive bias of kernel methods.
cs.LG / 101 / 2609.34475
Causal Routing for Unlearning
Bardh Prenkaj, Andrea D'Angelo, Davide Mottin, Federico Fontana, Davide Gabrielli, Paola Velardi, Stefano Faralli
cs.LG · cs.AI
Abstract
LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to change one thing, and none of them say which part produced that change. To address this, we introduce Causal Routing for Unlearning (CRU) by asking where the concept is expressed in the model and suppressing only that part. One untrained forward pass over the forget set ranks neurons by how their activations vary. Then, small routing modules on those neurons gate and suppress only the concepts that need to be forgotten. In CRU, the base model is frozen, and any change in behavior is caused only by the gated neurons; hence, why the routing is causal. Due to our parameter efficiency (only ~0.01% as many parameters as the base model), unlearning a concept costs 14 GiB, whereas the baselines require 71 GiB. On TOFU, CRU is indistinguishable from the retained model (p > 0.05, KS test) and is never Pareto-dominated, whereas every compared baseline matches its forgetting on the larger-forget batches only by collapsing utility. On RWKU, it achieves an adversarial-probe recall of 0.052, compared to 0.250 for the strongest baseline, meaning the knowledge is gone, not merely harder to reach. Thus, deciding on the intervention at query time, rather than fixing it beforehand, is the axis along which we argue that unlearning should proceed.
cs.LG / 102 / 2609.34478
Learn Here, Move Less Elsewhere: Input-Conditioned Plasticity from Retained-Domain Activation Atlases
Jiangtao Lin, Bangyang Wei, Yihang Ding, Siyi Liu, Yuhan Dong
cs.LG · cs.AI
Abstract
Task-specific fine-tuning can rewrite a language model's answers beyond the training task, complicating updates that must preserve existing behavior. We introduce ATLAS, which turns retained-domain representations into an input-dependent rule for task adaptation. An activation atlas supplies local reference centers and directional filters to a shared low-rank residual. Target supervision learns the residual, while retained geometry shapes its action throughout training and inference. On Qwen3-8B, ATLAS achieves lower mean retained-output Kullback-Leibler (KL) divergence than all seven published baselines at shared coding-performance requirements, with consistent advantages across multiple training seeds. Structural comparisons identify the contributions of retained reference states and directional conditioning, and answer-level analyses show fewer rewritten mathematical answers and more stable commonsense choices. Experiments spanning five backbones and two retained domains further demonstrate coding gains with reduced retained-output movement. With compact storage and modest decoding overhead, ATLAS provides a practical mechanism for acquiring specialized skills while maintaining continuity in existing responses.
cs.LG / 103 / 2609.34488
FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL
Xun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang
cs.LG · cs.AI
Abstract
Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on $30$ online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to $\bf{43\%}$ and achieves up to $\bf{1.34\times}$ wall-clock speedup in a stress test.
cs.LG / 104 / 2609.34489
When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation
Minjae Park, Taesun Yeom, Jaeho Lee
cs.LG
Abstract
Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering. Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization. Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.
cs.LG / 105 / 2609.34493
LLN: Learnable Lens Networks for Parameter-Efficient Long-Horizon Dynamical Prediction
Binbin Yong, Zhao Su, Lan Guo, Haoran Li, Jun Shen, Qingguo Zhou
cs.LG
Abstract
Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, while leaving the evolution of feature geometry across layers largely unconstrained. We introduce Learnable Lens Networks (LLN), a physics-inspired architecture that replaces direct feature-space residual accumulation with learnable optical transport in an augmented position-angle phase space. Each layer alternates between free propagation, which provides an implicit transport path, and a learnable lens field that performs nonlinear trajectory transformation and focusing. Theoretically, we establish that LLN transport is globally invertible and volume-preserving for any differentiable lens field, with the implemented coordinate-wise Gaussian transport further satisfying symplecticity. Importantly, these structural constraints do not limit expressivity: with unrestricted embeddings and readouts, LLN retain universal approximation of continuous end-to-end maps. Experiments across diverse dynamical systems demonstrate that LLN improves long-horizon prediction while using substantially fewer parameters than same-depth comparators. Further analysis reveals stable depth-wise gradient transport and interpretable learned dynamics under the coupled propagation and refraction design.
cs.LG / 106 / 2609.34497
QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou, Jiashu Hou, Ye Shi, Yanwei Fu
cs.LG
Abstract
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.
cs.LG / 107 / 2609.34499
Scalable GNN-based Knowledge Graph Representation Learning with Efficient Message Passing
Huu Tan Mai, Cuong Xuan Chu, Heiko Paulheim, Daria Stepanova
cs.LG · cs.AI
Abstract
Graph neural networks (GNNs) excel at representation learning on Knowledge Graphs (KGs), achieving stateof-the-art performance on tasks like link prediction or entity classification. However, their high computational complexity, inherent to their user-defined message passing (MP) algorithm, still prohibits their widespread adoption, especially for large KGs. Current efforts to mitigate the scalability bottlenecks of GNNs on KGs, such as subgraph sampling, are often task- and model-specific, and do not reliably guarantee lossless (if applicable) runtime/space reductions. To address this, we extend Relational Sparse Matrix Multiplication (RSPMM), originally designed to losslessly lower the space complexity of composition-based MP with pointwise composition functions, to support more expressive functions (e.g., 2x2 block-diagonal matrix multiplication, Givens rotation, circular correlation). Our method delivers significant task-independent reductions in runtime and space for current GNNs on KGs and facilitates efficient re-implementations of GNNs that maintain near state-of-the-art performance on challenging KG tasks, for a fraction of computational costs.
cs.LG / 108 / 2609.34503
Distribution-Conditioned Task Routing for Class-Incremental Learning
Longhuan Xu, Zhipeng Zhou, Wei Ji, Chunyan Miao, Peilin Zhao, Lijun Zhang
cs.LG
Abstract
Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all classes seen so far without access to its task identity. For learners equipped with task-specific parameter-efficient modules, this introduces a critical task-routing challenge beyond catastrophic forgetting. We study post-hoc task routing without retraining the learner or introducing a separately trained router. Such training-free inference-time calibration remains comparatively underexplored in parameter-efficient class-incremental learning. We identify three sources of routing error (feature-level, task-level, and class-level misalignment) and propose Feature Distribution Calibration (FDC). Its three components address these misalignments: Task Subspace Filtering (TSF) suppresses feature components outside each task's principal subspace, Residual Likelihood Calibration (RLC) evaluates the typicality of its subspace residual, and Prototype Affinity Calibration (PAC) measures compatibility with the task's class prototypes. Experiments demonstrate plug-and-play applicability to eight parameter-efficient class-incremental methods using a shared encoder. With one component configuration selected per method across all five benchmarks, FDC improves final accuracy in all 40 method-dataset pairs by 4.39 percentage points on average. Enabling all components improves 35 of the 40 pairs, with an average gain of 4.45 points. When applied to a simple baseline, FDC achieves strong overall performance.
cs.LG / 109 / 2609.34511
HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling
Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang
cs.LG
Abstract
Most existing time series generators rely on a two-stage modeling paradigm: the first stage learns discrete latent representations of time series; the second stage performs autoregressive modeling on these discrete latents through next token prediction. However, this paradigm suffers from two stage-specific limitations: the first stage can lead to information loss when discretizing continuous time series, while the second stage is prone to error accumulation during autoregressive generation. To address these limitations, our core idea is to perform generative modeling in a continuous latent space with a more efficient autoregressive framework. We propose HALO, which enhances time series generation via Hyperspherical Latents and Masked Autoregressive modeling to achieve this goal by tackling two key bottlenecks: (1) variance and scale heterogeneity of continuous latent representations; (2) the difficulty of balancing generation efficiency with temporal correlation modeling. HALO first introduces a hyperspherical VAE that constrains continuous latents to a fixed-radius hyperspherical shell, effectively stabilizing the numerical fluctuations of continuous latent representations. Secondly, we develop a masked autoregressive model that balances parallel decoding and temporal correlation learning, substantially reducing the number of inference steps required for generation and improving generation stability. Our extensive experiments demonstrate that HALO achieves state-of-the-art generation performance while offering significantly improved inference efficiency over existing advanced baselines.
cs.LG / 110 / 2609.34538
Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers
Guanghao Li, Zihan Su, Hao Yu, Jinyang Jiang, Tao Ren, Zehao Li, Feng Lu, Ming Tang, Chun Yuan
cs.LG · cs.AI
Abstract
Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96$\times$ mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.
cs.LG / 111 / 2609.34549
On Parameter Symmetries and Conservation Laws in Gradient Flow
Khang Nguyen, Guido Montúfar
cs.LG · math.OC
Abstract
Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient flow from parameter symmetries, but the scope and limitations of this connection remain unclear. We develop a unified geometric framework that clarifies the precise relationship between the two notions, including the conditions under which symmetries correspond to conservation laws. We introduce a notion of compositional identifiability and use it to establish a general inheritance principle for complete characterizations of symmetries and conservation laws in multilayer networks. We apply the framework to multi-head and grouped-query attention, polynomial neural networks, and square deep linear networks.
cs.LG / 112 / 2609.34553
Verifying Neural Networks with Reinforcement Learning
Hai Duong, Thanh Le, ThanhVu Nguyen
cs.LG · cs.SE
Abstract
Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently explore the verification space. However, existing branching heuristics make greedy decisions based on static scoring functions. They do not anticipate long-term efficiency or leverage the growing availability of verification data to improve performance. This work introduces RSB, a reinforcement learning framework that learns to refine baseline branching heuristics. It trains an actor-critic architecture to maximize cumulative future rewards rather than immediate scores. The actor generates attention weights from observations of raw neuron features and learned graph embeddings, which rescale baseline heuristic scores to guide neuron branching. Evaluation on 600 challenging instances demonstrates that RSB consistently outperforms state-of-the-art branching heuristics, solving 11% more instances while reducing branch exploration by 50%.
cs.LG / 113 / 2609.34555
PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding
Qiuyang Zhang, Kai Zhou, Kai Lu, Haocheng Lu, Jian Zhou, Yuanpeng Su, Kun Bao, Jiguang Wan, Fei Wu
cs.LG
Abstract
Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that existing offloading systems shift the bottleneck to CPU-GPU recall I/O: recall volume varies widely across layers, decode steps and requests, while headwise sparse selection fragments recalls into many small PCIe transfers. This paper presents PulseInfer, an I/O-centric sparse KV cache offloading system. PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine. Implemented on SGLang, PulseInfer improves decode throughput by up to 4.7x over SGLang and 2.6x over the best existing offloading baseline, while reducing TPOT by up to 76% and preserving near-lossless accuracy.
cs.LG / 114 / 2609.34561
Brain-Conditioned Action Policies for Neural Motor Decoding
Luyao Jin, Running Zhao, Huan Zhao, Vincent C. K. Cheung, Wei-Hsin Liao
cs.LG
Abstract
Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarce paired neural-action data. We propose BrainVLA, a framework that enables neural motor decoding by drawing on a pretrained vision-language-action (VLA) model through language-mediated alignment. BrainVLA mitigates reliance on scarce paired neural-action data by leveraging VLA policies. We first construct VLA-compatible datasets including paired neural activity, action signals, language instructions, and rendered visual observations. Then, we adapt the OpenVLA-OFT policy to the target action spaces through LoRA fine-tuning. To establish an effective interface through which neural activity can convey motor intention to adapted VLA policies and guide action generation, we train a neural encoder via neural-language alignment, using language representations as semantic targets to capture latent motor intent from neural activity. The resulting neural representations serve as an endogenous intention signal to guide VLA policies to generate executable actions, while visual observations provide complementary information about the evolving task state. BrainVLA is evaluated on two neural motor datasets with different action dimensionalities using causal rollout decoding. It outperforms the evaluated baselines in cross-session decoding $R^2$ and task success rate, while demonstrating high training data efficiency. These results establish a route for neural motor decoding to draw on large-scale robotic priors through brain-conditioned VLA policies.
cs.LG / 115 / 2609.34562
Single-Layer MeMo as a Randomized Hamming-Kernel Classifier
Alessandro Straziota
cs.LG
Abstract
MeMo (Zanzotto et al., 2025) is a recent language-model architecture that stores associations between token contexts and next tokens in a correlation matrix memory. In this work, we study its single-layer form and show that its ideal retrieval rule is a multiclass classifier based on the positional Hamming kernel. The MeMo architecture represents both the sequence features and the output labels with Gaussian random codes. Its score is therefore a doubly randomized sketch of the ideal classifier. Under independent input and output codebooks, we bound the errors introduced by context sketching and output decoding, characterize their dependence on model and data parameters, and give a margin-based guarantee for recovering the ideal prediction. Controlled simulations support the trends predicted by the analysis. On a restricted WikiText-2 next-token task, we compare single-layer MeMo with classical baselines and show that it can offer a useful trade-off among predictive accuracy, memory, and throughput, particularly on a GPU, where its matrix operations can be parallelized.
cs.LG / 116 / 2609.34566
Retracing Hodgkin and Huxley: State Recovery Does Not Certify Mechanism
Peiyu Zang, Jiayi Hao, Yongqiang Cai
cs.LG
Abstract
Predicting observed dynamics does not establish recovery of the underlying physical mechanism. Can machine learning retrace the hidden-state reasoning behind the Hodgkin-Huxley (HH) model? We train structured latent models on simulated current and voltage, withholding gate identities and trajectories from training and model selection. We then test response prediction, state recovery, protocol transfer, and agreement with HH dynamics. Prediction error and its cross-seed spread both drop sharply at three latent dimensions under the tested protocols, while gate recovery under new protocols improves through five to six coordinates. State recovery depends on which observations the chart uses. Observed voltage improves current-clamp decoding relative to freely predicted voltage. Under voltage clamp, adding latent state to command voltage raises m-state $R^2$ from 0.976 to above 0.99, yet the transported field disagrees with HH on identical smooth samples. Known invertible HH coordinates achieve high fast-m field agreement under the same audit procedure. An exact HH identity decomposes the discrepancy into time-scale-weighted state error and a residual in the transported field; these terms can cancel or reinforce. These findings concern the tested models and charts. They support evaluating state and dynamics recovery separately, including chart inputs and transported-field agreement across interventions.
cs.LG / 117 / 2609.34593
Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation
Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong
cs.LG · stat.ML
Abstract
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep $Q^*$ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
cs.LG / 118 / 2609.34599
The Low-Rank Structure of VLA Reinforcement Learning
Minjae Oh, Yoonah Park, Jongwon Lim, Yohan Jo
cs.LG
Abstract
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $π_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces substantially lower-rank parameter updates that are highly concentrated in the action expert's Timestep Modules, a small and previously overlooked component. Through systematic module-replacement experiments, we further show that these modules capture a disproportionate share of the performance gains from RL. We then characterize what is encoded in these Timestep Modules. First, we show that RL specializes them to the discrete denoising timesteps used during rollouts, and that this discrete-timestep training underlies the low-rank updates. Second, we find that among their outputs, the shift vector changes most distinctly under RL, and through probing, we show that shift update directions strongly predict task success (ROC-AUC up to $99.6\%$). Third, we find that the geometry of shift updates reflects task relationships, as their pairwise similarity correlates with cross-task transfer patterns. Building on these findings, we show that steering along shift update directions further improves RL-trained policies without additional RL training. Overall, we provide a systematic understanding of how RL reshapes VLA policies by studying how learned signals are encoded in parameter space, offering insights into more efficient and interpretable VLA post-training.
cs.LG / 119 / 2609.34602
When local gains fail to transfer: Frozen Earth-observation embeddings across wildfires
Philipp Stark, Alexandros Sopasakis, Ola Hall
cs.LG
Abstract
Frozen Earth-observation embeddings are judged almost entirely by spatially blocked cross-validation inside one study region. We show that this number does not predict accuracy in a new region; we show why; and we show the one setting in which such a model does keep working, using a protocol that needs only a linear probe and labels one already has. The testbed is wildfire, with Copernicus burned-area maps of six fires in Greece and Spain and descriptors from the year before each fire, comparing TESSERA and AlphaEarth with ESA WorldCover classes and annual Sentinel-2 index summaries. Inside a fire, the embeddings identify the burned land 0.05 to 0.13 ROC AUC better than the index summaries, and repeated fold allocations, spatial buffers, a block bootstrap, and gradient-boosted trees leave that margin unchanged. On a fire in another region, they lose 0.15 to 0.18 AUC, and the index summaries lose 0.06, so the three end within a few hundredths of each other. The representation is not the cause. Eight labelled blocks from the new region restore the embedding advantage and give a higher AUC than 59,000 labelled pixels from other regions, and the weight vector fitted in one region is nearly orthogonal to the vector fitted in the others, so the part that carries across regions is small and low-dimensional. Forecasting within a region is a different matter. Fitted on a fire that burned in 2023 and applied to a fire twelve kilometres away that burned in 2024, where nothing used postdates the target fire, TESSERA reaches 0.772 AUC and loses 0.04 against a classifier fitted inside the 2024 fire, while classifiers fitted in other regions lose 0.09 to 0.18. A region with one mapped fire can therefore forecast susceptibility for later fires there; a region without one cannot borrow a model from elsewhere, and every evaluation of a frozen embedding should report a held-out region.
cs.LG / 120 / 2609.34604
Shaping Persistent Representations from Independent Interactions
Ji Dai, Quan Fang, Junyu Gao, Rongfeng Guo, Haoyan Rong, YipingHuang, Yongxi Li
cs.LG
Abstract
World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone, without organizing persistent information into reusable context. We introduce SPRII, a training principle that uses relations between interactions as weak supervision for persistent context while retaining the learner's native objective. For example, different trajectories of the same system share persistent properties even when their states and actions differ. SPRII uses such relations to guide context learning without numerical property labels. Two composable components encourage contexts from related interactions to agree (Align) and use one interaction's context to predict another's future (Cross). Our analysis distinguishes three linked questions: what persistent information is accessible in the learned context (Formation), how that context influences a fixed predictor (Use), and whether it reduces task error (Value). Success at one stage does not guarantee success at the next. Controlled experiments show that more reliable relations improve representation organization, but adding a shared-property constraint can reduce access to a property that remains shared. Context substitutions change predictions at fixed model weights, while the benefit from history depends on prediction horizon and readout. Evaluations span thirteen settings, including controlled physical systems, public dynamics tasks, robotic and tactile data, and partner interaction, across multiple learner families. Relative to the corresponding baselines, SPRII yields average gains of over 10% in downstream task performance and over 15% in persistent-property readout. The project page is available at https://persistent-learning-review.netlify.app/interactive.html.
cs.LG / 121 / 2609.34605
PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
Youzhi Liu, Ruobing Zheng, Boyuan Tong, Tianqi Li, Pingqi Li, Hanbo Bi, Yi Yuan, Jingdong Chen
cs.LG
Abstract
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
cs.LG / 122 / 2609.34611
Beyond Site Agreement: Re-estimation for Brain Network Generalization
Yingxu Wang, Kunyu Zhang, Yanwu Yang3, Thomas Wolfers, Yujie Wu, Siyang Gao, Nan Yin
cs.LG
Abstract
Cross-site out-of-distribution (OOD) generalization in resting-state functional magnetic resonance imaging (rs-fMRI) often relies on learning task-discriminative representations from full-scan functional connectivity (FC) graphs and promoting invariance across source sites. However, FC graphs are estimated from finite, temporally correlated blood-oxygen-level-dependent (BOLD) sequences. Cross-site agreement therefore does not necessarily imply that predictive evidence remains supported under FC re-estimation within the same scan. In this paper, we propose Brain Network Re-estimation-Informed OOD Learning (BRIO), a framework that uses within-scan FC re-estimation to guide cross-site alignment. BRIO maps fullscan graphs and their re-estimates into consistently indexed connectome factors, enabling comparisons of their predictive contributions. It assesses re-estimation support from changes in these contributions relative to within-class subject variability and class separation. For each source-site pair and class, this task-calibrated support from both sites is combined with predictive relevance to form pairwise qualifications, which determine relative factor weights and overall alignment strength. Leave-one-site-out experiments on four real-world datasets (ABIDE, REST-metaMDD, SRPBS, and ABCD) show that BRIO consistently outperforms competitive baselines, with relative improvements of up to 3.8% in accuracy. These gains also persist under an alternative brain parcellation on ABIDE.
cs.LG / 123 / 2609.34614
Learning Regional Snow Water Equivalent and Snow Height Variations from Sentinel-1 InSAR Acquisitions
Luca Barco, Lorenzo Innocenti, Bianca Bartoli, Claudio Rossi, Edoardo Arnaudo, Paolo Garza
cs.LG
Abstract
Managing water resources in mountainous regions depends heavily on reliable Snow Water Equivalent (SWE) and Snow Height (HS) data, yet these variables remain difficult to track at scale. This study evaluates three machine learning architectures (XGBoost, U-Net and SegFormer) for the joint estimation of SWE and HS variations from Sentinel-1 InSAR data over the Italian Alps, using the IT-SNOW reanalysis as reference. SegFormer achieves the best results on both targets, with an MAE of 10.391 cm for HS and 27.113 mm w.e. for SWE and the lowest variability across initializations. A feature sensitivity analysis shows that including all available features does not guarantee the lowest error, with model- and task-specific sensitivities. Spatial metrics (R2, Pearson's r) separate the three architectures far more clearly than mean error (MAE, RMSE) does, and decomposing the error per window attributes most of it to a systematic offset in the estimated mean variation rather than to the spatial pattern.
cs.LG / 124 / 2609.34627
Evaluating Dynamical Fidelity through Predictive Structure in Physical Representations
Oskar Bohn Lassen, Joao Paulo de Souza Boger, Simon Driscoll, Stephen I. Thomson, Sebastian Schemm, Filipe Rodrigues, Francisco C. Pereira
cs.LG
Abstract
Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy selected physical requirements, but provide limited insight into whether learned trajectories reproduce the underlying dynamics. Domain experts examine such relationships through physical representations that expose relevant processes, interactions, and responses, but these analyses are often separated from typical machine-learning evaluation. We introduce a practical framework for evaluating dynamical fidelity through predictive structure in physical representation spaces. Experts define the representations, while reference trajectories determine which relationships are predictive and retained as evaluation tests. We demonstrate the approach in atmospheric forecasting using ERA5 representations of planetary-wave activity and Northern Annular Mode evolution, and evaluate Pangu-Weather, GraphCast, and FengWu. The models exhibit distinct departures from reference predictive structure that are not reflected by conventional forecast errors. The framework thereby turns domain-expert representations into systematic tests of learned physical dynamics without prescribing the relationships in advance.
cs.LG / 125 / 2609.34629
DisKO: Deep Koopman Learning in Distribution Space from Unpaired Snapshots
He Ma, Xiaochen Liu, Wanfeng Lu, Ying Wang, Wei Lin, Qunxi Zhu
cs.LG · math.DS
Abstract
Many complex systems are observed only through temporally unpaired distribution snapshots, making trajectory-based dynamical learning difficult without additional assumptions. We therefore formulate the problem directly in distribution space, treating the distribution itself as the dynamical state. The challenge is that distribution space is infinite-dimensional, making compact and approximately closed representations difficult to learn from finite snapshots. We introduce DisKO, which extends deep Koopman learning to distribution dynamics by jointly learning predictive distributional observables, a finite-dimensional Koopman representation, and a generative map back to the full distribution. Across seven diverse benchmarks, DisKO achieves state-of-the-art extrapolation performance, with substantially slower error accumulation on long-horizon prediction tasks. DisKO further recovers leading Koopman eigenvalues and eigenfunctions on systems with analytic spectra, revealing meaningful dynamical structure in the learned representation.
cs.LG / 126 / 2609.34633
GenMem: Generative Symbolic Memory for Self-Evolving Harness
Xinke Jiang, Tao Feng, Weixuan Xu, Zhixin Zhang, Zhibang Yang, Wentao Zhang, Runchuan Zhu, Xu Chu, Junfeng Zhao, Yasha Wang
cs.LG
Abstract
Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address the sparse, hierarchical, and highly redundant structure of reusable experience: only a small, task-dependent subset of trajectories and memories warrants retention, retrieval, or revision. Learning these operations is further complicated by sparse, delayed, and indirect task-level feedback, with weak supervision across the memory lifecycle. Moreover, continual memory evolution introduces an architectural tension as addressing invariance: stored experience is perpetually revised, yet the addressing interface consumed by learned retrieval policies must remain stable. To address, we present GenMem, which reformulates memory management as generative symbolic addressing. Its core mechanism is the Symbolic Identifier (SID), a multi-level discrete token tuple drawn from a Cartesian-product address space that factorizes a million-scale sparse memory space using fewer than one hundred discrete symbols. Instead of generating ever-changing raw content, the memory agent learns to generate SIDs, while memory evolution rewrites the payload at a fixed address without shifting the address itself. Architecturally, GenMem couples a MemRetriever and a MemEvolver within a multi-agent harness, trained via GRPO with dense process and outcome rewards with two-channels optimization. Under offline memory evolution, experiments spanning ALFWorld, WebShop, multi-hop QA, medical reasoning, and deep research evaluate GenMem against strong memory-augmented baselines...
cs.LG / 127 / 2609.34638
Correction-space Cross-variate Interaction for Test-time Adaptation in Time Series Forecasting
Yuanyuan Deng, Mykola Pechenizkiy, Songgaojun Deng
cs.LG
Abstract
Test-time adaptation (TTA) is a promising paradigm for handling distribution shift in time-series forecasting (TSF), where models adapt at inference time, often leveraging delayed observed data to refine predictions. In the multivariate setting, distribution shifts often exhibit cross-variate dependencies, yet existing TSF-TTA methods adapt each variate independently and ignore this cross-variate structure. Exploiting such structure motivates cross-variate interaction, but coupling variates through backbone predictions introduces direct pathways for mixing uncorrected errors across variates, a concern under the delayed supervision of TSF-TTA. We identify the \emph{interaction space} as a key design choice, and show that acting on adapter corrections that refine backbone outputs, the \emph{correction space}, rather than on the predictions themselves, avoids directly propagating backbone errors across variates. We build on this to propose \textsc{CoRe} (\textsc{Co}rrection-space Interaction \textsc{Re}finement), realizing correction-space interaction through (i) Shared-anchor Correction Refinement (SCR), which combines each variate's correction with a shared anchor through a parameter-efficient bottleneck, and (ii) input-conditioned spectral gating, which adaptively modulates the refinement from the current input window. Across seven backbones, six datasets, and four prediction horizons, \textsc{CoRe} reduces MSE by 25.82\% on average over backbones and 10.57\% over the state-of-the-art TSF-TTA method, with stronger gains at medium-to-long horizons and modest computational overhead. Data and code are available at: https://github.com/yyddou/CoReTTA
cs.LG / 128 / 2609.34639
Uniform Race: Parameter-Free Approximate Rejection Sampling
Seiyun Shin, Juhyeong Pang, Kwang-Sung Jun
cs.LG · stat.ML
Abstract
We study approximate sampling: given $N$ independent samples from a proposal distribution $μ$, the goal is to select one whose distribution is close to a target $π$ specified only up to a normalizing constant. Block and Polyanskiy (2023) provide finite budget error bounds for approximate rejection sampling (RS) as a function of the acceptance threshold $M$. The threshold $M$ giving the smallest bound, however, depends on properties of $(π,μ)$ that are typically unavailable from the observed sample. This raises a natural question: Can one attain the best RS guarantee without taking $M$ as input? We answer affirmatively by proposing a parameter-free sampling algorithm called uniform race (UR), based on importance weights, which are ratios of target to proposal probabilities (or densities). It divides each observed weight by an independent uniform random variable to form a score and returns the candidate with the largest score. For every budget $N$, its total variation error satisfies the RS upper bound for every fixed threshold $M$ simultaneously, thereby achieving the best such bound in hindsight. We also characterize its output distribution conditional on the largest score, identifying when it is exactly the target $π$. Uniform race has no larger total variation error than a natural budget-calibrated RS derived from Rohatgi et al. (2025) and sampling importance resampling (SIR). In particular, we exhibit instances where UR's error is exponentially smaller in $N$ than that of either baseline. Furthermore, we establish conditions under which attaining this RS guarantee for every $(π,μ)$ uniquely determines the selection probabilities as those of UR. Finally, test-time scaling experiments on LLM math-reasoning tasks corroborate the theoretical comparisons and demonstrate that UR remains competitive in ground-truth accuracy without requiring threshold selection.
cs.LG / 129 / 2609.34643
Universal Dynamic Portfolios
Yu-Jie Zhang, Yu-Xiang Wang, Peng Zhao, Kevin Jamieson
cs.LG
Abstract
Cover's Universal Portfolio (Cover, 1991) matches the performance of the best constant rebalanced portfolio in hindsight. We generalize this framework to compete with an arbitrary comparator sequence $\mathbf{u}_1,\ldots,\mathbf{u}_T$, leading to a dynamic regret minimization problem for the log loss where existing methods break down due to potentially unbounded gradients. The log loss is exp-concave, a curvature property that classically yields fast rates for static regret, yet we show that this advantage generally disappears in the dynamic setting. In particular, a linear-loss-type $\sqrt{TP_T}$ dependence is unavoidable, where $P_T=\sum_{t=2}^T\lVert\mathbf{u}_t-\mathbf{u}_{t-1}\rVert_1$ is the standard path length. This limitation stems from the coarse nature of $P_T$, which obscures finer spatial and temporal structure of the comparator sequence. We therefore introduce two structure-aware measures---the Jensen-Shannon distance for spatial structure and the JS$^q$-path length for temporal structure---under which faster rates are attainable when the comparator sequence has favorable structure. To achieve sharp bounds for both measures simultaneously, we develop Universal Dynamic Portfolio, a parameter-free method that combines a new Dirichlet Hedge algorithm with a fixed-share update, while retaining a near-optimal $P_T$ guarantee in the worst case. Finally, under an additional bounded-gradient assumption, we show that OPS admits the faster $T^{1/3}P_T^{2/3}$ dynamic regret rate over all comparator sequences. We attain this rate with a tractable proper algorithm that applies more broadly to general online exp-concave optimization over arbitrary compact convex domains.
cs.LG / 130 / 2609.34650
When Can Attention Heads Be Statically Defined?
Weixian Waylon Li, Yintao Tai, Marcio Fonseca, Shay B. Cohen
cs.LG · cs.CL
Abstract
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance and replaces their attention weights with fitted post-softmax means halfway through training. We represent these fixed patterns with absolute-position and relative-distance preferences, reducing storage from quadratic to linear in sequence length. A fused kernel reconstructs the patterns and executes ordinary-attention and replaced heads together. At matched training-token budgets, replacing 25% of attention heads gives 1.056x faster post-replacement optimiser updates at 124M parameters and 4K context, with a 0.77% perplexity increase. At 1B and 8K context, post-replacement updates are 1.068x faster on four GPUs including communication, with a 0.51% perplexity increase. The resulting models also accelerate long-input finetuning and causal prefill. After associative-recall adaptation, the 124M model with 25% replacement generalises to more key-value pairs at a fixed length better than ordinary attention and two pruning controls.
cs.LG / 131 / 2609.34656
Minimax Last-Iterate Convergence in Matrix Games with Observed Actions
Yuheng Zhang
cs.LG · cs.GT
Abstract
We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneously at every round $t$. This improves the dimension dependence of the best previously known guarantee by a factor of $d^{3/2}$. The rate matches a standard bandit lower bound, establishing minimax optimality in both the number of actions and the number of rounds, up to logarithmic factors. The algorithm is computationally efficient, requiring only $\mathcal{O}(d)$ time and memory per round. Our technical contribution is a joint design of adaptive averaging and corrected exponential weights that absorbs estimation variance, together with a potential argument that bounds phase durations.
cs.LG / 132 / 2609.34673
FestDPO: Few-step Generator Alignment with Direct Preference Optimization
Jaewoo Lee, Kyuil Sim, Hyeongyu Kang, Kanghoon Lee, Woocheol Shin, Jinkyoo Park
cs.LG
Abstract
Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher $β$-sheet fraction and better structural designability than the baselines.
cs.LG / 133 / 2609.34683
AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin, Yao Lai, Haoran Wu, Nicholas D. Lane, Robert D. Mullins, Ilia Shumailov, Yiren Zhao
cs.LG
Abstract
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks primarily focus on simple single-turn chatbot workloads. LLM applications are increasingly agentic: coding agents, terminal execution systems, and tool-use agents issue multi-turn requests with growing context lengths. We introduce AgentPerfBench, a benchmark suite for agentic inference. It uses real traces from agentic benchmarks, such as SWE-Bench and TerminalBench, alongside standard chat baselines. This enables benchmarking of models on multi-turn tasks involving tool calling, skill utilization, and increasing context lengths. AgentPerfBench also samples from empirical distributions of input length, output length, and turn count derived from the real traces, generating representative synthetic profiles for cheap and accurate measurements on new hardware. In addition, we further find that several existing benchmarks fail to accurately reflect real hardware performance for two key reasons: 1) they do not account for realistic context-length growth, and 2) they measure inference performance without operating at hardware saturation. We discuss these issues in detail and provide rich kernel-level Nsight Compute (NCU) traces to construct a new multi-dimensional roofline model that captures hardware-system limitations in both memory bandwidth and memory capacity footprint. The benchmarking suite then includes automated scripts to identify potential bottleneck conditions on emerging hardware when evaluated with diverse agentic traces. Together, these contributions quantify the chat-to-agentic gap in current inference benchmarks and characterise per-kernel GPU resource utilisation via roofline analysis.
cs.LG / 134 / 2609.34692
Predicting Delayed Train Trajectories on the Dutch Railway Network: Explainable AI Evaluation of Topological, Operational and Weather Features with Tree Based Ensemble Methods
Jia Long Bao, Ali Mohammed Mansoor Alsahag, Seyed Sahand Mohammadi Ziabari
cs.LG
Abstract
The reliable prediction of passenger train delays is a critical component of railway management. While contemporary research frequently attempts to maximize absolute accuracy by deploying opaque deep learning architectures, the underlying data mechanics driving longitudinal predictive decay remain underexplored. Consequently, this study provides an explainable temporal robustness analysis of network-wide railway delay prediction. Focusing on the Dutch railway network, this research utilizes interpretable tree-based ensembles to integrate granular topological, environmental, and operational features. The overarching finding establishes that while feature-rich tree-based models improve simultaneous (within-month) prediction, predictive performance systematically degrades when evaluated across non-simultaneous (future) months. Furthermore, multi-horizon SHAP and dispersion analyses explicitly link this degradation to environmental feature volatility and instability within the statistical target definition. Ultimately, this thesis demonstrates that richer feature sets alone are insufficient to resolve long-term forecasting constraints, underscoring the necessity to transition toward dynamic, season-aware architectures anchored by absolute operational boundaries.
cs.LG / 135 / 2609.34711
Learning Propagation Geometry from Message-Passing Feedback
Yingxu Wang, Kunyu Zhang, Xinwang Liu, Mengzhu Wang, Siyang Gao, Chang Tang, Nan Yin
cs.LG
Abstract
Learning local geometry enables graph neural networks (GNNs) to adapt how they compare and integrate neighborhood information. However, estimating geometry from aggregated representations can overlook variation among individual messages and dependencies across feature dimensions. We propose GeoF, a recurrent framework that jointly evolves node features and propagation geometry through message-passing feedback. Each node maintains a local symmetric positive-definite geometry, initialized from a structure-aware prototype atlas and parameterized in block log-triangular coordinates. At each step, the geometry determines neighborhood weights, while triangular frame transport maps transformed source messages into the target node's local coordinates before aggregation. Weighted second-order statistics of residuals between aligned messages and the transformed target state capture directional variation and within-block dependencies, yielding a geometric update target. A shared controller learns complementary corrections through task supervision. A bounded log-triangular update combines these corrections, the target, and the previous geometric state while preserving positive definiteness. The geometry governs subsequent propagation, closing the feedback loop. With parameters shared across recurrent steps, task-specific readouts support node classification, link prediction, and graph classification. Experiments on benchmark datasets show that GeoF consistently outperforms state-of-the-art GNN baselines.
cs.LG / 136 / 2609.34718
Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning
Zijun Weng, Zhongan Bi, Xuanang Gao, Xiaohui Hu, Shuangyong Song, Yongxiang Li, Kaidong Yu, Xuanjing Huang
cs.LG · cs.CL
Abstract
Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks lack a natural success boundary for deciding when efficiency should be prioritized, and (iii) dense, graded rewards often yield small within-group quality margins, making quality-induced advantages especially sensitive to reward-level length shaping, which can perturb their magnitudes and even reverse their signs. We therefore adopt an asymmetric principle: quality should determine the direction of reinforcement, while length should only shape its magnitude. We instantiate this principle with Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged. The bonus strength is further adapted to within-group quality separation, allowing conciseness to matter more when quality-favored responses are similar and less when their quality differences are clear. Across different model families, open-ended benchmarks, and reward sources, QGLAS consistently achieves a stronger quality--length trade-off than representative baselines. At approximately 30% compression, QGLAS retains 98.4--102.0% of the macro-average quality gains achieved by quality-only RL over the base model, compared with 68.3--75.5% for these baselines at comparable compression.
cs.LG / 137 / 2609.34729
Edge-Level Automorphism in GNNs: A Quantitative Framework and Effective Designs For Link Prediction
Chen Shao, Donald Loveland, Tobias Käfer, Danai Koutra
cs.LG · cs.GT
Abstract
Graph Neural Networks (GNNs) are effective for learning node and link embeddings through permutation-equivariant aggregation. However, standard GNNs collapse automorphic nodes, i.e., those with identical structural roles (or orbits) into indistinguishable representations, leading to the node automorphism problem. This collapse limits their expressive power and degrades link prediction performance. Existing approaches to characterize GNN expressiveness rely primarily on Weisfeiler-Lehman (WL) analyses, but these methods are typically qualitative and often misaligned with empirical results. To address this gap, we begin by introducing a novel quantitative framework to assess GNN expressiveness for link prediction. We first formalize edge-level automorphism through edge orbits, which capture the set of structural role pairs for nodes that share a link. Then, we introduce the edge automorphism ratio (EAR), a scalar metric that quantifies a GNN's ability to distinguish links in a given graph. We empirically demonstrate that EAR correlates strongly with performance, validating its practical benefit. Building on this insight, we design EDGE-ORBIT EQUIVARIANT GRAPH NEURAL NETWORK (EO-GNN), a GNN architecture that addresses automorphism collapse while preserving equivariance and incurring minimal computational overhead. EO-GNN accomplishes this through two core designs combined with WL-based node hashes: (i) automorphism-aware dropouts and (ii) subgraph orbit-biased aggregation. Empirical evaluations on synthetic and real graphs show improvements of up to 42.36% and 28.44%, respectively, in predicting links in scenarios with high automorphism.
cs.LG / 138 / 2609.34740
Predictive Dual Smoothing for Column Generation
Senne Berden, Noah Schutte, Andrea Lodi, Tias Guns
cs.LG · cs.AI
Abstract
Solving large-scale linear programs efficiently is an important challenge in many optimization settings. A key technique is column generation, which alternates between solving the master problem over a restricted subset of the variables, and using a pricing subproblem to identify new variables to add. The pricing subproblem is guided by the dual solution of the current restricted master problem, but oscillations in these dual solutions can substantially slow convergence. Dual stabilization methods address this issue. Dual smoothing is a common stabilization method, which guides the pricing subproblem using a combination of the current dual solution and duals from previous iterations. However, while past dual solutions can stabilize the dual trajectory, they do not necessarily guide pricing towards useful new variables. We therefore introduce predictive dual smoothing, which instead combines the current dual solution with a learned prediction of future duals to steer pricing towards variables that are more useful in subsequent iterations. The predictor is trained offline using supervision extracted from standard column generation trajectories and is used only to modify the pricing subproblem's objective function, while exact reduced-cost checks and fallback pricing with the unsmoothed duals preserve correctness. Experiments on cutting stock and generalized assignment problems show that predictive dual smoothing substantially reduces generated columns and wall-clock time relative to standard column generation and existing classical and learned stabilization methods. These gains extend to out-of-distribution instance sizes, and predictive smoothing provides further improvements when combined with strong classical stabilization.
cs.LG / 139 / 2609.34745
No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation
Seonghyeon Kim, Chaeyun Jang, Noah Lee, Boseop Kim, Juho Lee
cs.LG · cs.AI
Abstract
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.
cs.LG / 140 / 2609.34750
A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
Vojtěch Kůr, Adam Kukučka, Tomáš Brázdil, Vít Musil
cs.LG · cs.AI · cs.CV
Abstract
Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
cs.LG / 141 / 2609.34786
Instance-Adaptive Prompts as Context for Time-Series Foundation Models
Zehao Xiao, Shifeng Xie, Lei Zan, Jianfeng Zhang, Lujia Pan, Ievgen Redko, Malik Tiomoko, Keli Zhang
cs.LG
Abstract
Longer histories can improve time-series foundation models (TSFMs), but require substantially higher inference cost. We therefore ask whether contextual information can be provided more efficiently through a compact set of learned token embeddings. We introduce PaCTS, which generates a small set of instance-adaptive latent prompts in the form of continuous embedding tokens conditioned on the visible context. These prompts serve as compact context surrogates for frozen TSFMs. PaCTS constructs them from instance-specific global statistics and further refines them with segment-level temporal information, capturing both global characteristics and local temporal variations. The prompt module is jointly trained and deployed across heterogeneous time series with the frozen backbone. Extensive experiments demonstrate the effectiveness of prompts as context, consistently improving forecasting across context lengths and model architectures. With a shorter input context, PaCTS can outperform the same frozen backbone using double context while requiring substantially less inference computation. Compared with weight-space adaptation methods, PaCTS achieves stronger improvements and better out-of-distribution generalization.
cs.LG / 142 / 2609.34801
Separating personal from population gains when calibrating EEG foundation models for new users
Xilin Tao, Kani Chen
cs.LG · eess.SP · q-bio.NC
Abstract
Foundation models are increasingly adapted to individual users, but an apparent personalization gain can simply reflect a stronger population model. This distinction matters for brain-computer interfaces, where every new user must be calibrated. We evaluated personal adaptation of three frozen EEG foundation models (CBraMod, REVE and LaBraM) in 235 held-out subjects from three motor-imagery datasets, comparing each subject's adapter with the population model and with adapters fitted to other subjects. Using all first-half session labels, personal adapters improved mean balanced accuracy over the population model by 1.5-5.4 percentage points and outperformed exchanged adapters by 2.3-7.3 points in all nine model-dataset combinations. The size of this benefit depended on population training: with four times the original budget, median gains remained positive (1.0-2.0 points) but were smaller for every model, and no population model reached a confirmed plateau. Acquiring the benefit cheaply was unreliable: few-label calibration was consistently non-negative on only one dataset, and in CBraMod neither unlabeled context nor meta-learned initialization outperformed matched controls. Personalization should therefore be evaluated against both a population reference and exchanged parameters, across population-training budgets.
cs.LG / 143 / 2609.34812
Polylogarithmic Nash Regret in Matrix Games with Bandit Feedback
Yuheng Zhang
cs.LG · cs.GT
Abstract
We study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. We develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent $\mathcal{O}(\log^2 T)$ Nash regret against arbitrary adaptive opponents, including games with nonunique equilibria. This resolves the open problem posed by Maiti et al. (2025), extending their polylogarithmic guarantee under bandit feedback from $2\times2$ games to arbitrary finite dimensions. To handle nonunique equilibria, we construct a reference strategy that leaves room for local adjustments. We order independent payoff differences by estimation accuracy and scale these adjustments by uncertainty, allowing the learner to exploit the opponent's imbalance to offset estimation costs. Our result thus shows that observing opponent actions suffices for polylogarithmic Nash regret in general finite matrix games.
cs.LG / 144 / 2609.34825
Gaussian Neural Networks
Peter Kuhn, Victoria Heusinger-Heß
cs.LG · cs.AI
Abstract
Gaussian neural networks (GaNNs) are proposed as a novel regularization mechanism for neural networks. From a Bayesian perspective standard regularization techniques can be viewed as imposing priors over weight-space. Assuming priors over activation-space remains a largely unexplored possibility. GaNNs assume such priors. They do this by treating activities from earlier layers like signals with Gaussian noise and predicting the properties of the noise distribution using an additional unsupervised loss. While training, the unsupervised loss acts as a penalty on unexpected activities, allowing greater weight updates in less surprising directions. The paper demonstrates the superiority of Gaussian neural networks over standard neural networks on a variety of classification and regression tasks. We also investigate the ability of GaNNs to quantify uncertainty.
cs.LG / 145 / 2609.34838
DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents
Hanyang Wang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan, Chaoxu Pang, Zhongda Su, Wulin Xie, Zhizhao Zeng, Ke Zeng, Tianxiang Zhao
cs.LG · cs.AI · cs.CL
Abstract
On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, wasting already-generated experience. To address this problem, we introduce DivOPD, a simple learner-side batch-selection method that spreads a fixed turn budget across more rollouts and, within each rollout, prioritizes turns with larger cumulative teacher-student disagreement. Turns without usable teacher feedback are excluded. The per-turn loss and optimizer remain fixed; selection only changes which student-visited turns receive training weight. For no-progress rollouts, an optional extension briefly hands control to the teacher before returning it to the student. Across six teacher-student settings on the simulated ALFWorld, ScienceWorld, and WebShop benchmarks, with 1.5B-7B students, DivOPD raises cross-setting mean peak success rate from 77.4 to 84.4 and mean success over the last five evaluations from 71.5 to 78.6. It reaches all reported setting-specific targets with geometric-mean speedups of 1.84x in training tokens and 1.87x in learner GPU time relative to vanilla OPD. Teacher intervention further raises this last-five mean to 82.4 while retaining about 1.7x learner-GPU speedup over vanilla OPD. Code will be released at https://github.com/HanyangWang0418-oss/DivOPD.
cs.LG / 146 / 2609.34842
QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities
Hanyin Cheng, Linfeng Wang, Zhengbo Qu, Yang Shu, Zhongwen Rao, Meng Wang, Yijie Li, Xin Jiang, Bin Yang, Chenjuan Guo
cs.LG
Abstract
Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two types of modalities separately. For endogenous modalities, to capture how they evolve along with the underlying temporal dynamics, we introduce an Endo-Multimodal Predictor and Endo-Multimodal Supervision to explicitly learn their evolution from history to the future. For exogenous modalities, to generalize across domains and across various modality types and numbers under the scarcity of exo-multimodal pretraining data, we propose an Exo-Multimodal Retrieval Enhancer that enables rapid downstream adaptation without updating the TSFM parameters. We further introduce Endo-Modality Proxy Training to train this retrieval module without exogenous multimodal pretraining data. Extensive experiments across unimodal and multimodal benchmarks demonstrate strong forecasting performance in scenarios both with and without exogenous modalities.
cs.LG / 147 / 2609.34847
Context-dependent time-series prediction via HyperReservoirs
Kohei Tsuchiyama, Takatomo Mihana, Ryoichi Horisaki, André Röhm
cs.LG · nlin.CD
Abstract
Time series prediction is a common application of reservoir computing. When the training and testing time series data contains multiple dynamical regimes, because an underlying parameter is changing, or the data in fact consists of multiple distinct systems, simple application of the reservoir computing principle produces high prediction errors. Here, we propose a HyperReservoir as an extended model of reservoir computing especially designed for such cases. The HyperReservoir combines a main reservoir with a smaller context reservoir, where the latter modulates the output weights of the former. This structure resembles the hypernetworks from deep neural network literature. However, in contrast, HyperReservoirs retain the simple training via linear regression of standard reservoir computing. We compare the proposed architecture with a conventional ESN, in which context acts at the input, and a full-matrix Conceptor, in which context modulates the reservoir state space. We evaluate all three models on time-series prediction tasks based on Lorenz and Rössler systems, including for varying bifurcation parameters and time sampling scales. We find that the HyperReservoir achieves the lowest mean test error in all three tasks, and particularly outperforms conceptors on data that is sampled from the same attractor but at different time scales.
cs.LG / 148 / 2609.34849
When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation
Xinke Jiang, Tao Feng, Zhibang Yang, Zhixin Zhang, Weixuan Xu, Haoyu Zhang, Xu Chu
cs.LG
Abstract
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK $K_{DR}(n)$, a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token ($K_{DR}(n)\!<\!0$). The severity of these effects depends on the optimization regime: the gradient-norm ratio $κ\!=\!\|\nabla\mathcal{L}_R\|/\|\nabla\mathcal{L}_D\|$ varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...
cs.LG / 149 / 2609.34851
Learning High-Risk High-Precision Motion Control
Nam Hee Kim, Markus Kirjonen, Perttu Hämäläinen
cs.LG · cs.RO
Abstract
Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of high-risk, high-precision motion control problems where actions carry irreversible outcomes, driving sharp peaks and ridges to plague the state-action reward landscape. Using computational pool as a representative example of such problems, we propose and evaluate State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications: 1) Performing policy optimization only using elite samples, allowing the policy to better latch on to the rare high-reward action samples; 2) Utilizing a mixture-of-experts (MoE) policy, to allow switching between reward landscape modes depending on the state; 3) Adding a distance regularization term and a learning curriculum to encourage exploring diverse strategies before adapting to the most advantageous samples. We showcase our features' performance in learning physically-based billiard shots demonstrating high action precision and discovering multiple shot strategies for a given ball configuration.
cs.LG / 150 / 2609.34857
Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts
Chenxiao Fan, Chongming Gao, Gangyi Zhang, Leyang Shen, Yaxin Gong, Jiamin Wang, Jiakai Wang, Dong Wang, Yang Liu, Fuli Feng, Xiangnan He
cs.LG · cs.AI · cs.CL
Abstract
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.
cs.LG / 151 / 2609.34860
Attention-based Hierarchical Variational Information Bottleneck for Robust Multi-Agent Communication under Variable Bandwidth
Lukas Koch Vindbjerg, Qi Zhang, Yury Brodskiy, Lukas Esterle
cs.LG · cs.AI
Abstract
Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first part of each message is received. To address it, we propose \textbf{AH-VIB}, an attention-based autoregressive variational communication model that combines a variational information bottleneck (VIB) with sequential message generation and a hierarchical robustness loss. We evaluate AH-VIB on a custom cooperative object-inspection and occupancy-mapping task, where agents equipped with a limited field-of-view sensor coordinate to scan inspection objects in an occupancy-grid world, under variable and fixed bandwidth conditions, and compare it against MADDPG, CommNet, a flat VIB baseline, and an autoregressive MLP ablation. AH-VIB achieves competitive mean return while improving performance reliability under the most constrained bandwidth conditions. These results indicate that AH-VIB improves the reliability and graceful degradation of learned communication under bandwidth constraints.
cs.LG / 152 / 2609.34866
From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers
Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei, Milad Hosseini, Adrian Weller
cs.LG · cs.AI
Abstract
Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention output instead, over all three projections at once, and reuse it throughout the pipeline. JAB defines one scalar loss over the joint Q, K, V weights of a block, evaluated against the block's real causally-masked attention output, and uses it twice: to fit the quantized weights (GPTQ warm start, then STE with learnable scales), and to score the block for a multiple-choice knapsack allocation. On attention-only quantization of Mistral-7B this works. At 3 bits JAB recovers 77-90% of the gap between uniform GPTQ and full precision, and its sensitivity estimate tracks an oracle costing 73 forward passes to within a fraction of a point. It stops working once MLP layers enter the allocation. A role-aware offset rule needing no sensitivity estimate at all beats JAB on GPT-2's MLP and on the full Mistral-7B model: with a 3-bit floor it quantizes 96.4% of the weights to 4.5 bits per parameter at 6.933 perplexity, within 4.4% of full precision (6.643) at 3.56x compression, against 7.158 for JAB at the same budget. Which matrix a weight sits in matters more than any sensitivity estimate we computed. Two things came out sideways. Block-local reconstruction is an unreliable proxy for end-to-end perplexity: one run improved a block's own objective 4.6x while perplexity rose 32x, which is why every allocation here is validated end-to-end. And on attention-only quantization, fine-tuning moved weights farther from their pretrained values while pulling attention outputs closer, with net gains. Post-training seems to recover attention behavior, not weights.
cs.LG / 153 / 2609.34904
Reference-Tail Trust:Certified Probability Floors for Learned Updates Inside a Deployed Network
Abdolvahab Khalili Sadaghiani, Jose Nunez-Yanez
cs.LG · cs.AI
Abstract
Graph neural networks (GNNs) need to exploit improved message passing without surrendering control over predictions already trusted in deployment. We introduce Reference-Tail Trust (RTT), a framework that admits learned updates inside a frozen GNN and certifies the prediction actually served. RTT couples graph-based proposal states with a constrained internal optimizer: each displacement is charged for its worst-case terminal cross-entropy increase through the incumbent's remaining message-passing layers. A trajectory-validated tube and an independent checker enforce per-node probability floors, $p^{\mathrm{s}}_{ic} \ge e^{-H_{\mathrm{row}}} p^{\mathrm{r}}_{ic}$, and a call-level budget, $\sum_i w_i D_\infty(p^{\mathrm{r}}_i \| p^{\mathrm{s}}_i) \le H^+$, uniformly over labels. Calls whose adapted outputs pass certification require no separate full incumbent rollout; failed certificates trigger whole-call fallback. We derive the exact probability-floor frontier by water-filling, characterize architecture-constrained efficiency, and establish conditions under which internal propagation exploits evidence unavailable to restricted output correctors. In the reported ogbn-arxiv audit, RTT achieves $6.5\times 10^{-3}$ nats of mean gain per call, with a one-sided 95% regression-rate upper bound of 0.95% and a 95% negative-flip upper bound of 0.51% on the uninspected part of the reserved node population. Its mean gain is 61% of a cross-fitted posterior-based frontier estimate and exceeds the strongest matched one-pass corrector by $+0.9\times 10^{-3}$ nats. Reported experiments span eight proposals, six graph-incumbent families, structural and temporal graph shifts, and molecular prediction, with additional image and tabular evaluations. RTT makes GNN adaptation a budgeted, certifiable inference decision rather than an unconditional model replacement.
cs.LG / 154 / 2609.34915
Muon Sublates the Edge of Stability in LLM Pretraining
Yanzhe Chen, Qifang Zhao, Xiaoxiao Xu, Fanghui Liu
cs.LG
Abstract
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary $2ρ_b/η$, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
cs.LG / 155 / 2609.34921
Drug-Target Interaction Prediction via Hierarchical Sequential Cross-Attention over Chemical and Protein Language Models
Khadidja Henni, Hamza Abdelali, Abdelkrim Aries, Neila Mezghani, Brigitte Vannier, Sara Magdouli, Lina Abou-Abbas
cs.LG · cs.AI
Abstract
Predicting Drug-Target Interactions~(DTIs) is a central task in computational drug discovery, with direct applications in virtual screening, drug repurposing, and therapeutic candidate prioritization. Although recent deep learning methods have improved DTI prediction, many sequence-based models still process drugs and proteins independently and only combine their representations at a late prediction stage. This limits their ability to explicitly model cross-molecular dependencies between chemical substructures and protein sequence regions. In this paper, we propose a sequence-only DTI prediction architecture that combines two pre-trained language models, ChemBERTa for drug SMILES strings and ESM-2 for protein amino acid sequences, with a hierarchical interaction module. The proposed model first extracts contextual representations using pre-trained encoders, then applies 1D convolutional layers to condense local sequence patterns, followed by a sequential bidirectional cross-attention mechanism inspired by the induced-fit view of molecular recognition. Finally, attention-based pooling constructs fixed-size interaction-aware vectors for binary prediction. Experiments on BIOSNAP, Davis, and BindingDB show that the proposed model achieves the best performance on BIOSNAP, matches the best AUROC on Davis, and remains competitive on BindingDB while using only 25.2 million trainable parameters. Ablation results confirm the contribution of both the CNN and cross-attention modules, and cold-start experiments indicate promising generalization to unseen proteins and drugs.
cs.LG / 156 / 2609.34924
Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Sebastian Bobadilla-Suarez, Bob Suh, Ryan Fortin
cs.LG · cs.AI · cs.SE
Abstract
An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, so frozen weights buy an eventual ceiling but no stationarity along the way. The criterion also separates three regimes usually merged: search within a fixed class, test-time training that raises the ceiling itself, and scaffold rewriting between them. Audit the scaffold, not the checkpoint. The same ceiling binds sideways. Best-of-$k$ orchestration realizes the best worker's ceiling exactly: width buys rate, not budget. Re-consulting a fixed pool has a horizon computable in advance, decided by the pool alone, and the one arrangement that would beat it, a weighted vote, needs diversity real workers lack: on 30 same-family workers the failure overlap sits at its maximum, and a majority fails 23/55 (42%) of tasks. We obtain the criterion by reading refinement as gradient boosting on the residual error between draft and target, a patch or git diff, and then measuring where that reading breaks: patches compose instead of standing beside each other to be voted on, and failures overlap. What we measure is saturation. Per-round improvement decays toward zero on SWE-bench, and churn decays geometrically across 401 production sessions, a shape shared with a pre-AI human baseline that establishes the regime without identifying its cause. Both breaks are engineering choices rather than laws about code, so together they specify a harness worth building.
cs.LG / 157 / 2609.34929
Sample What You Say: Aligning Language Models to Sample the Distributions They State
Kasra Arabi, Virginia Smith, Chhavi Yadav
cs.LG · cs.CL
Abstract
Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting and changes to decoding reduce this mismatch only partly, which motivates training with policy optimization. Group relative policy optimization (GRPO) is a natural fit for this problem because it already samples a group of rollouts per prompt, and the group's empirical distribution can be compared with the target. However, scoring the group as a whole gives every rollout the same reward. Group-relative centering then sets all advantages to zero, and the model receives no learning signal. To give each rollout its own signal, we introduce the witness advantage, a per-rollout advantage derived from maximum mean discrepancy (MMD). It trains a model to match a target distribution over a finite set of outcomes. The MMD between the model's distribution and the target has a witness function that measures how over- or under-produced each outcome is. Each rollout's advantage estimates the negative witness at its outcome, so a rollout is rewarded for an outcome the group under-produces and penalized for one it over-produces. The witness advantage is computed in closed form from the group's outcome counts, and we use it as the reward in GRPO. On unseen target distributions, training with the witness advantage substantially reduces the total variation distance to the target while largely preserving the model's general capabilities.
cs.LG / 158 / 2609.34933
Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models
Florian Eichin, Philipp Mondorf, Andrei Mircea, Yupei Du, Barbara Plank, Michael A. Hedderich
cs.LG · cs.CL
Abstract
Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.
cs.LG / 159 / 2609.34939
XMatch: Enhancing Covariate-Aware Time Series Forecasting through Tree-Structured Exogenous Matching
Ziyang Zhang, Hanyin Cheng, Xiangfei Qiu, Yang Shu, Bin Yang, Chenjuan Guo
cs.LG
Abstract
Future exogenous variables provide valuable information for forecasting endogenous time series. Existing covariate-aware methods primarily learn the direct influence of exogenous variables on endogenous variables. However, these effects can be complex and change with the pattern of the exogenous variables, making them difficult to capture. Beyond this perspective, we observe that a given exogenous pattern often co-occurs with only a small set of endogenous response patterns. These associations motivate a strategy that matches future and historical exogenous patterns and uses the corresponding endogenous patterns to enhance forecasting. However, in real-world forecasting scenarios with multiple exogenous variables, each exogenous variable provides a distinct dimension for matching, creating a dilemma for this strategy between precise matching and sufficient historical support. To bridge this gap, we propose XMatch (EXogenous MATCHing), a covariate-aware forecasting model that realizes the aforementioned strategy through a tree-structured matching process that adaptively adjusts the number of exogenous variables used as matching conditions. Specifically, we first introduce the ProtoTree Creator, which organizes historical correspondences between exogenous and endogenous patterns into a ProtoTree, whose deeper levels incorporate additional exogenous variables for matching. For forecasting, we then design the ProtoTree Matcher, which uses future exogenous variables to query the ProtoTree and adaptively determines how many exogenous variables to use for matching based on exogenous pattern similarity and historical support. Finally, the matched endogenous patterns are used as explicit historical evidence to enhance forecasting. Extensive experiments on 12 real-world datasets demonstrate that XMatch outperforms state-of-the-art baselines.
cs.LG / 160 / 2609.34944
Adjoint Guidance Flow: Amortized Critic Guidance for VLA Policies
Jeongsol Kim, Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Kwanyoung Kim, Jong Chul Ye
cs.LG · cs.CV · cs.RO
Abstract
Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic through a one-step surrogate of the sampler and back-propagate a critic ensemble at every flow step. In contrast, here we propose Adjoint Guidance Flow (AGF), which amortizes trajectory-aware critic guidance into a lightweight guidance network while preserving the pretrained VLA policy. Specifically, we formulate critic-guided flow generation as a deterministic optimal control problem, whose optimal guidance is a costate that carries the terminal critic gradient back through the remaining flow, and regress the guidance network onto this costate while keeping both the VLA and critic frozen. This design provides favorable memory and throughput scaling during training, and inference needs one guidance-network forward pass per step, without the critic ensemble, back-propagation, or adjoint computation. Across LIBERO, RoboCasa, and LIBERO-Pro, AGF consistently improves pretrained VLAs, remains competitive with critic-guidance and policy-fine-tuning baselines, and is the most robust method when a single guidance strength is deployed across tasks. Compared with QGF, AGF runs $3.6\times$ faster per guidance step with $7.0\times$ fewer parameters, with comparable and even better performance, showing that critic guidance can be trajectory-aware and lightweight.
cs.LG / 161 / 2609.34962
ALICE: In-context, Zero-shot, Mutual Information Estimation
Giulio Franzese, Simone Rossi, Pietro Michiardi
cs.LG
Abstract
Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are moreover tied to specific data types. These constraints limit their adoption in many applications where per-distribution training is impractical and sample sizes are small. We present ALICE, a foundation model that removes per-distribution training, while achieving competitive estimation accuracy. Trained exclusively on a broad family of synthetic distributions, ALICE acts as an in-context estimator of rectified-flow velocity fields: conditioned on samples of an unseen distribution, it estimates that distribution's velocity field without any explicit training. MI is then obtained through a fixed identity that integrates the squared difference between the joint and conditional fields. We validate ALICE on a standard, challenging benchmark and apply it in three domains, biology, genetics, and neuroscience, whose data the model has never seen. For the first time, we show that a single model closes the gap with neural estimators trained separately for each distribution, while natively supporting different data dimensionality and sample cardinality, enabling zero-shot MI analysis across scientific domains.
cs.LG / 162 / 2609.34970
See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs
Weiqiao Que, Ruizhe Li, Chengyu Wang, Dakan Wang, Emine Yilmaz, Xiaofeng He
cs.LG · cs.CL
Abstract
Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: https://github.com/WeiqiaoQUE/mechanistic-emergent-misalignment.
cs.LG / 163 / 2609.34975
Teach to Learn: Hint Annealing for Self-improving LLM Reasoning
Zile Wang, Zijian Li, Haodong Wang, Jian Liu, Qianli Liu, Lucas Muli, Blaze Chen, Song Guo
cs.LG
Abstract
Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based methods construct auxiliary hints from solution evidence and use them to re-solve failed queries, recovering learning signal. Yet the resulting trajectories are typically treated as ordinary solution trajectories despite being generated under an assisted condition unavailable at evaluation. We discover hinted reward shift: recovered reward contrast can concentrate policy updates on hinted trajectories, limiting improvement without hints. This also creates a trade-off: increasing hinted trajectories can accelerate early learning but intensify reward shift later. To address this problem, we propose HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance. To mitigate hinted reward shift, we introduce online weighting to anneal the contribution of hinted trajectories. However, learning to generate hints can conflict with improving query solving. We therefore use gradient projection to remove the opposing component of hint-generation updates. Together, these designs support self-improvement by enabling the policy to create learning opportunities for itself and turn them into stronger reasoning without hints. We evaluate our method on mathematical reasoning benchmarks and outperform state-of-the-art methods by 1.02 pp on Llama-3.2-1B-Instruct, 2.84 pp on Qwen3-1.7B, and 4.32 pp on Qwen3-8B.
cs.LG / 164 / 2609.34976
Inspector: Conversational and Lightweight Analyzer of Analog Circuit Layouts Using LLM and CNNs
Abril Cano Castro, Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni
cs.LG · cs.CV
Abstract
The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes a novel framework that combines fine-tuned LLMs and CNNs to analyze GDSII files of analog circuits, enabling a conversational interface between the tool and the designers. Experimental results using thousands of analog designs across four realistic tasks demonstrate that the proposed solution outperforms state-of-the-art general-purpose massive VLMs by a significant margin (up to 81%), thus providing a lightweight solution to the problem of GDSII analysis.
cs.LG / 165 / 2609.34985
ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients
Shicheng Fang, Yiwen Zhao, Wenbo Tian, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu
cs.LG · cs.CL
Abstract
Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.
cs.LG / 166 / 2609.35011
Price Stability in the European Union: A Systemic Approach Using Random Matrix Theory
Sami Diaf
cs.LG · econ.EM · stat.AP · stat.ME
Abstract
Price stability remains a pillar in monetary policy practices and carries a special importance within monetary unions. Mainstream economics tried to leverage price stability using price indices and several metrics to shed light on specific dynamics and optimal macroeconomic levels. The wide availability of data led researchers to consider the study of systems using Random Matrix Theory, based on inner correlation patterns. This aims to enhance the multivariate analysis by removing noisy patterns from the signal and improve data quality for further inferences. This work considers the collection of monthly inflation indices in the Eurozone as a \textit{system} of prices to analyze its eigenvalues' statistical and asymptotic properties and uncover inner country-level insights. Results confirm the system cannot assumed to be randomly generated, and the data exhibit noise-dominated patterns, due to small and persistent variations at the country-level. The latter make the inter-country correlations more dynamic and the separation of the signal from the noise quiet difficult. Findings identified two countries as distorting inflation dynamics besides three other distinct, regional-based groups of countries. Variability sources might stem from economic episodes fueling inflation spikes in some countries, as well as methodological aspects used to ensure data quality and representativeness in the European Union. Despite being complex, the system demonstrates a certain stability, in terms of self-organization; while large monthly fluctuations cannot be considered as rare events, but part of the data-generating process.
cs.LG / 167 / 2609.35021
Addressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided Masking
Guangyu Wang, Jiawei Tong
cs.LG · cs.AI
Abstract
Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent futures, severely degrading forecasting performance in real-world sensor networks. Existing embedding-based and graph neural network (GNN)-based approaches can partially detect such ambiguous nodes but rely on historical similarity, struggling to capture \emph{future behavioral divergence}. We propose \textbf{STOT} (\textbf{S}patio\textbf{T}emporal \textbf{O}ptimal \textbf{T}ransport), a self-supervised framework that resolves spatiotemporal ambiguity via structured masking guided by optimal transport. Our key idea treats indistinguishability as a \emph{disambiguation} problem: future states are inferred by exploiting concurrent spatial correlations and their time-varying similarity. We design a similarity-aware metric for dynamic inter-node relationships and an optimal transport-based masking strategy to emphasize ambiguous positions during pre-training. A batch consistency constraint preserves semantic coherence, while a random-walk masking mechanism promotes structured context exploration. Experiments on six real-world datasets show that STOT performs competitively with state-of-the-art baselines on the evaluated benchmarks and improved interpretability through transport-plan visualizations.
cs.LG / 168 / 2609.35029
Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
Mana Sakai, Masaaki Imaizumi
cs.LG · stat.ML
Abstract
Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as $n,T\to\infty$ whenever $T=o(\sqrt{n})$. (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.
cs.LG / 169 / 2609.35035
THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout
Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni
cs.LG
Abstract
The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes THEIA, a novel dataset containing thousands of layout images paired with question-answer conversations, along with a benchmark that employs a fine-tuned vision-language model (VLM) to analyze GDSII files of analog circuits, enabling designers to interact with and query physical layouts as intuitive, meaningful entities. Experimental results using thousands of analog designs across five realistic tasks demonstrate that the proposed fine-tuned VLM outperforms state-of-the-art general-purpose VLMs by a significant margin (up to 73%), highlighting a fundamental gap between general-purpose multimodal reasoning and domain-specific layout understanding.
cs.LG / 170 / 2609.35044
BA-DPO: Bias-Adjusted Direct Preference Optimization for Language Model Alignment
Antonio Ferrara, Alberto Rumi, Francesco Bonchi
cs.LG · cs.AI
Abstract
Preference-based alignment methods such as Direct Preference Optimization (DPO) use pairwise preferences labeled by human annotators to fine-tune language models. However, annotators carry systematic biases toward some attributes: a name that signals a gender or an ethnicity, a persona, a language variety, a formatting convention, or length. If not properly addressed, these systematic biases can be absorbed and amplified during alignment. Existing methods address length bias or annotator disagreement, but fail to eliminate biases toward arbitrary attributes. To address this limitation, we propose Bias-Adjusted DPO (BA-DPO), a generalization of DPO that adds one bias parameter per annotator toward responses carrying a declared attribute. We prove that the objective is convex in the bias parameters and that the votes identify each annotator's bias up to a shared constant. The remaining constant is what fixes the aligned model's attribute rate: by default the rate of the reference model, or a target rate, which we use to bring a biased policy to statistical parity. On a corpus with planted biases, DPO drives the attribute from a balanced start to probability 0.96 and BA-DPO removes 81 to 95\% of that shift; on MultiPref with real annotators it removes about half of DPO's lengthening. Both hold at 0.5B with full fine-tuning and at 8B with LoRA, at no higher KL than DPO and no loss in judged quality.
cs.LG / 171 / 2609.35055
Universality and Generalization of Causal Transformers Across Context Lengths
Takashi Furuya, Maarten V. de Hoop, Gabriel Peyré
cs.LG · math.OC · stat.ML
Abstract
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $α$-Hölder sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $β$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{β/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the Hölder-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.
cs.LG / 172 / 2609.35058
TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL
Yibin Huang, Xinming Xu, Conghui Zhu
cs.LG · cs.AI
Abstract
Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.
cs.LG / 173 / 2609.35066
Depot-Closed Multi-Component Construction for Neural Vehicle Routing
Shinichiro Hamada, Hisashi Kashima
cs.LG
Abstract
Most neural constructive solvers for the vehicle routing problem (VRP) use route-by-route construction, extending one route until completion before starting the next. This commits route membership early and hinders global coordination across routes. We propose multi-component construction, which maintains many route components simultaneously and merges them in an arbitrary order. This removes the depot-return cue that route-by-route construction obtains from the remaining capacity; to compensate, we introduce an interpretation in which every component is treated as an implicitly depot-closed route. Under this depot-closed interpretation, every intermediate state of standard CVRP construction is a complete feasible solution, and the exact cost reduction of a merge is the Clarke-Wright saving. The neural policy combines this CW-saving signal with the evolving component state to learn what to connect and when to connect. A policy trained only on CVRP100 outperforms the reported results of representative neural solvers on CVRP100-500 with greedy inference and, reused for ruin-and-reconstruct, performs strongly at all evaluated sizes up to CVRP1000. In a zero-shot Constraint Tightness evaluation with capacities from $C=10$ to $500$, it outperforms the reported neural solvers at every capacity. Controlled analyses show that robustness persists without CW grounding and point to learned route-closing behavior as a plausible contributor to the tight-regime degradation of learned route-by-route solvers.
cs.LG / 174 / 2609.35072
Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning
Xuesong Jia, Ziao Yang, Zhanhe Huang, Hongfu Liu
cs.LG
Abstract
We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.
cs.LG / 175 / 2609.35073
Interrelating Fruchterman-Reingold Graph Visualization and Agglomerative Clustering
Alexandre Benatti, Luciano da F. Costa
cs.LG
Abstract
Graph visualization methods and agglomerative clustering have been frequently considered in data analysis and pattern recognition. Because these approaches are interrelated and complementary, it is of particular interest to investigate their associations. In this work, we study the possible relationship between the Fruchterman-Reingold graph visualization method and four types of agglomerative clustering adopting single- and complete-linkage, average, and Ward's linkage criteria. Three types of datasets have been considered in 2 and 10 dimensions, as well as the PCA projection of the latter to two dimensions. The results obtained suggest that the relationship between the methods considered did not vary much for the three types of data mentioned above. At the same time, the agglomerative methods tended to yield results that are mostly similar to each other, while presenting moderate similarity with the original data. The Fruchterman-Reingold visualization resulted similar to the original data, but exhibited relatively smaller similarity to the agglomerative methods.
cs.LG / 176 / 2609.35075
ReCo: When to Relocate Sensor Kits under Deployment Constraints -- A NILM Case Study
Haokun Chen, Yu Tong, Yehai Chen
cs.LG
Abstract
Many sensing tasks obtain training labels only by deploying instruments in the field. With a limited number of sensor kits, a collection deadline, and measurement downtime at every move, the collector must repeatedly decide whether to stay at the current site or relocate. We study this decision in non-intrusive load monitoring (NILM), which estimates the power drawn by individual appliances from a home's main meter and is trained on data from homes temporarily fitted with appliance-level sub-meters. In NILM, appliance usage varies with the appliance, season and climate, and the value of new data depends on how diverse the combinations of target operation and background load are. To address this, we propose a constraint-based relocation framework and instantiate it for NILM as ReCo (Relocation by Coverage gain). ReCo counts new operating regimes in a joint target-background feature space, forecasts each home's future gain from the data collected so far, and each night weighs the gain of staying against the gain of moving elsewhere after the downtime. In replayed deployments on the Plegma dataset under two kit counts and two downtime costs, ReCo outperforms fixed-dwell and count-based schedules and a threshold rule using the same metric in every setting. Its advantage is not explained by collecting more days alone and reflects allocating the days to more valuable homes and periods.
cs.LG / 177 / 2609.35080
Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers
Preben Johnsen Bentdal, Nello Blaser, Xue-Cheng Tai
cs.LG · stat.ML
Abstract
We study post-hoc refinement of frozen node classifiers: given only the graph $G$ and class distributions $Q$ predicted by a frozen model, can we improve accuracy without access to node features, model parameters, or gradients? APPNP answers this by propagating logits with a restart towards the initial predictions, minimizing the anchored Dirichlet energy. Instead, we consider the Potts energy, and decompose it into a Dirichlet term, which penalizes disagreement between neighbouring nodes, and a Gini term, which penalizes indecision within each node. This decomposition motivates Propagate, Then Sharpen (PtS), which alternates between propagation of class probabilities and node-wise, mass-preserving sharpening, with only one additional hyperparameter selected using labelled validation nodes. Across nine homophilic graphs, with a frozen MLP backbone, PtS improves mean test accuracy over independently tuned APPNP by $1.71$ percentage points on clean inputs and $3.90$ under severe Gaussian feature corruption. Gains over APPNP become smaller, but remain positive with frozen GCN and GraphSAGE backbones. Sharpening also removes most of the accuracy loss of deep propagation: on clean inputs without restart, accuracy falls by $2.2$ points between $2$ and $100$ propagation steps under PtS, compared with $33.8$ for APPNP.
cs.LG / 178 / 2609.35082
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
Yangyang Ren, Haodong Zhu, Linlin Yang, Sheng Xu, Peichao Lai, Baochang Zhang
cs.LG
Abstract
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. We introduce Cross-Rollout Bellman Closure (CRBC), which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest evaluated baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct.
cs.LG / 179 / 2609.35097
SpikeLite: Lightweight Spiking Neural Networks for Time-Series Forecasting
Bang Hu, Changze Lv, Mingjie Li, Xiaoqing Zheng, Wei cao, Fan Zhang
cs.LG · cs.AI
Abstract
Spiking neural networks (SNNs) offer an energy-efficient paradigm for time-series forecasting through spike-driven computation. However, recent SNN forecasters often pursue higher accuracy through increasingly complex attention mechanisms, or specialized neuronal dynamics, weakening the lightweight motivation of SNNs. We introduce SpikeLite, a spiking forecasting framework built around two modules: a Frequency-Selective Spiking Encoder (FSSE) for frequency-sensitive temporal encoding and a Sparse Spiking Channel Attention (SSCA) module for selective cross-channel interaction. FSSE exploits the low-pass filtering behavior of LIF dynamics to reorganize each input sequence into frequency-sensitive components while collectively preserving the input at the decomposition stage. SSCA then learns a binary mask from encoded channel representations and uses it to selectively exchange information within spike-driven self-attention, retaining informative cross-channel interactions while suppressing redundant ones. When explicit channel interaction is unnecessary, SpikeLite uses the lighter FSSE-only channel-independent path. Experiments under the SeqSNN and SpikF protocols cover four standard multivariate and eight long-term forecasting benchmarks. SpikeLite achieves the best aggregate performance under both protocols, with an average $R^2$ of 0.790 and RSE of 0.440, and lowest average MSE/MAE of 0.343/0.345 in long-term forecasting. Moreover, evaluation on the ECL dataset shows that SpikeLite achieves the lowest reported energy consumption, further demonstrating its potential for energy-efficient time-series forecasting.
cs.LG / 180 / 2609.35099
E3J: An Efficient and Open-Source Backend for Euclidean Equivariant Operations on GPU and TPU
Olivier Peltre, Armand Picard, Adrien Pichard, Miguel Bragança, Luca Giacomoni, Valentin Heyraud, Zachary Weller-Davies, Christoph Brunken, Jules Tilly
cs.LG · cs.DC · cs.MS · physics.comp-ph
Abstract
We present e3j, a fast Euclid-equivariance backend for geometric deep learning applications with JAX bindings for GPU and TPU. Leveraging both optimized CUDA and Pallas kernels and algorithmic improvements, the library achieves state-of-the-art throughput and runtime on both forward and backward paths. On a machine learning interatomic potential (MLIP) use case, it outperforms established backends, measuring up to 34% speed-up over cuEquivariance on water box NPT simulation using MACE, while remaining fully open source. E3j achieves over 80% efficiency over the H100 maximum memory bandwidth on tensor product operations, and in many cases more than doubles throughput of message passing convolutions forward compared to previously available backends. In addition, with the release of dedicated Pallas TPU kernel, e3j opens the possibility of large scale equivariant deep learning workloads on TPU architectures, which has so far been difficult to achieve. Our benchmarks show that e3j also achieves over 80% of a TPUv6e memory bandwidth, up to one order of magnitude more than e3nn-jax. The library is available on GitHub, PyPI and is released under an open source Apache 2.0 license.
cs.LG / 181 / 2609.35106
DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
Mustapha Bounoua, Giulio Franzese, Pietro Michiardi
cs.LG
Abstract
Predicting cellular responses to perturbations is a central problem in cellular biology, with broad applications in systems biology and drug discovery. This task is challenging because cellular responses can be complex and cell-state dependent, intrinsic cell-to-cell variability can be confounded with perturbation effects, and destructive single-cell RNA sequencing precludes paired measurements of the same cell before and after treatment. Flow matching transports control cells to perturbed states flexibly, but acting on the full cell state can confound perturbation effects with pre-existing cell-to-cell variability. Disentangled approaches separate responsive from invariant components, but model perturbations through prescribed mechanisms, such as latent shifts or graph edits, limiting their flexibility. We address both limitations in a unified framework. A variational encoder disentangles each cell into an invariant block, capturing state unaffected by the perturbation, and a responsive block, capturing state it changes, through conditional priors and an information-theoretic invariance constraint. Conditional flow matching transports only the responsive block, conditioned on the perturbation and invariant state, yielding a flexible, data-driven model of perturbation effects without confounding pre-existing variability. Across several benchmarks, our method outperforms the strongest published method in settings involving combinatorial and unseen perturbation prediction.
cs.LG / 182 / 2609.35113
SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression
Ziwen Zhang, Xiju Wu, Yuheng Jing, Runxiang Wang, Boxiao Wang, Yifan Zang, Yifan Zhang, Yang Wang, Kai Li, Yifan Zhang, Huilin Xu, Jian Cheng
cs.LG
Abstract
Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack systematic evidence of preserved task diversity and algorithm discriminability. SymbolicArena provides a unified infrastructure for benchmark distillation and dynamic evaluation. The framework standardizes 664 heterogeneous tasks with executable ground truth expressions and distills the Full Task Set into Core50, a validated benchmark of 50 tasks. The distillation process preserves task coverage and algorithm discrimination under explicit balance constraints. SymbolicArena applies a unified execution protocol to heterogeneous SR algorithms and produces comparable outputs and search trajectories. Multi Axis Evaluation characterizes numerical quality, symbolic quality, and search behavior. Core50 reduces evaluation workload by 92.5% and maintains agreement with Full Task Set evaluations. Experiments show that SymbolicArena achieves 72.6% to 86.7% lower approximation error than alternative selectors, further supporting its fidelity to the Full Task Set. Evaluation reveals a substantial gap between numerical fitting and symbolic recovery across current SR methods, suggesting that reliable equation recovery remains an open challenge.
cs.LG / 183 / 2609.35121
A Multimodal Autonomic Sensing Framework for Objective Assessment of Patient Responses to Dental Pulp Stimulation
Youngsun Kong, Yubin Choi, Dongjin Song, Dong-Guk Shin, I-Ping Chen, Ki Chon
cs.LG · eess.SP
Abstract
Patient responses to dental pulp testing, ranging from no sensation to intense pain, provide important information for assessing pulp status in endodontic diagnosis. However, pain is a subjective sensory and emotional experience that varies considerably across individuals and can be difficult to communicate. We investigated whether complementary autonomic signals could support objective assessment of responses during dental examination. Forty-nine patients underwent cold pulp testing, yielding no-response, mild-response, and intense-response conditions. The framework integrated ECG-derived skin nerve activity (SKNA) and R-R intervals (RRI), together with electrodermal activity (EDA), using temporal convolutional network encoders with attention-based mid-level fusion. Individual baseline signals and subject-level covariates, including anxiety scores and biological sex, were also incorporated. The framework achieved 80.2% balanced accuracy, 75.2% sensitivity, and 85.2% specificity for binary classification of no response versus mild or intense response. For three-class classification, it achieved 60.0% balanced accuracy and a 58.8% macro-averaged F1 score. Ablation and attention-weight analyses indicated that EDA contributed most strongly to model performance, followed by RRI, while SKNA improved balanced accuracy by approximately five percentage points. Age was significantly associated with model performance. These findings support the feasibility of multimodal autonomic sensing for objective, non-invasive assessment of responses to dental pulp stimulation.
cs.LG / 184 / 2609.35128
Explaining Hyperbolic Neural Networks via Geometry-Aware Relevance Propagation
Ping Xiong, Shanglin Li, Yi Ding, Thomas Schnake, Shinichi Nakajima
cs.LG
Abstract
Hyperbolic neural networks introduce geometric operations that require explicit treatment in relevance propagation. Equivalent geometric realizations can produce different feature attributions, even when local relevance is conserved. We study this problem through Geometric Representation Invariance (GRI), a specialization of Implementation Invariance, and zero-curvature consistency, which requires identity relevance propagation when a geometric module approaches the identity. We propose LRP-radial-all for origin-centered radial modules, treating geometric scaling as modulation and assigning relevance entirely to the signal branch. The rule conserves relevance, is invariant to equivalent radial factorizations, and satisfies zero-curvature consistency, yielding GRI for a specified Poincaré-Lorentz logarithmic-map construction. In contrast, a conservative LRP-half baseline can violate both consistency criteria. Experiments on hyperbolic MNIST, sEEG, and CIFAR-10 classifiers assess attribution fidelity, qualitative explanations, and runtime. LRP-radial-all achieves competitive attribution fidelity across datasets with runtime comparable to Gradient$\times$Input and substantially lower than Integrated Gradients. These findings motivate geometry-aware propagation rules that distinguish relevance conservation from consistency across equivalent computations.
cs.LG / 185 / 2609.35130
CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning
Junkang Liu
cs.LG · cs.AI
Abstract
Communication-efficient federated optimization commonly spends several gradient evaluations between server updates. Existing local-update methods use this computation to advance an independent model on each client. Under heterogeneous data, however, these models evaluate gradients at different locations, making the aggregated update difficult to interpret as a gradient of the global objective. We study an alternative use of the same computation budget: \emph{evaluate the global objective along a shared, predicted path}. We propose Common-Trajectory Predictive Federated Learning (\texttt{CTP-FL}). At each round, all clients construct the same sequence of query points from the current global model and the previous aggregated direction, evaluate $K$ stochastic gradients along this sequence, and upload their average. The server then performs a single global update. Thus, \texttt{CTP-FL} uses $K$ mini-batch gradients per client and one model-sized vector in each communication direction, matching the per-round computation and communication of full-participation FedAvg-M. Shared query points make the aggregated direction an unbiased estimator of the average \emph{global} gradient along the predicted path. The remaining discrepancy from the gradient at the current model is controlled by the path length, without assuming bounded client-gradient dissimilarity or bounded gradients. For smooth non-convex objectives, we establish an $\mathcal{O}\!\left( \sqrt{LΔσ^2/(NKR)}+LΔ/R \right)$ average-stationarity bound under full participation. The analysis isolates a testable trade-off: extending the prediction path provides more forward-looking gradient information but increases its displacement bias.
cs.LG / 186 / 2609.35138
FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu
cs.LG
Abstract
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.
cs.LG / 187 / 2609.35139
CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion
Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu, Guoliang Xing, Zhenyu Yan
cs.LG
Abstract
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.
cs.LG / 188 / 2609.35161
Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism
Amir Mohammadpour, Michael Franke
cs.LG
Abstract
We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the model from the causal structure of the true data-generating process, (ii) orients more towards natural language prediction by considering inference of latent common causes, and (iii) considers whether and how Bayesian evidence accumulation for latent common causes can be implemented in representations and mechanisms that allow for cross-context generalization to novel test cases.
cs.LG / 189 / 2609.35167
EdgeCraft: Automated Model Crafting for Edge IoT
Genglin Wang, Kaiwei Liu, Liekang Zeng, Wangsong Yin, Shangcheng Jin, Guoliang Xing, Zhenyu Yan
cs.LG · cs.DC
Abstract
Machine learning (ML) increasingly powers Internet of Things (IoT) applications at the edge. Yet producing a deployable edge ML artifact for a specific scenario requires navigating a huge search space spanning data representation, model design, training on domain-specific data, and runtime customization. This workflow is fragmented and difficult to scale across diverse edge applications. We present EdgeCraft, an LLM-driven system that turns high-level intent into deployable edge ML artifacts. Building such a system raises two challenges: (1) How can an LLM be guided to find high-quality solutions that meet dynamic SLOs for task quality, latency, and energy? (2) How can trustworthy target-device verification be obtained at low cost? EdgeCraft addresses these challenges with two designs. (1) A constraint-aware synthesis tree explores alternative candidates and uses measured SLO gaps to guide each improvement. (2) A multi-fidelity verifier progressively combines low-cost checks with full target-device verification to reduce verification cost while preserving reliable verification results. It also records verified failures for reuse, avoiding repeated device work. To support concurrency, EdgeCraft provides a multi-tenant runtime that runs cloud training and target-device verification in parallel while isolating requests. Across 50 public tasks, EdgeCraft exceeds the task-specific Reference in best-observed quality on 40 tasks and finds an SLO-feasible artifact on 45, with the two outcomes overlapping on 38 tasks. Moreover, EdgeCraft achieves competitive performance on our self-collected SEN dataset, suggesting its generalizability to real-world IoT sensing tasks.
cs.LG / 190 / 2609.35168
QAM: Quadratic-Accurate Checkpoint Merging via Sequential Consistency
Shihao Wang, Rui Kong, Xinran Chen, Hui Wu, Qipeng Qian, Jinman Zhao, Jiashu Zhao, Yuchen Li, Jimmy Huang, Dawei Yin
cs.LG · cs.AI
Abstract
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size $h$ can achieve $o(h^3)$ endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by $Ω(h^3)$. \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform $O(h^3)$ endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.
cs.LG / 191 / 2609.35174
ProtoSeam: Lifting Classifier Training with Latent Gaussian Mixture Models
Robert Lampel, Timon Klein, Sebastian Sager
cs.LG
Abstract
We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network $N=N_2\circ N_1$ is split at a single semantic interface and one learnable prototype per class is inserted there. Training combines a quadratic consensus penalty that pulls $N_1(x)$ toward the prototype of its class with a classification loss of $N_2$ evaluated on samples drawn around the prototypes, whereat no gradient crosses the interface. At inference the prototypes are discarded and the unmodified network $N_2\circ N_1$ is used. Across CIFAR-10, CIFAR-100, and TinyImageNet with ResNet and vision transformer backbones, lifted training improves test accuracy by up to five percentage points over variants without lifting under a shared tuning protocol. Moreover, we provide theoretical justification of those results.
cs.LG / 192 / 2609.35177
Subgroup Rank-1 Lattice for Practical High-dimensional Black-box Integral Approximation
Yueming Lyu
cs.LG
Abstract
Estimating integrals of black-box, high-dimensional functions, from expectations and kernel mean embeddings to the softmax kernel in self-attention, is a basic subroutine in machine learning. Rank-1 lattice rules suit this setting: they query the integrand only at a fixed point set and need no gradients. When the $n$ points serve as a design matrix $X\in\mathbb{R}^{n\times d}$ for a feature map, however, computing $Ψ(X)^\top v$ or $Ψ(X)w$ for an elementwise nonlinearity $Ψ$ costs $O(nd)$ time and memory for any standard quasi-Monte Carlo point set. We study subgroup rank-1 lattices, whose Korobov generator $(1,t,\dots,t^{d-1})$ uses a scalar $t$ of fixed multiplicative order $m$. Splitting $\mathbb{F}_n^\times$ into cosets of $\langle t\rangle$ reduces both maps to short cyclic correlations evaluated by FFT, giving exact results for arbitrary $Ψ$ in $O(n\log m)$ time and $O(n)$ memory, without forming $X$. Since fixing $m$ falls outside classical component-by-component theory, we prove convergence directly: via resultants with the cyclotomic polynomial $Φ_m$, the squared worst-case error in the Korobov space decays as $O(n^{-(α-1)/(m-1)})$ for prime $m\ge d+1$, and this threshold is exact. Using the splitting of $n$ in $\mathbb{Q}(ζ_m)$, averaging over the $m-1$ admissible generators improves the constant by a factor $Θ(m-1)$. Empirically, the subgroup lattice beats Gaussian and orthogonal random features and scrambled Sobol' and Halton points in 49 of 54 synthetic kernel-estimation settings and all 45 softmax-attention settings on nine real datasets, and builds a sample set with $d=2048$, $n\approx4.1\times10^7$ in 2.3 ms.
cs.LG / 193 / 2609.35193
ConRAG: Lightweight inference of multi-hop relations
Kilian Bänziger, Sonia Laguna, Markus Kreft, Robert Jakob, Kevin O'Sullivan, Lasse B. Strand, Julia E. Vogt
cs.LG
Abstract
Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and other knowledge-intensive analyses. We formalise this setting as multi-hop relation inference: given two known endpoint entities, we aim to recover the bridge entities and evidence-grounded reasoning chains that connect them across a document corpus, and to generate an explanation grounded in the retrieved evidence. Existing multi-hop RAG systems typically seek an unknown answer entity rather than explicitly recovering the connection between two known endpoints and graph-based approaches often rely on costly LLM-extracted knowledge graphs that limit scalability to large document collections. We introduce ConRAG, which builds a lightweight entity-document graph from entity co-occurrence and LLM-based entity filtering. Its connective retrieval infers and semantically ranks paths between two endpoints. On MuSiQue and 2WikiMultiHopQA, ConRAG consistently improves bridge entity and reasoning chain recovery over strong RAG baselines, while reducing graph-indexing token cost by up to roughly 1.5 orders of magnitude. Our results show that endpoint-constrained path retrieval provides an effective and index-efficient approach to evidence-grounded relation discovery.
cs.LG / 194 / 2609.35212
Adversarial Consistency-Guided Representation Learning for Multi-view Clustering
Yuchen Lin, Kunpeng Xu, Ying Fang, Lifei Chen
cs.LG
Abstract
Multi-view clustering aims to capture cross-view consistency while exploiting view-specific information. However, shared representations learned to capture cross-view consistency may still retain view-identifying information, potentially compromising the consistency of cross-view clustering structures. To address this issue, we propose ACGRL, an adversarial consistency-guided representation learning framework for multi-view clustering. ACGRL employs a gradient-reversal view discriminator to reduce view identifiability and obtain invariant reference representations. These representations are then frozen to provide fixed references for disentangling view-specific information from cross-view common information in the subsequent learning stage. The fixed reference representations are concatenated with the learned view-specific representations for reconstruction and clustering, with cross-view cluster alignment encouraging consistent clustering assignments. Experiments on four benchmark datasets demonstrate the superior clustering performance of ACGRL compared with representative multi-view clustering methods.
cs.LG / 195 / 2609.35219
Temporal Heterogeneous Graph Pretraining for Relational Deep Learning
Yixin Peng, Er Jin, Diego Collarana, Stefan Decker
cs.LG
Abstract
Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work often treats time as a single signal or studies temporal representation and pretraining separately. We investigate how explicitly encoding both signals affects temporal pretraining for downstream tasks. Our framework combines Multi-scale Time Encoding, which captures record age using learnable time scales and type-specific projections, with Rotary Time Encoding, which represents signed inter-record intervals through rotary transformations during graph propagation. We pair these encodings with three self-supervised objectives: historical relation recovery, horizon-aware future relation activity prediction, and temporal subgraph contrast. All inputs respect their observation cutoffs. Pretraining proceeds in two stages: subgraph contrast first learns neighborhood representations, followed by refinement through either relation recovery or future activity prediction. We evaluate on five RelBench datasets across 11 classification and regression tasks using heterogeneous GNN and graph Transformer backbones. With both encodings, the best evaluated staged schedules improve over supervised training with the same encodings by 3.02% and 1.06% on the two backbones, respectively, and over controls without pretraining or either encoding by 3.24% and 2.37%.
cs.LG / 196 / 2609.35236
Long-Horizon Scaling: How Model Capabilities Shape the Returns to Computation
Haoyu Zheng, Zhengyu Chen, Huaisheng Zhu, Ruishan Fang, Teng Xiao, Yiwei Li, Jingang Wang, Wenqiao Zhang
cs.LG
Abstract
Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two long-horizon benchmarks. We find that starting performance and subsequent growth are associated with different capabilities: within a task category, similar early scores can precede different later gains. To formalize this finding, we model capability-time scaling with category-specific logistic power laws shared across models. Fitted to early trajectories, these curves extrapolate the observed models' category-average scores to later computation. However, rising average scores mask narrowing improvement opportunities: later gains concentrate among fewer improving models. High final scores and continued improvement also have distinct capability profiles. Predicted mean gains estimate each model's fraction of improving tasks; averaging these estimates forecasts the average share of improving models. These uneven returns motivate deciding whether a specific run should continue. We therefore derive a continuation policy to save time and compute with limited score loss. The policy conditions growth predictions on the run's observed progress and weighs immediate and delayed gains against computation costs. In replay with training and price calibration based on other models' histories, the policy saves roughly one-third of full-run time, with relative score losses of 2.4% on AutoLab individual runs and 3.3% on EdgeBench published mean curves. Our repository is available at https://github.com/Chihaya-Anon-chan/long-horizon-scaling.
cs.LG / 197 / 2609.35240
Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
Surajit Das
cs.LG
Abstract
Artificial-intelligence studies using computed tomography (CT) for lung cancer are often broadly labelled "prediction" despite addressing clinically distinct tasks. We systematically mapped CT/low-dose CT (LDCT)-centered lung-cancer AI using five-database retrieval, full-text eligibility assessment, role-aware modality/omics extraction, clinical-task classification, and a Multi-Tier Evidence Graph (MTEG). The final corpus comprised 293 studies (2016-2026): 230 Detection, 8 future Risk-prediction, and 55 Other studies. Clinical variables (96.2%), 3D CT/LDCT (73.0%), and radiomics (63.5%) predominated, whereas external validation (29.0%), calibration (20.5%), decision-curve analysis (13.0%), longitudinal CT (17.7%), and saliency/attribution XAI (21.5%) were less frequent. The MTEG comprised 377 nodes and 3,444 edges; only 31 studies (10.6%) completed the six-tier substantive evidence chain, with greatest attrition at reasoning/explanation. Overall, the literature is detection-dominated, genuine future risk prediction remains uncommon, and complete translational evidence chains are rare.
cs.LG / 198 / 2609.35257
AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning
Yue Xie, Zhi Zheng, Yunpeng Ba, Xuyang Wu, Xialiang Tong, Zhichao Lu, Tao Zhong, Zhenkun Wang
cs.LG
Abstract
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly informative directions. To make these evaluations more informative, existing ZO methods restrict perturbations to low-dimensional subspaces. Yet the quality of these subspaces is critical: overly compressed or poorly maintained spaces can miss useful update directions. To obtain a high-quality subspace for ZO updates, this paper proposes AIM-ZO, a ZO fine-tuning method based on Activation-Informed Subspace Maintenance. AIM-ZO uses forward activations as local directional information and continuously integrates them into a broad, evolving subspace over training. To access broader gradient-relevant structure while keeping individual perturbations low-dimensional, AIM-ZO activates only a smaller set of shared and sampled directions, decoupling the maintained width from the active width. We evaluate AIM-ZO across 5 LLMs and 11 downstream tasks under matched forward-evaluation budgets; its six-task average exceeds the strongest fully evaluated ZO baseline by 1.26 percentage points on OPT-2.7B and MeZO by 2.85 percentage points on OPT-30B. Our code is available at https://github.com/EkkoXy/AIM-ZO
cs.LG / 199 / 2609.35259
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar
cs.LG
Abstract
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
cs.LG / 200 / 2609.35260
Latency and accuracy tradeoffs in Spiking Neural Networks
Zhanglu Yan, Zixuan Zhu, Kaiwen Tang, Yuyang Cai, Qianhui Liu, Weng-Fai Wong
cs.LG
Abstract
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.
cs.LG / 201 / 2609.35268
SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards
Yingchao Yu, Pengfei Sun, Wenxuan Pan, Wei Chen, Yitian Hong, Kuangrong Hao, Yaochu Jin
cs.LG
Abstract
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
cs.LG / 202 / 2609.35288
$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
Berker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello
cs.LG · cs.CV
Abstract
Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $λ$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $λ$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $λ$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose $λ$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $λ$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.
cs.LG / 203 / 2609.35291
Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
Shunchang Liu, Lukas Fluri, Xin Chen, Francesco Croce
cs.LG · cs.CV
Abstract
Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in text-only tasks, leaving its manifestation in multimodal models unclear. In this paper, we define and analyze EM in the context of vision-language models. We first induce EM via fine-tuning on narrow multimodal tasks targeting vulnerable code, careless household-object use, and conspiratorial interpretations of ordinary scenes. Across fifteen commercial and open-source models with different scales, we find that narrow multimodal fine-tuning can induce coherent and broadly misaligned behavior that transfers to unrelated tasks, including misaligned opinions, visual factual dishonesty, unsafe image generation, vulnerability to visual jailbreaks, and risky agentic actions. We further find that multimodal EM does not depend on the apparent harmfulness of training data but is sensitive to training-evaluation modality alignment. EM can arise under both supervised fine-tuning and preference optimization and can propagate through intermediate reasoning. Finally, we explore several mitigation strategies, including prompt inoculation, benign continued training, and activation-level steering, which can partially reduce EM. Overall, our findings suggest that multimodal EM reflects a behavioral shift rather than a general loss of capability, extending beyond text to the visual modality.
cs.LG / 204 / 2609.35297
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horváth, Martin Takáč, Aleksandr Beznosikov
cs.LG
Abstract
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per $P$ steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with $P=2$ and $P=5$ reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon
cs.LG / 205 / 2609.35314
When Should a Satellite Estimate Be Changed? Stress-Testing Neural Corrections for Evapotranspiration
Marco Trotta
cs.LG
Abstract
Neural residuals can improve satellite evapotranspiration (ET) estimates, but selectors must predict when a correction helps and reject unsupported inputs. We evaluate ten-member models on 16,366 flux-tower observations from 151 stations paired with OpenET, across nine rolling years and five spatial folds. At one held-out station, Gain accepted corrections on all 32 physically invalid records: it predicted a mean benefit of 0.83 mm/day, but the corrections increased mean absolute error by 21.6 mm/day versus OpenET. On spatially held-out unit errors, SupportGain reduced station-macro MAE versus Gain by 0.148 mm/day under wind x3.6 (simultaneous 95% interval, 0.070 to 0.226), with 9.3% acceptance versus Gain's 51.8%; on clean inputs, its 0.006 mm/day advantage had an interval that includes zero. These fault analyses are exploratory; none of 40 preplanned temporal comparisons passed Holm correction, while a separate predeclared cropland contrast found 0.041 mm/day lower station-macro MAE with crop-only training (95% interval, 0.009 to 0.079).
cs.LG / 206 / 2609.35319
Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao, Xuchen Li, Honglin Lin, Tao Cheng, Zhihan Yu, Kai Tang, Xiaoxi Jiang, Guanjun Jiang
cs.LG · cs.AI
Abstract
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.
cs.LG / 207 / 2609.35322
Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
Jérémie Klinger, Raphaël Urfin, Giulio Biroli, Marylou Gabrié
cs.LG · cond-mat.dis-nn
Abstract
Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weighting schedule $w(t)$. Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the \textit{speciation time}. In this work, focusing on high-dimensional data, we decompose the integrated loss into its single-time contributions and analyze each at fixed signal-to-noise ratio $Λ(t)$: we show that $Λ(t)$ sets the rate at which each feature of a multimodal target - the mode directions and their relative weights - is acquired during training. Crucially, at high $Λ(t)$ all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are not learned at all. Only near the speciation time, where $Λ(t)$ becomes of order one, do all features become learnable, each on its own timescale: the weights are acquired jointly with the directions, and the directions at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is then governed by how much of the weighting effectively sits near the speciation time, which provides insights on $w(t)$ design choices. These results follow from an exact high-dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures. Numerical experiments on image and human genome haplotype generation recover the predicted hierarchy of learning timescales in more complex settings.
cs.LG / 208 / 2609.35333
Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation
Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao
cs.LG
Abstract
Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.
cs.LG / 209 / 2609.35338
NeuronDiscover: Agent-in-Twin for Mechanistic Discovery in Neuronal Microenvironments with World Action Models
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
cs.LG · q-bio.NC
Abstract
Mechanistic discovery in neuronal microenvironments requires interventions and measurements that separate competing explanations of solute transport and neuronal response. Predictive accuracy cannot settle the question: a real mechanistic change and an error in the computational twin leave the same signature in sparse observations. We formalize this twin confounding and reason over a joint mechanism--discrepancy belief, designing experiments that separate the two. NeuronDiscover is an Agent-in-Twin framework whose shared, mechanism-grounded World Action Model (WAM) couples prediction, intervention proposals, and observation design; independently adjudicated outcomes revise a scoped Mechanism--Intervention--Observation--Outcome (MIOY) graph, whose supported relations compile into executable programs carrying discrepancy-adjusted acceptance bounds. We evaluate on simulated brain-fluid tracer-transport worlds adjudicated by an independently frozen finer-mesh reference solver, and on donor-disjoint public current-clamp recordings of cortical neurons. Counting only relations that reach a certified terminal status, and scoring abstentions as unresolved for every method, at a matched budget of 16 experiments over 32 source units NeuronDiscover resolves 4.0 relations per assigned world against 3.4 for the strongest baseline and 3.2 without graph revision, at 5% false support and 82% scope accuracy. Joint mechanism--discrepancy acquisition resolves 3.8 relations versus 2.9 for plug-in expected information gain; discrepancy-adjusted verification lowers accepted-program failure from 15% to 9% at 60% acceptance coverage; and transfer to the recordings yields 1.94 versus 1.53 relations per assigned world. Correctness is adjudicated within declared model worlds and archival recordings.
cs.LG / 210 / 2609.35347
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen
cs.LG
Abstract
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
cs.LG / 211 / 2609.35348
From Data to Program: Fast & Direct Generative Program Inference from Empirical Data
Simon Klüttermann, Xueying Ding, Leman Akoglu
cs.LG · cs.AI
Abstract
Estimating probability densities from a finite set of samples typically requires dataset-specific model fitting. We introduce PRODiGI, a pretrained data-to-program model that infers an explicit, executable generative program in a single forward pass. Pretrained on synthetic datasets paired with their ground-truth programs, PRODiGI accommodates diverse generative families and data dimensionalities through template prediction and non-autoregressive program parameter decoding. Its inferred programs support direct sampling, density and score evaluation, and inspection independently of the pretrained model. We further introduce program-space fine-tuning, which refines differentiable program parameters by matching generated and empirical samples while keeping model parameters intact. Experiments show that PRODiGI achieves lower average density and score MAE than existing pretrained models, while offering multi-fold speedups over its closest competitors. Program-space fine-tuning further reduces generation MMD by 84%. By turning empirical data into explicit, reusable programs, PRODiGI introduces a new direction for fast, interpretable tabular generative modeling.
cs.LG / 212 / 2609.35349
Quasi Linear Kernel Attention with Infinite Capacity
Nicolaj Rux, Johannes Hertrich, Sebastian Neumayer
cs.LG · math.NA
Abstract
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi linear computation. To quantify expressivity, we introduce a capacity for each kernel, measuring the maximum sequence length for which the attention matrix can approximate the identity. A higher capacity thus indicates greater expressivity. We show that expressive kernels like softmax, Gauss, and Laplace have infinite capacity. In contrast, common quasi linear kernels, such as those derived from finite dimensional feature maps, exhibit finite capacity. As a solution, we propose additive kernels constructed from univariate spline and polynomial exponential kernels. We prove that these maintain infinite capacity while allowing quasi linear computation via sorting. Finally, we implement additive sorting kernels efficiently and benchmark them against modern softmax backends, demonstrating advantages for long sequences.
cs.LG / 213 / 2609.35351
Interference Beyond Geometry in Concept Extraction
Valérie Costa, Bahareh Tolooshams
cs.LG
Abstract
Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms: feature orthogonalization, bias compensation, gain adaptation, and encoder-decoder separation. Experiments with sparse autoencoders show that constrained architectures selectively reduce overlap among co-active features, while bias, gain, and encoder freedom allow constructive cross-contributions to remain. Together, these results show that interference in learned representations depends not only on feature geometry, but also on how features are used and on the architecture that produces their codes.
cs.LG / 214 / 2609.35352
Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights
Yiteng Peng, Zhibo Liu, Dongwei Xiao, Shuai Wang
cs.LG
Abstract
Fully homomorphic encryption (FHE) enables neural network inference directly on encrypted inputs, but it remains orders of magnitude slower than plaintext in- ference. Applying the server's plaintext weights to encrypted activations involves plaintext-ciphertext multiplications (PMult) and accounts for more than half of inference time in recent systems. Ternary quantization can replace these multipli- cations with additions and subtractions, but the savings rarely materialize under packed execution. A single PMult applies a weight group fixed by the packing layout and can be avoided only when all its weights share the same ternary value. Ternarizing all groups, however, largely degrades accuracy. We present FIONA, an offline optimizer that selectively ternarizes weights within a given packing layout based on the estimated effect of ternary conversion on the model's performance. FIONA encourages a shared ternary value within each weight group and retains full-precision weights for sensitive groups, so ternar- ized and full-precision paths coexist within a layer. It then compiles these hybrid operators exactly, applying common scaling factors once to accumulated inputs and reusing sums across outputs. Weight ternarization can also narrow the input ranges of downstream polynomials. FIONA fits lower-degree replacements under a cumulative accuracy budget, reducing multiplicative depth and bootstrapping. On VGG11, ViT, and BERT, FIONA reduces PMult operations by 53.4-79.5% and accelerates end-to-end encrypted inference by 2.38x, 1.68x, and 1.84x, re- spectively, with less than 1% accuracy loss across all three models.
cs.LG / 215 / 2609.35364
Do Temporal Link Predictors Need Learned Memory? A Smoothed-Count Baseline with a Handful of Parameters
Lisi Qarkaxhija, Ingo Scholtes
cs.LG · cs.AI
Abstract
Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal link predictor based on statistical language modelling. It pools transition and co-occurrence counts across sources to predict links that a source has never formed. We smooth sparse estimates using destination frequencies or Kneser-Ney continuation counts. A shared log-linear rule combines these estimates with popularity, source history, and recency, without node embeddings. In our main evaluation, the model achieves the highest MRR among the compared methods on 7 out of 16 datasets from TGB and TGB-Seq. It also outperforms EdgeBank and Base3 on all 16 datasets and the heuristic family on 14. These gains extend to datasets designed to limit repeated edges. With only 9--13 learned parameters, our model provides a simple and competitive baseline for evaluating future neural temporal link predictors.
cs.LG / 216 / 2609.35379
Identifying Neural Source Dynamics from Unknown Local Interventions
Ayana Mussabayeva, Jiaqi Sun, Anuar Aimoldin, Olivier Oullier, Kun Zhang
cs.LG
Abstract
Electroencephalography (EEG) records mixtures of brain-source activity. Even with a known anatomical forward model, experiments that excite only part of the source-state space leave the dynamics unidentified, and repetition cannot resolve the ambiguity. We show that unknown local mechanism changes can supply the missing information. We consider linear dynamics among fixed anatomical sources with known source-state initialization patterns. Changing one source's update rule for one transition leaves a rank-one, source-specific signature in subsequent EEG: subtracting matched baseline responses isolates it, and the forward model identifies the source and calibrates its response history. Combining these histories with initialization responses recovers source interactions without baseline reachability and without first identifying the intervention coefficients. We establish sufficient recovery conditions, a direct estimator, and a noise-sensitivity bound conditional on correct source labels. Simulated EEG on anatomy derived from magnetic resonance imaging confirms the information gain: with baseline excitation confined to four of twelve source coordinates, eight unknown changes recover all dynamics in 32/32 systems, whereas baseline realization, baseline regression through an invertible forward model, and changes that leave the tested states unexposed all fail, and explicitly constructed alternative dynamics reproduce every baseline mean. Where baseline information suffices, direct reconstruction is also more reliable than a matched-information spectral estimator. Nonlocal changes and forward-model error limit accuracy even when source labels are correct.
cs.LG / 217 / 2609.35390
Inductive Feedback for Mixed-Policy Distillation
Amir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang
cs.LG
Abstract
Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.
cs.LG / 218 / 2609.35392
The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
Yihe Zhou, Tongtian Zhu, Yingxiao Huo, Satya Prakash Dash, Can Wang, Samuel Kaski, Mingfei Sun
cs.LG · cs.AI
Abstract
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$β$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$β$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
cs.LG / 219 / 2609.35402
Persistent Partners Raise Prices Among Learning Agents
Paul-Peter Arslan, Yubin Kim, Xiao Xiao
cs.LG · cs.MA
Abstract
When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand duopoly of Calvano et al., each agent's price is set by a tabular Q-learning module, not by the small language model attached to it, and we randomise whether each agent keeps its partner, sees its rival's prices and can send messages. Keeping the same partner raises the level of profits, averaged over training, by 0.27 of the gap between competitive and monopoly profit (95% CI 0.20 to 0.35, all twenty paired runs positive), our registered primary result, and the resting price by 0.17 of the Nash-to-monopoly range (post hoc). A plain tabular learner reproduces the effect in all 25 further blocks, and there one permanent partner raises the level more than about three do (+0.23 against +0.05, exploratory). Where rival prices are hidden, the price-setting module cannot see a cut, so cannot punish it, yet the resting price rises as much and the rise lasts to the end of training, while with visible rivals it shrinks with longer training (post hoc). Where the rival is visible, a static best responder accounts for a third to a half of what a forced-deviation probe reads as punishment, on the starts where the rival can see the cut, and net of it the registered test of learned punishment is inconclusive. A test that looks only for punishment would thus miss the rise where the rival is hidden, while a check for profitable deviations flags most of those prices (post hoc). In an exploratory extension, untrained Qwen2.5 7B and 14B models under one prompt show the effect when the rival's price is left out of the prompt and inconsistently when it is shown, the 7B result replicating on fresh blocks, while two other model families show none.
cs.LG / 220 / 2609.35427
LLMs are General Asynchronous Agents
George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov, Vladimir Kaurkin, Vadim Pastushenko
cs.LG · cs.CL
Abstract
Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.
cs.LG / 221 / 2609.35433
ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh
cs.LG · cs.AI
Abstract
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $α$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.
cs.LG / 222 / 2609.35441
Riccati State Space Models: Non-iterative Parallelization for Nonlinear Sequence Modeling
Mónika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu
cs.LG · cs.AI
Abstract
State space models (SSMs) achieve efficient sequence processing because their affine state updates are closed under composition and can therefore be evaluated with an associative parallel scan. Nonlinear recurrent models can provide richer, state-dependent dynamics, but generally lose this compositional structure: parallel evaluation then requires iterative methods that repeatedly linearize and scan the recurrence. We ask, what state-dependent nonlinear dynamics can be designed to remain exactly composable? We answer by introducing RiccatiSSM, a nonlinear SSM, in which each state dimension follows an input-conditioned Riccati differential equation. Its quadratic state dependence makes the local Jacobian explicitly state-dependent, while its exact per-step flow under piecewise-constant inputs is a Möbius transformation. Since Möbius maps are closed under composition and compose through $2\times 2$ matrix multiplication, the complete nonlinear state trajectory can be evaluated exactly with a single associative parallel scan, without iterative linearization. We further derive a constrained parameterization that ensures bounded, contractive dynamics, and avoids poles in the fractional-linear state update. Across long-sequence classification, regression, and forecasting tasks, RiccatiSSM achieves competitive predictive performance while reducing runtime by $22{-}33\%$ compared to the nonlinear LrcSSM under matched architectures. These results demonstrate that state-dependent nonlinear dynamics can retain exact composability and be evaluated efficiently within a single parallel scan.
cs.LG / 223 / 2609.35454
Manifold-Stable Flow Matching
Amirhossein Nazerian, Ali Pezeshki, Jianguo Zhao
cs.LG · cs.RO · eess.SY
Abstract
Flow matching (FM) learns generative dynamics through velocity regression. Geometric FM variants commonly assume a prior supported on the data manifold, requiring geometric knowledge that is often unavailable. Without such knowledge, low regression error alone does not guarantee manifold adherence. Adherence keeps generated samples within valid configurations and is empirically associated with better task performance. We introduce manifold-stable flow matching (MSFM), which can start from an arbitrary ambient prior, not necessarily supported on the manifold. Using tools from nonlinear dynamics, namely contraction theory, MSFM combines learned tangential transport with prescribed normal contraction. The construction uses analytical projectors for known manifolds and local affine proxies estimated by principal component analysis for unknown data geometry. By implementing contraction theory in both cases of known and unknown manifolds, we guarantee manifold invariance and transverse convergence to the manifold within a desired time window (e.g., one second). We derive a family of compatible probability paths and decompose the training loss into a learnable tangential term and a normal residual. An ellipse experiment attains a mean terminal off-manifold error of order $10^{-6}$. In Push-T robotic experiments, MSFM raises success from $74\%$ to $82\%$. In the Robomimic Square task, success increases from $60\%$ to $72\%$, while rotation-manifold deviation decreases from order $10^{-2}$ to $10^{-7}$. The MSFM terminal geometric errors are controlled by the chosen numerical tolerance. These results demonstrate stronger geometric adherence and higher observed task performance, supporting prescribed normal contraction as a complement to learned generative transport.
cs.LG / 224 / 2609.35462
CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski, Alexandra V. Kulinkina, Kristina Keitel, Akhil Arora, Lars Henning Klein, Mary-Anne Hartley
cs.LG · cs.AI · cs.CL
Abstract
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.
cs.LG / 225 / 2609.35465
Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter
Pier-Jean Malandrino
cs.LG
Abstract
Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on the same lattice. A 24-weight block still takes 48 bits, most of which index a 64-state trellis of the Golay code and one shared 16 KiB table. The kernel decodes a block with six table loads and two small lookups inside the matrix-vector product, and reads 2.148 bits per weight. For full models, we retrain one scale per matrix row, store the matrices that lose the most as 4-bit integers, and pay for them with 4-bit embedding tables. Our Qwen3-4B, 8B and 14B files hold 2.73, 2.70 and 2.73 bits per parameter over the whole model. They score 63.37, 69.58 and 75.66 on the full MMLU test set, 4.76, 4.21 and 2.46 points below 4-bit AWQ at 5.3 to 6.0 bits per parameter. They generate 113.8, 95.0 and 57.2 tokens per second in our engine. On GSM8K, through the served kernel, they lose 9.63, 4.62 and 3.26 points to FP16. At 4B our file scores 23.6 points above llama.cpp's IQ2_XXS (2.48 bits per parameter). Every number we measured for a table or figure comes from one NVIDIA L40S GPU. We preregistered the main experiments.
cs.LG / 226 / 2609.35466
An analysis of Mirror-Descent Soft Actor-Critic
Denis Zorba, Michal Valko
cs.LG · math.OC
Abstract
Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor objective, characterised by the curvature of the $Q$-function estimate through the Legendre differential operator, and establish an $\mathcal{O}\!\left(N^{-\frac{1}{5}}\right)$ best-iterate finite-time convergence rate up to actor and critic approximation errors. Moreover, the mirror-descent step size $λ$ directly controls the target drift and hence actor tracking error, whereas the analogous Gibbs bound contains a non-vanishing tracking term.
cs.LG / 227 / 2609.35483
Universal Approximation of Measure-to-Measure Operators by Pushforwards
Takashi Furuya, Nicholas H. Nelsen, Frank Cole
cs.LG · math.NA · stat.ML
Abstract
Many learning tasks map an input distribution to an output distribution. A natural way to model such an operator is to transform each input sample using a continuous function that may depend on the entire input distribution, and then take the distribution of the transformed samples. This defines a measure-dependent pushforward model and includes measure-theoretic formulations of transformers. We ask when such models can approximate arbitrary continuous operators between spaces of probability measures. We first show that universal approximation fails when atomic inputs are allowed: some continuous measure-to-measure operators that split or redistribute atomic mass cannot be approximated arbitrarily well by deterministic pushforward models. We then introduce the uniform level set condition, which requires a continuous measure-dependent scalarization whose shrinking level set neighborhoods carry uniformly vanishing mass over the input family. This condition is satisfied, in particular, by compact families of absolutely continuous measures. On every compact family satisfying this condition, we prove that any continuous measure-to-measure operator with outputs of finite $p$-th moment can be uniformly approximated, in the $p$-Wasserstein distance, by continuous measure-dependent pushforwards. Combining our theorem with existing approximation results for measure-dependent in-context maps yields universal approximation by measure-theoretic transformers. We also extend the framework to continuously-varying source measures, yielding a corresponding universality result for a class of pushforward models that are closely aligned with cross-attention architectures.
cs.LG / 228 / 2609.35502
Structured Latent Modeling for Supervised Multimodal Information Decomposition
Wanting Huang, Sanvesh Srivastava, Weiran Wang
cs.LG
Abstract
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.
cs.LG / 229 / 2609.35505
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
cs.LG
Abstract
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
cs.LG / 230 / 2609.35512
Improving Generative Model Self-Training with Geometrically Modified Outputs
Patrick Batsell, Thomas Walker, Richard Baraniuk
cs.LG · cs.AI
Abstract
Self-training generative models - the continued improvement of a model using its own outputs - is becoming increasingly important as high-quality training data becomes scarce. However, naively finetuning on model-generated samples leads to degradation through model collapse and the model autophagy disorder. Negative-guidance self-training methods turn this degradation into a useful signal, using a model finetuned on its own outputs to guide the original model toward improved generation. Existing methods, however, take the negative signal in standard model outputs as given. We instead ask whether this signal can be explicitly strengthened. We introduce Geometrically Modified Outputs (GMOs), which reweight the singular values of the generator's input-output Jacobian to increase the influence of its leading singular directions. This geometric modification amplifies the mode-seeking behavior and distortions of standard outputs, providing a stronger and more targeted negative signal for self-training. Across a range of one-step generative models, GMOs consistently improve the performance of negative-guidance methods, including Neon and SIMS, compared with using standard model outputs.
cs.LG / 231 / 2609.35514
One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices
Ruishuo Chen, Weijia Li, Xun Wang, Yu Chen, Leheng Cai, Longbo Huang
cs.LG · stat.CO
Abstract
In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count estimator, but its efficiency depends critically on the proposal distribution. Existing proposals are analytically designed, and their accuracy can vary substantially with the margins. We show that the ideal SIS proposal, under which every weight equals the count and the variance vanishes, is exactly the policy of a generative flow network (GFlowNet) with unit reward on every matrix that has the given margins. We therefore propose MarginFlow, a framework that turns the design of the proposal into a learning problem and amortizes it across margins by exploiting their self-similarity. Every partial matrix is itself an instance with reduced margins, so one set transformer that reads the remaining margins serves every margin. We train MarginFlow on a pool of 1904 margins and evaluate it zero-shot on 1190 held-out margins, synthetic and real, from $3\times3$ to $870\times6$. On 1187 of the 1190 margins it matches or beats the best of 31 analytically designed configurations, chosen post hoc for each margin, and its median effective sample fraction is 99.8%. On the 56 margins where that best loses more than one nat of effective sample size, MarginFlow wins every one and raises the median effective sample fraction from 10.3% to 94.1%.
cs.LG / 232 / 2609.35517
Reward-Aligned Reweighting for On-Policy Distillation
Haofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou, Chuan Wu
cs.LG
Abstract
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R$^{2}$-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R$^{2}$-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.
cs.LG / 233 / 2609.35525
Deep Epistemic Value Functions for Optimistic Exploration
Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause
cs.LG
Abstract
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
cs.LG / 234 / 2609.35527
GeoGAE: Scalable Graph-Level Autoencoding via Hyperball Cloud Representations
Radosław Nowak, Anna Bielawska, Bogusz Stefańczyk, Maciej Sanocki, Paweł Wawrzyński
cs.LG
Abstract
Embedding structured objects into Euclidean spaces has enabled a wide range of successful machine learning applications. Such objects include words, documents, image patches, time series, and graph nodes. In contrast, embedding entire graphs remains a challenging problem. Existing methods either sustain the original order of the graph nodes or match the output nodes to the input ones, both of which create scalability issues. In this work, we propose a graph representation as a cloud of hyperballs, which allows us to define a specific, typically unique, node ordering. Based on this representation, we propose GeoGAE, an autoencoder, in which the Transformer encoder translates a hyperball cloud into a graph-level embedding, and the Transformer decoder translates the graph-level embedding back into the graph. This formulation enables the model to capture both the global graph structure and local relational patterns. We evaluate our method on multiple graph datasets, spanning various domains. The results demonstrate effectiveness of our method in encoding and reconstructing graphs from their embeddings.
cs.LG / 235 / 2609.35528
Let the Neurons Die: Exploiting ReLU-Induced Model Degradation
Kexin Li, Wenjun Qiu, Joshua Abraham, Aditi Maheshwari, David Lie
cs.LG · cs.AI
Abstract
Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example that minimizes the target layer's post-update weight sum, aiming to push ReLU units toward negative pre-activations without modifying training samples or labels. We then develop two poisoning attacks, IG-DOA and IG-SKA, which use gradient inversion to synthesize class-conditioned samples by matching reference gradients in adverse model states constructed through data ordering or soft knockout, respectively. Soft knockout rearranges weights across adjacent layers to concentrate negative contributions. On a fully connected ReLU network trained on MNIST, ordering 100 of 60,000 training examples reduces test accuracy from 96% to 95% after only five epochs. Adding 200 poisoned samples from a single class reduces test accuracy to approximately 86-88% after five epochs in most evaluated conditions, compared with approximately 96% under clean training. These results demonstrate that ReLU-targeted data ordering and poisoning can impair learning without directly modifying the victim model's parameters.
cs.LG / 236 / 2609.35537
Optimal Networks for Agentic Information Aggregation
MohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Shayan Taherijam
cs.LG · cs.GT · econ.TH
Abstract
We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a subset of the features and its parents' predictions, fits a linear predictor to minimize mean squared error, and passes only its prediction forward. The global predictor is the best linear predictor using all features. Kearns, Roth, and Ryu show that the output agent's error approaches the global predictor's error along sufficiently deep paths with suitable feature coverage, while insufficient depth can prevent aggregation even in large networks. In contrast to their main focus on a given graph and feature allocation, we consider the limits of the model under two settings. In the adaptive designer setting, a designer chooses the graph, feature allocation, and output agent knowing the distribution. In the oblivious designer setting, the designer fixes all three before an adversary chooses the distribution. Each agent observes one feature and receives predictions from a limited number of parents. We call the aggregation exact when the output agent matches the global predictor exactly. For $d\ge3$, we show that no finite depth guarantees exact aggregation for every distribution with one parent per agent, even when the designer knows the distribution. In contrast, two parents per agent suffice for exact aggregation even in the oblivious designer setting. A fixed graph, feature allocation, and output agent achieve this for every distribution at depth $O(d\log d)$. Knowing the distribution reduces the depth to $O(d)$. Both constructions use $O(d^2)$ agents, with a very large constant for two parents. We show the bounds on the depth and number of agents are all optimal up to constant factors.
cs.LG / 237 / 2609.35541
Learning the Robustness Mechanism with Bilevel Optimization
Yiyang Shen, Qihang Lin, Weiran Wang
cs.LG
Abstract
We propose a distributionally robust learning framework where parameters defining the robustness mechanism are learned from held-out data instead of extensively tuned. Using bilevel optimization with both upper and lower level minimax problems, we create two instances of our framework to tackle setups with and without group labels in the training set. Theoretically, we provide sample complexity analysis for our robustness mechanism learning paradigm, showing that it achieves generalization guarantees comparable to exhaustive grid search while being more computationally efficient. Empirically, we evaluate our framework under a challenging setup when both intra-group and inter-group test distribution shifts occur at the same time, thereby demonstrating the efficacy and scalability of our method.
cs.LG / 238 / 2609.35545
Graph World Models for Constrained Epidemic Policy Planning
Yiqi Su, Rashed Shelim, Lingyi Wang, Walid Saad, Naren Ramakrishnan
cs.LG · cs.AI
Abstract
Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We present EpiMind, a graph world model framework for constrained epidemic policy planning across regions. A graph-factored recurrent state-space model generates joint policy-conditioned rollouts from regional latent beliefs, while graph-temporal ADMM optimizes regional interventions, enforces shared-resource feasibility through projection, and evaluates temporal specifications under the learned model. EpiMind reduces admission RMSE by 29% relative to graph-free dynamics modeling, plans within 1-5% of the best feasible constant policy with guaranteed shared-budget feasibility, and outperforms all deployable baselines across three resource budgets in real-context evaluation. These results demonstrate that graph-structured policy imagination with explicit constrained coordination supports effective epidemic interventions from learned dynamics.
cs.LG / 239 / 2609.35568
From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis
Longxiao Fan, Tao Zhang, Han Yan, Jiajun Li, Mingcong Song, Guoping Long, Hongjie Si, Weiwei Sun
cs.LG
Abstract
High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarchies differ substantially from those of GPUs. To address this transfer gap, post-training methods adapt LLMs to NPU programming but depend on scarce expert data and substantial training compute. Memory-learning agents instead adapt through external memory, but their uniform credit assignment gives adopted and unused experiences the same reward target, potentially biasing subsequent retrieval rankings. Moreover, when learned values guide only retrieval, high-value experiences that generalize across operators must be retrieved repeatedly rather than retained in context, thereby increasing retrieval overhead and weakening cross-task guidance. We therefore present SAGE, a persistent self-improving agent for NPU kernel synthesis. Adoption-Traced Utility estimation (ATU) combines explicit adoption records with kernel evaluation outcomes for adoption-aware credit assignment. Utility-Gated Consolidation (UGC) uses positive utility and repeated adoption across operators to select and abstract reusable rules into a bounded resident context. On NPUKernelBench, SAGE achieves a 95.5% execution rate versus 84.1% for the strongest controlled baseline, with 86.9% of solved operators outperforming torch_npu. With GLM-5.3, SAGE achieves a 43.99x speedup over the torch_npu reference on sparse flash attention. These results show that adoption-aware credit assignment and selective consolidation enable agents to accumulate and reuse hardware-specific knowledge across tasks.
cs.LG / 240 / 2609.35587
Hardware-Aware Features for CUTLASS Kernel Selection
Shriram Chandran, Dominic Rinderer, Yakup Budanaz, Alexandru Calotoiu, Marcin Copik, Torsten Hoefler
cs.LG · cs.PF
Abstract
GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. Existing analytical selectors require hand-designed performance rules, while learned selectors operate on raw configuration parameters and must infer hardware consequences from data. We introduce a hardware-aware representation for CUTLASS kernel selection that augments candidate configurations with statically computable estimates of induced hardware behavior. We construct a dataset of 4.9 million CUTLASS kernels and train gradient-boosted and neural learning-to-rank models to rank candidates within each problem. On held-out exhaustive evaluation problems, hardware-aware representations reduce selection regret by up to 40\% relative to structural baselines and 64.2\% relative to NVIDIA's matrix-multiply heuristics. We further evaluate data-efficient cross-precision and epilogue-fusion transfer within CUTLASS GEMM, showing that explicitly representing candidate-induced hardware behavior provides a useful inductive bias for learned kernel selection.
cs.LG / 241 / 2609.35603
Control-Geometry Straightening for Sampling-Based Latent Planning
Ziang Fu, Ning Ning
cs.LG · stat.ML
Abstract
Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly representations by directly straightening control geometry for sampling-efficient planning. CGS matches pairwise cosine similarities among actions to those among corresponding latent differences only using local transitions from pixel-action pairs. The loss can be applied across world-model architectures using end-to-end learned or pretrained representations. Under linear-dynamics, our theoretical analysis connects this objective to temporal straightening and more balanced terminal-cost curvature across the full planning horizon, yielding finite-budget guarantees for MPPI, local contraction results for CEM, and convergence bounds for gradient descent. Across four control environments and multiple planners, CGS improves planning with fewer sampled candidates and refinement steps, achieving success-rate gains up to 20 and 12.6 percentage points over LeWorldModel (LeWM) and its temporal-straightening variant (LeWM+TS), respectively, with sampling-based planners using 128 candidates per update. Probes, comparisons with DINO-WM architecture, and planner-side ablations clarify how latent motion organization, state dependence, and dynamical context shape planning behavior. Straightening control geometry thus makes good action sequences easier to find under limited planning budgets.
cs.LG / 242 / 2609.35614
EvE: An Alternate Optimizer to Adam
Shashank Raj, Kalyanmoy Deb
cs.LG · cs.NE
Abstract
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and since gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven scalable benchmarks up to one million variables. On three real neural-network tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE finishes the same charged budget 1.7-3.9x faster, at a modest cost in final quality (about one accuracy point on MNIST, 9-11% higher relative test loss on the two fine-tuning tasks; on GSM8K, Adam is about 5 accuracy points more accurate, and fine-tuning lowers accuracy below the base model for both). Inside successive halving on UCI Adult, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, ranking configurations about as consistently with Adam as Adam does with itself across seeds (Kendall's tau 0.66-0.69). EvE is not a total replacement for Adam as a final-stage trainer, but a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up.
cs.LG / 243 / 2609.35620
Attention Graphons: A Graph Limit Perspective on Graph Transformers
Caio F. Deberaldini Netto, Moshe Eliasof, Luana Ruiz
cs.LG
Abstract
Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-dependent? We answer this using dense graph limit theory, treating each attention matrix as a finite sample from an underlying kernel---an \emph{attention graphon}---and studying concentration around this limit under the cut-distance. We derive a worst-case variance bound requiring no assumptions on the graphon, and a sharper regularity-aware bound based on nonparametric estimation theory. To operationalize the theory, we propose a canonicalize-then-block-average pipeline for estimating dataset-level attention graphons, and a variance-based diagnostic for testing whether attention admits a stable continuum description. Experiments across multiple graph benchmarks show that learned attention stabilizes to dataset-specific graphon structure on several datasets; that empirical cut-distance and cut-norm variance decreases with $n$ consistent with our bounds; and that attention graphons transfer to larger graph sizes with error decreasing in $n$.
cs.LG / 244 / 2609.35628
Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
Zilan Cheng, Li-Lian Wang, Zhongjian Wang
cs.LG · cs.NE · stat.ML
Abstract
We study the minimum number of hidden neurons required for arbitrary-accuracy approximation of multivariate Hölder-continuous functions on $[0,1]^d$ and the associated encoding complexity. For $d\geq 2$, we construct a fixed, explicitly defined activation function for which a closed-form network with two hidden layers of widths $d$ and $1$ achieves arbitrary accuracy in the uniform norm. We prove that $d+1$ is the exact minimum total number of hidden neurons among standard feedforward networks with locally integrable activations and affine outputs. We further give a simpler construction using a single elementary activation that combines the floor and exponential functions. This construction requires three hidden layers of widths $d$, $1$, and $2$, only two neurons above the minimum. If a skip connection is allowed, widths $d$, $1$, and $1$ suffice. These constructions use explicit grid addressing and integer encoding of quantized function values. For a bounded $α$-Hölder class, they require $O(\varepsilon^{-d/α}\log(1/\varepsilon))$ bits, matching the metric-entropy lower bound up to a logarithmic factor.
cs.LG / 245 / 2609.35629
SANTA++: Sampling Attention through Representative Keys
Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari
cs.LG · cs.CL
Abstract
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusion probability. This importance sampling correction estimates attention over the full cache, with a sampling budget that lets us trade memory reads for accuracy. Remarkably, with 32 or 64 sampled teams, SANTA++ uses 16% to 22% of dense attention's KV reads and retains 94% to 99% of the dense-attention baseline's scores on LongBench v2 and HELMET's retrieval-augmented generation subset, and 85% to 91% on RULER, with Qwen2.5-7B-Instruct at 32K context. With 31 sampled teams, our GPU implementation delivers a $1.69\times$ attention speedup over the dense FlashAttention baseline at 32K context. By reducing the number of cache entries read, SANTA++ in principle complements architectures with compressed KV representations, such as multi-head latent attention. Our kernels are available at: https://github.com/OPUSLab/santapp-kernel-demo.git.
cs.LG / 246 / 2609.35635
Bounding Retraining Equivalence and the Deletion Floor in Materials Machine Unlearning
Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban
cs.LG · cond-mat.mtrl-sci · quant-ph
Abstract
In materials machine learning, closely related retained structures can sustain accurate property predictions even after removing a specific record, rendering post-deletion prediction error an ambiguous metric for machine unlearning. To resolve this ambiguity, we define the deletion floor as the expected target loss under a specified retraining procedure at the deleted request. Standard indistinguishability constraints yield a sharp interval bounding an update's target loss around this baseline reference. Theoretically, a conditional neighbor bound links a low deletion floor directly to retained fit, prediction regularity, and local label agreement, while an exact ridge identity isolates residual fit from the prediction change induced by record deletion. Empirically, controlled redundancy sweeps show an $\approx 8\times$ drop in median normalized retraining loss when one retained relative remains after deletion. Across two distinct fitting regimes in a paired Materials Project study, the lower-floor regime also exhibits a larger prediction change on more than 50% of the shared requests. Systematic comparisons against approximate updates and the original model decouple deliberate target suppression from preserved overall model utility. Consequently, request-level unlearning evaluations should report reference loss, prediction change, and retained utility together, interpreting post-deletion accuracy against what retraining itself leaves behind.
cs.LG / 247 / 2609.35649
Transferable Mass Spectrum Prediction via Reference-Guided Test-time Specialization
Yunhua Zhong, Runting Li, Yifan Li, Pan Liu, Zhiwen Yang, Zikun Wang, Yixuan Tang, Jun Xia
cs.LG
Abstract
Tandem mass spectrum prediction supports compound identification across metabolomics, natural-product discovery, and environmental analysis. However, pretrained predictors often degrade under shifts in chemical space and acquisition conditions, while retraining domain-specific models from scratch is costly. We introduce SPARC, a retrieval-guided test-time specialization framework that adapts a pretrained predictor using a spectral reference library without accessing test-query spectra. For each target query, SPARC retrieves chemically related reference spectra to recalibrate fragment intensities within the learned fragmentation space. During Transfer, SPARC combines reference-guided spectral adaptation with reliability-aware consistency, using reconstruction behavior on retrieved spectra to selectively preserve trustworthy predictions during continual specialization. Across MassSpecGym, NPLIB1 and application-specific GNPS libraries, SPARC improves spectral prediction under multiple transfer settings. These results establish retrieval-guided test-time specialization as a practical strategy for extending pretrained MS/MS predictors to specific chemical and acquisition domains, with continual test-time training providing further refinement during deployment.
cs.LG / 248 / 2609.35686
Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si
cs.LG · cs.AI
Abstract
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
cs.LG / 249 / 2609.35698
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang
cs.LG
Abstract
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}ε\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.
cs.LG / 250 / 2609.35703
A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
Fred Xu, Thomas Markovich, Florence Regol, Yizhou Sun
cs.LG · cs.AI
Abstract
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filters capture structural variation, and a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. The resulting doubly-spectral stochastic (DSS) expansion supplies task-matched readouts from one representation: the mean coefficient encodes class evidence for the energy-based OOD score, the higher-order coefficients encode structured logit variation, and quadrature averaging over the chaos coordinate defines the single predictive distribution used for prediction and calibration. A capacity theorem shows that, under a full-rank feature assumption, a restricted subfamily matches the chaos coefficients of any Gaussian-latent random graph signal, with exponentially decaying truncation error under a growth condition; the task-level claims are established empirically. DSS-GNN has two deployment modes: standalone, or as a residual branch beside a deterministic encoder (DSS-Hybrid). Standalone DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD detection, and the strongest shifted accuracy on all 7 GOOD concept-shift benchmarks under standard empirical risk minimization (ERM). Cross-evaluating both modes on all three tasks shows that each remains effective on the other's tasks, with documented exceptions, and yields explicit deployment guidance.
cs.LG / 251 / 2609.35707
ScAn-Bench: Evaluating Scaling Analysis Methodology
Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog, Neeratyoy Mallik, Jenia Jitsev, Danny Stoll
cs.LG
Abstract
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.
cs.LG / 252 / 2609.35715
X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
Prithwish Dan, Chenyang Ma, Wei Zhan
cs.LG · cs.AI · cs.RO
Abstract
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.
cs.LG / 253 / 2609.35750
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre Côté, Laurent Charlin, Guillaume Lajoie
cs.LG · cs.AI
Abstract
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
cs.LG / 254 / 2609.35751
How to Loop MoE: Flatten the Experts, Untie the Attention
Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly, Wang Yang, Xiaoqing Tong, Qianying Liu, Xiaotian Han, Vipin Chaudhary
cs.LG · cs.AI · cs.CL
Abstract
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.
cs.LG / 255 / 2609.35752
Neural Harmonic Measure Operator
Jinjin He, Sinan Wang, Yuchen Sun, Bo Zhu
cs.LG · math.NA
Abstract
We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes the density of this measure as a transformer-based boundary kernel supervised by Walk-on-Spheres exit samples, so one trained kernel handles different boundary values on a shape with no retraining. We extend it to Poisson via a classical decomposition, with an auxiliary network amortizing the source-induced correction and avoiding the singular volume quadrature that breaks direct evaluation. At inference, new boundary values and new sources both yield PDE solutions by re-integration against the fitted kernel and lift, with no retraining. NHMO improves over four prior baselines on the MCB-B 3D variable-shape Poisson benchmark across all five categories, and is competitive with major neural-operator baselines on a controlled 2D testbed.
cs.LG / 256 / 2609.35763
Unifying Distributional Training for One-Step Visual Generation
Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
cs.LG
Abstract
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/
cs.LG / 257 / 2609.34018
Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control
Denis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz, Wojciech Samek, Marc Toussaint
cs.RO · cs.LG
Abstract
Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.
cs.LG / 258 / 2609.34261
RoboICL: Embodied In-Context Learning with GPT-6 Astra
Fangcheng Liu, Yeqing Shen, Anda Cheng, Weishi Mi, Chao Tang, Chenyuan Liu, Yushun Xiang, Tingguang Li, Yong-Lu Li, Yehui Tang
cs.RO · cs.LG
Abstract
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the $π_{0.5}$ + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48\%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.
cs.LG / 259 / 2609.34484
ARS: Agentic Reward System for Robot Learning
Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng, Weiwei Zhang, Kai Xie, Xiaofeng Mou, Yi Xu
cs.RO · cs.LG
Abstract
Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars
cs.LG / 260 / 2609.34666
On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation
Xintong Yang, Minglun Wei, Yu-Kun Lai, Ze Ji
cs.RO · cs.LG
Abstract
Differentiable physics is increasingly used in robotic material manipulation for system identification, trajectory or skill optimization, demonstration generation, and robot or end-effector design. These applications depend on gradients propagated through long, contact-rich simulation rollouts. We study the numerical reliability of those gradients using two Material Point Method (MPM) system-identification benchmarks derived from elastoplastic and granular manipulation. The benchmarks provide controlled cases for three effects that also arise in broader differentiable physics-based optimization. GPU many-to-one sums whose order depends on thread scheduling changed long-horizon gradients and reversed the sign of one parameter gradient relative to a deterministic reference. Finite-difference checks became less reliable for longer rollouts because repeated-run loss variation grew much faster than the loss change produced by the tested parameter perturbations. Observation and loss definitions changed optimization behaviour and the solution preferred by an independent metric. These results motivate reproducible accumulation, finite-difference validation that compares perturbation-induced loss changes with repeated-run variation, and explicit reporting of objective construction when differentiable simulation is used for robotic optimization.
cs.LG / 261 / 2609.34684
Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
Hyungjoon Kim, Wonbin Son, Mi Young Lee, Jun Young Lee, Seungmin Rho
cs.RO · cs.CV · cs.LG
Abstract
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.
cs.LG / 262 / 2609.34806
On Temporal Binding in Large Audio Language Models
Paul Primus, Gerhard Widmer
cs.SD · cs.LG
Abstract
Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model components that may need improvement. Using mechanistic interpretability, we investigate how temporal information is represented and bound to sound events in three open-source LALMs. We find that across all three, event-specific location becomes concentrated in event name representations at intermediate modality integration layers. These representations encode coarse event position along a low-dimensional, curved relative time trajectory. Steering event name representations along this trajectory systematically shifts before/after beliefs, providing evidence that these representations contribute to coarse temporal reasoning. In contrast, the same interventions do not reliably shift predicted onset timestamps, suggesting that coarse temporal reasoning and precise metric event localization rely on distinct mechanisms.
cs.LG / 263 / 2609.35005
Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański, Szymon Klimaszewski
cs.SD · cs.LG
Abstract
Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.
cs.LG / 264 / 2609.35054
Perceptual Quality Loss or Loss of Perceptual Quality?
Danilo de Oliveira, Tal Peer, Maurício do V. M. da Costa, Timo Gerkmann
eess.AS · cs.LG
Abstract
Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily correlate with an improved listening experience. Through objective and subjective experiments, we assess the performance of SE models trained with two different types of auxiliary PESQ loss terms. The numerical evaluation on a suite of standard metrics suggests that, while models optimized for PESQ naturally obtain higher PESQ scores in the test set, for most other metrics the scores do not significantly change. In some cases, the PESQ loss even results in worse PESQ scores on mismatched data. A formal listening experiment reveals that the models without a PESQ loss were generally preferred over models that include it, across all settings. Finally, we analyze the relative importance of PESQ in the composite metrics CSIG, CBAK and COVL, and find that PESQ dominates all of them. Our study highlights the perils of over-reliance on PESQ and stresses the importance of a complete evaluation procedure for SE.
cs.LG / 265 / 2609.35295
Simulation-Based Inference for Plate Reverb System Identification
Dylan Sechet, Marc Evrard, Matthieu Kowalski
eess.AS · cs.LG
Abstract
We address Task A of the 1st DAFx Parameter Estimation Challenge, which aims to retrieve the physical parameters of a plate model from an impulse response. To do so, we use the Simulation-Based Inference (SBI) framework, in which we train a neural network to estimate a density over plate parameters given an impulse response, using a dataset generated by the simulator. Inference for a new impulse response then requires only a forward pass through the network, without involving the simulator. For each test observation, we fine-tune a specific network: additional simulation rounds are performed by sampling parameters from the current estimated distribution, simulating the corresponding impulse responses, and fine-tuning to produce the specialized network.
cs.LG / 266 / 2609.33880
Neuron-Level Architecture Growth: A Controlled Evaluation for EEG Time-Series Decoding
Adam Mounir, Stella Douka, Arnault H. Caillet, Bruno Aristimunha, Sylvain Chevallier
eess.SP · cs.HC · cs.LG · cs.NE
Abstract
Convolutional EEG decoders are trained at a fixed width, usually set by their authors on other data. Growing methods add neurons during training where the loss could decrease the most, but whether they improve compared to a reference width is untested on EEG. Here, we grow three convolutional backbones on 12 motor-imagery datasets under three protocols and compare each with its reference model per subject. The growing ShallowFBCSPNet scores 2.9 points above its reference model with only half the parameters (0.57x), SCCNet changes by at most 1.2 points. Deep4Net growing models show decreased accuracy, but they require adaptation that prevent to compare faithfully the results. These differences follow the selection step, which keeps a candidate neuron relying on a dynamic threshold from singular values decomposition. Overall, these results suggest that growth helps when its criterion can rank the candidate neurons, and that the rate of skipped neuron addition tells where a decoder can be grown small from scratch.
cs.LG / 267 / 2609.34670
Hierarchical Clustering and Signal Denoising on Digraphs
Yi Wang, Sippanon Kitimoon, Hrushikesh N. Mhaskar, Xiaosheng Zhuang
eess.SP · cs.LG
Abstract
In this paper, we propose a representation of a digraph (directed graph) as a Hermitian matrix derived from its adjacency matrix. This representation characterizes both the connectivity and the edge orientation of the digraph. Based on the spectral decomposition of the Hermitian matrix, a digraph clustering algorithm with $k$-means is introduced to produce a partition on the graph. Applying this algorithm (bottom-up) recursively to a digraph with partially labeled vertices yields a spectral hierarchical digraph clustering (\myproj) algorithm that produces consistent nested partitions of the digraph, or equivalently, a tree structure. Furthermore, based on the in-degree and out-degree of each cluster in the digraph clustering, a pair of hierarchical interval partitions (filtrations) can be derived in a top-down manner to produce a pair of nested knot sequences. These knot sequences facilitate the construction of multilevel spline quasi-interpolants, enabling a noisy graph signal to be decomposed into a coarse approximation and inter-level details, followed by adaptive thresholding and reconstruction. Experiments on synthetic and real-world digraphs demonstrate the superiority of our {\myproj} algorithm for digraph clustering across diverse graph structural properties (homophily and heterophily) and supervision settings. Moreover, experiments on digraph signal processing using multilevel spline quasi-interpolants further demonstrate the effectiveness of signal recovery on digraphs in terms of RMSE and SNR.
cs.LG / 268 / 2609.35758
Statistical Learning of Contractive Dynamical Representations for Composite Adaptive Control
Min Kim, José Leonardo Brenes, Fred Hadaegh, Soon-Jo Chung
eess.SY · cs.LG · cs.RO
Abstract
We present a representation-learning framework for composite adaptive tracking control under dynamically coupled disturbances. The framework connects classical disturbance-accommodating control (DAC) to recent last-layer adaptive disturbance-rejection methods. Specifically, we introduce a statistically principled hard expectation-maximization (hard-EM) procedure, with a Kalman smoother in the hard E-step, to identify dynamical representations of disturbance whose latent evolution is uniformly contractive. The learned representation evolves a latent disturbance-excitation state from measured plant features and control inputs and decodes that state into the time-varying disturbance acting on the nominal plant, thereby extending prior "fixed-decay" last-layer adaptive methods to a learned, predictive DAC-style formulation. Combined with Bayesian filtering of the learned latent state, this representation yields a composite adaptive tracking controller with predictive capability and provable exponential convergence to a bounded neighborhood. We validate our approach experimentally on a slippery ground vehicle carrying a liquid-sloshing tank and a pendulum load, and we further assess its robustness on a system of coupled Duffing oscillators. Across both settings, the method achieves accurate disturbance prediction and improved overall tracking performance relative to fixed-decay representation-learning ablations, LTI disturbance-accommodating baselines, and model-based PD baselines.
cs.LG / 269 / 2609.33814
Annealed Sinkhorn with Momentum: Certified Unregularized Optimal Transport in Linear Memory
Samuel J. K. Chin, Maximilian Schiffer
math.OC · cs.LG
Abstract
We characterize Bregman Douglas-Rachford splitting (BDRS) for unregularized discrete optimal transport and develop an anytime primal-dual certificate in linear memory. We first establish that BDRS coincides with warm-started Inexact Proximal point method for exact Optimal Transport (IPOT) using a single inner Sinkhorn iteration. By eliminating the primal transport plan from the updates, we derive an equivalent dual formulation that reveals BDRS as annealed Sinkhorn under an implicit inverse-linear temperature schedule, with an additional log-scaling momentum term and a cooler kernel. While this explains the role of the temperature parameter in BDRS as an initial temperature, it also reduces the solver's memory requirement from quadratic to linear. Utilizing this annealing perspective, we introduce overrelaxed BDRS, which combines annealing and overrelaxed scaling within a single recursion. We derive a primal-dual certificate for both methods that can be evaluated in linear memory without transport plan construction, thus providing a computable stopping rule. On the DOTmark benchmark, combining momentum with the cooler kernel produces substantially smaller optimality gaps than annealed Sinkhorn under the same schedule. For pixel-level color transfer between $1024\times1024$ images, BDRS attains a lower repaired transport cost than MDOT-TNT with a 9$\times$ speed up, reaching a relative duality gap of $1.59\%$ in 24 minutes. We further demonstrate a color transfer with $4238\times2365$ images, yielding 10 million pixels per image and approximately one hundred trillion implicit transport entries, reaching a best relative duality gap of $2.41\%$ and $2.80\%$ within 35 hours in each direction on a single NVIDIA L40S GPU.
cs.LG / 270 / 2609.33879
DCEmbed: Scalable Optimization over Neural Surrogates
Akshay Sreekumar, Nicolas Christianson, Priya L. Donti, Ellen Vitercik, Ram Rajagopal
math.OC · cs.LG
Abstract
Neural surrogates can accelerate large-scale optimization by replacing expensive or intractable model components with efficient learned approximations, but solving the resulting embedded problems can remain prohibitively costly. For instance, standard exact encodings of neural networks with ReLU activations allow the problem to be solved by mixed-integer solvers, but add large numbers of binary variables to accommodate the nonlinearity of the activations, which can render the problem computationally prohibitive. To address this, we propose DCEmbed, a heuristic for optimization problems with embedded neural surrogates that leverages the difference-of-convex (DC) representation of the network and avoids adding activation binaries. Exploiting shared structure within the DC representation of a ReLU network, we derive a reduced-size, exact formulation for its convex components that can be embedded in optimization problems using just two linear inequalities and one continuous auxiliary variable per hidden neuron. Using this, the problem is solved via an iterative penalty convex-concave procedure, where only the concave portions of the neural terms are approximated at each stage. The original objective, constraints, and any discrete decisions are retained, allowing standard convex or mixed-integer optimization solvers to optimize the host and surrogate jointly at each iteration. In experiments on quadratic programs, mixed-integer resource allocation, and neural two-stage stochastic programming, our method demonstrates much faster progress toward high-quality feasible solutions than approaches using exact mixed-integer embeddings. In particular, DCEmbed achieves $4\times$ lower normalized primal integral than the best exact baseline on the resource allocation problem, while in two-stage stochastic programming it reaches the global surrogate optimum $\sim 5\times$ faster than Gurobi ML.
cs.LG / 271 / 2609.34791
Finite-Time Concentration and Convergence Rates for Projected Two-Time-Scale Stochastic Approximation with Markov Noise
Rahul Singh, Vivek S. Borkar, Eric Moulines
math.OC · cs.LG
Abstract
We study finite-time concentration and convergence rates for projected two-time-scale stochastic approximation driven by a controlled Markov chain. The averaged fast map is contractive, while the slow iterate is projected onto a compact convex polyhedron. The associated projected ordinary differential equation may have a discontinuous vector field at the boundary, preventing a direct application of standard analyses based on Lipschitz vector fields. Using the Skorokhod map, we establish explicit high-probability bounds for tracking the moving fast equilibrium and the projected slow dynamics. These bounds separate martingale fluctuations, Markov-noise residuals, and the bias due to time-scale separation. A Lipschitz Lyapunov function satisfying a uniform decrease condition over fixed time intervals yields almost-sure convergence, with explicit last-iterate rates when the decrease admits a power lower bound. Under uniform Lyapunov contraction, polynomial step sizes yield joint fast-tracking and slow Lyapunov-error exponents arbitrarily close to $1/3$. Under the additional assumption that the reduced slow update map is a Euclidean contraction, logarithmically separated step sizes improve the joint rate to $O(n^{-1/2}\log n)$ almost surely, including for boundary equilibria. The same rate holds under a distinct geometric condition involving a strictly attracting face of a box and a fast equilibrium that is constant on that face. An actor-critic application achieves an almost-sure value-gap rate of $O(n^{-1}\log n)$ relative to the optimum within the constrained policy class. Further applications include projected TD(0) and projected stochastic gradient descent. We also extend the analysis to projection of the fast recursion under Euclidean contractivity.
cs.LG / 272 / 2609.35152
Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach
Guodong Ma, Baofeng Sun, Wenyu Yang, Zhihong Yao
math.OC · cs.LG
Abstract
Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transferability of model-based control. We propose a hybrid framework to coordinate lane-level variable speed limits (VSLs) and ramp metering across SWSs. First, we reconstruct L-METANET, a lane-level macroscopic traffic flow model that captures free and forced lane changes. Second, we combine XGBoost-SHAP with a random-parameters binary logit (RPBL) model to derive analytical equations for merging and diverging collision risks and formulate system cost and reward functions. Third, we develop MPC-STMAPPO, a hierarchical controller integrating model predictive control (MPC) and multi-agent reinforcement learning (MARL). Its upper MPC layer uses L-METANET for long-horizon rolling optimization and generates baseline commands; its lower spatiotemporal MAPPO (ST-MAPPO) layer, enhanced with Mamba cells and graph attention, produces residual actions for short-horizon adjustment. Real-world experiments on the 18-km Eastern Expressway in Changchun, China, show that L-METANET accurately reproduces lane-changing-induced flow redistribution and capacity drops, with state evolution aligned with ground truth. XGBoost-SHAP-RPBL achieves AUCs above 0.80 in most tasks, outperforming conventional logit models. MPC-STMAPPO converges faster and performs better across multiple metrics than MPC- and MARL-based baselines. Under randomly fluctuating demand, it also significantly outperforms pure MARL in generalization, demonstrating strong potential for industrial deployment.
cs.LG / 273 / 2609.35418
Convex Optimization Is Free When Accuracy Is Expensive
Arthur Paing, Arthur Jacot
math.OC · cs.LG · stat.ML
Abstract
This paper studies convex optimization when the gradient cannot be evaluated exactly, but only approximated by a hierarchy of algorithms whose compute grows like $δ^{-γ}$ in the accuracy $δ$. When $γ>2$, falling into the Harder-Than-Monte-Carlo (HTMC) regime, the price of accuracy outruns the variance reduction that Monte Carlo would buy and we show that minimizing a loss function costs no more, up to a factor depending only on $γ$, than a single evaluation of its gradient at the accuracy the problem demands. A randomized multilevel oracle replaces the deterministic approximation of accuracy $δ$ by an unbiased estimator of it, whose variance $σ^2$ becomes a second, independently priced dial: the cost of one call drops from $δ^{-γ}$ to $δ^{2-γ}σ^{-2}$. Plain inexact gradient descent driven by that oracle reaches loss $\varepsilon$ at expected compute $Θ(\varepsilon^{-γ})$ in the convex case, against $Θ(\varepsilon^{-(γ+1)})$ for the same method run at a fixed accuracy: randomization buys a full power of $\varepsilon$. Under $μ$-strong convexity the exponent halves, to $\varepsilon^{-γ/2}$, because the iterates settle at a noise floor and the bias budget relaxes accordingly. Both bounds are independent of the step size, and hence of the smoothness constant, and we show that the cost is a functional of the underlying gradient flow rather than of any discretization of it.
cs.LG / 274 / 2609.35665
Learned Preconditioning for a Primal-Dual Interior-Point Method
Abhinav Madabhushi, Jialin Liu, Minxin Zhang
math.OC · cs.LG
Abstract
Interior-point methods (IPMs) are among the most widely used algorithms for constrained optimization, yet their Newton-based search directions require costly second-order information and large linear-system solves. Learning to optimize offers cheaper updates learned from data, but the singular behavior of logarithmic barriers near constraint boundaries makes IPMs highly sensitive to perturbations, complicating both warm starting and learning reliable updates. We introduce pdLIP, an IPM for smooth nonlinear programs that integrates learned preconditioning with pdProj, an all-shifted primal-dual projected-search IPM. A shared coordinate-wise recurrent network predicts a positive diagonal preconditioner that scales the right-hand side of the reduced Newton system for the primal step, and the remaining slack and multiplier directions are recovered analytically. The learned iterations avoid Hessian evaluations and Newton-system solves, using only first-order and coordinate-wise operations amenable to GPU parallelization. Training is self-supervised, with a loss based on a penalty-barrier merit function and the residual of perturbed optimality conditions, requiring neither target directions nor precomputed solutions. Primal and dual shifts mitigate the barrier's sensitivity to perturbations near constraint boundaries, enabling effective warm starting. Across four classes of 200-dimensional convex and nonconvex constrained problems, pdLIP warm starts reduce pdProj refinement iterations by 63-67% compared with cold starts at the same KKT residual tolerance of $10^{-8}$, with negligible warm-start generation cost relative to the subsequent pdProj solve. Improvements persist on box-constrained QPs with 1000 variables and extend to applications including portfolio optimization, support vector machines, and a nonlinear control example.
cs.LG / 275 / 2609.34836
MW-Nowcast: Six-hour ensemble nowcasting of extreme precipitation
Ning Wang, Zuliang Fang, Weixin Jin, Zhongjian Lv, Shuang Qin, Pengcheng Zhao, Siqi Xiang, Jiang Bian, Haoyi Xiong, Nan Guan, Bin Zhang, Liangjie Zhang, Denvy Deng, Qi Zhang, Matt Corey, Jitu Keshri, Sridhar Iyer, Hongyu Sun, Kit Thambiratnam, Jonathan Weyn, Richard E. Turner, Haiyu Dong
physics.ao-ph · cs.LG
Abstract
Extending reliable nowcasting of extreme precipitation could provide critical additional time for warnings and emergency response during high-impact events such as flash floods. Radar-based generative machine-learning models have enabled skilful hyperlocal precipitation nowcasting, but accurate prediction of intense precipitation remains confined to the first few hours. Because storm-scale structure is predictable for longer than individual cells, a natural strategy is to predict that structure while generatively modelling only the uncertain local growth, decay, reorganisation and initiation of storms. Here we present Microsoft Weather Nowcast (MW-Nowcast), a six-hour ensemble radar nowcasting model that jointly learns a deterministic predictor to capture organised precipitation structure shared across ensemble members, and a generator to produce diverse local residuals around this shared prediction. Across independent test data from the United States, Europe and China, MW-Nowcast achieves higher detection skill than leading methods for heavy and extreme precipitation throughout the 6 h horizon. For the most intense rainfall, MW-Nowcast doubles the available warning time across all three regions, delivering 6 h forecasts with skill previously limited to 3 h for the leading generative baseline. A cost-loss decision analysis shows that MW-Nowcast retains substantial value for a broad range of applications even at 4-6 h, where alternative methods offer little benefit. These additional hours can give forecasters and emergency managers the time to warn and act before extreme rainfall strikes, helping to protect lives and property.
cs.LG / 276 / 2609.34916
Physics-Informed Neural Networks for Depth-Averaged Avalanche Dynamics
Pradyumn Singh Sikarwar, Vishal Sharma, Gaurav Bhutani
physics.flu-dyn · cond-mat.soft · cs.LG
Abstract
Accurate prediction of avalanche motion is essential for hazard assessment in mountainous terrain. This study develops and evaluates a physics-informed neural network (PINN) framework for the Savage-Hutter model of depth-averaged granular flow, progressing from 1D analytical verification to 2D experimental validation. First, three 1D problems of increasing complexity were verified against the analytical solution: height prediction with prescribed velocity, velocity prediction with prescribed height, and coupled prediction of both fields using the conservative formulation. The decoupled tests accurately reconstructed the spatio-temporal evolution of each field when the other was prescribed. The coupled formulation learned both fields without prescribed data, achieving mean height and velocity RMSEs of 0.043 and 0.079 in non-dimensional units. A hyperparameter sensitivity study evaluated the effects of network depth, width, collocation density, learning rate, and epochs. The framework was then extended to 2D and validated against laboratory experiments of a cylindrical granular pile collapsing on an inclined plane, with TITAN2D providing numerical comparisons. Purely physics-based training converged to the trivial zero solution; augmenting the loss with 10 sparse training points from final deposit profiles produced a physics-informed, data-assisted hybrid framework. Peak flow depth, depth-averaged velocity, RMSE, and wetted-area IoU evaluated global and local agreement. Global height RMSE ranged from 2.7 to 6.7 mm across four experimental cases, while mean wetted-area IoU ranged from 69 to 81 %, demonstrating consistent performance across variations in pile mass and slope angle.
cs.LG / 277 / 2609.35083
Continuous Variational Synthesis
Alan N. Amin, Mattia G. Gollub, Andrei Slabodkin, Elizabeth B. Wood, Eli N. Weinstein
q-bio.BM · cs.LG
Abstract
Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models is challenging: constraints on chemical synthesis can force many parameters into a discrete space, limiting the ability to pre-train and fine-tune. In this article we train ``free'' variational synthesis models using stochastic gradient descent in continuous space, and then discretize with post-training quantization to impose hardware and wetware constraints. This enables variational synthesis models to satisfy stringent reward criteria, while still synthesizing diverse designs, achieving a strictly dominating quality-diversity Pareto frontier. We demonstrate by training variational synthesis models of enzymes, peptides, antibody CDRH3s, and regulatory DNA elements. In silico performance is maintained in vitro.
cs.LG / 278 / 2609.35214
Uncertainty Quantification in Cardiac Model Personalisation from Ultrafast Ultrasound
Camilla Ferrario, Maelys Venet, Olivier Villemain, Maxime Sermesant
q-bio.QM · cs.LG · physics.med-ph
Abstract
Cardiac model personalisation requires inferring mechanical parameters that are not directly measurable in vivo. Ultrafast ultrasound shear wave elastography (SWE) enables non-invasive tracking of myocardial stiffness dynamics over the cardiac cycle, providing a target for personalisation. However, mapping these observations to subject specific model parameters remains ill-posed, as multiple parameter sets can reproduce the same stiffness dynamics. We formulate SWE-informed personalisation as a statistical inference problem using simulation-based inference (SBI). Using a subject-adapted 0D cardiovascular model and neural posterior estimation, we estimate model-conditional posterior distributions over active stiffness scale k0, contraction rate kATP, and relaxation rate kSR, conditioned on SWE-derived curve features and subject specific context. Among six healthy volunteers, four passed objective prior-support diagnostics and were retained for quantitative posterior analysis. Curve-level RMSE against the observed SWE target decreased from 12.61 $\pm$ 5.55 kPa for the prior predictive median to 1.14 $\pm$ 0.38 kPa for the posterior predictive median, an 89.7 $\pm$ 4.2% reduction. Posterior analysis revealed parameter-specific uncertainty, k0-kATP compensation, weaker constraint of kSR, and the importance of prior-predictive diagnostics for assessing whether each subject is represented within the modelled SWE feature space. These results support SBI for uncertainty aware SWE-based personalisation, while identifying prior support and forward-model adequacy as key diagnostics.
cs.LG / 279 / 2609.34995
Simulation-Based Quantum System Inference with Neural Posterior Estimation
Hang Zou, Anton Frisk Kockum, Martin Rahm, Simon Olsson
quant-ph · cs.LG
Abstract
Models of quantum systems faithfully map system parameters to observations, but the inverse problem of parameter inference from measurement data presents a fundamental challenge: computationally intractable likelihoods due to an exponentially large Hilbert space. Here, we introduce simulation-based quantum system inference, a unified, likelihood-free framework that learns parameter posteriors directly from classical simulation data. The central idea is to pair polynomial-cost classical simulators, such as Pauli propagation and tensor networks, with normalizing flows or other neural density estimators for accurate, reusable inference. A single model, trained once, maps any new measurement record to its posterior in one forward pass---turning per-experiment inference into a fixed, up-front cost. We numerically demonstrate the framework's versatility across Pauli noise learning, quantum error mitigation, quantum state tomography, and Hamiltonian learning, with examples involving 81-qubit shallow circuits and 735-parameter inference. In each case, the approach yields accurate estimates of identifiable parameters, while posterior uncertainty provides additional diagnostics of non-identifiability and indicates where further characterization is needed. Our framework reduces data-acquisition requirements in quantum experiments and accelerates parameter inference, providing a practical route to characterizing and improving large-scale quantum systems.
cs.LG / 280 / 2609.33844
ViBR-WM: Visual Bayesian Regression for World Modeling
Jifan Li, Ning Ning
stat.ME · cs.CV · cs.LG
Abstract
Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle dynamics. Visual compression reduces representation dimension, while Bayesian variable selection reduces active regression dimension. Posterior prediction combines forecasts across predictor subsets using their posterior probabilities as weights and accounts for parameter uncertainty and future disturbances. The model forecasts joint visual--physical states recursively and physical targets directly. Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task. Repeated fitting and resampling support these overall gains.
cs.LG / 281 / 2609.34115
Forecast-Necessary Causal Discovery for Nonlinear Political Panel Data: Feedback, Functional Form, and the Dynamics of Democratization
Michael Coppedge, Dmitry Zaytsev, Valentina Kuskova
stat.ME · cs.LG
Abstract
A non-significant coefficient in a dynamic panel model need not imply the absence of a relationship. It may instead reflect heterogeneous effects averaged toward zero, reciprocal dynamics overlooked by a recursive specification, or relationships masked by the omission of correlated covariates. Standard linear estimators cannot distinguish among these possibilities. We develop an inferential workflow for political panel data that resolves this ambiguity by combining flexible autoregressive estimation, forecast-necessity testing, functional characterization, and same-data linear benchmarking. The workflow first identifies relationships required for out-of-sample prediction, then characterizes their functional form across political contexts, and finally, distinguishes differences arising from estimator flexibility from those due to model specification. Applied to the causal sequence model of democratization on the V-Dem panel of 113 countries, the workflow reproduces the model's central finding - the protective belt of civil society, the rule of law, and institutionalized parties - while recovering reciprocal relationships from democracy to its institutional supports that a linear model cannot detect. Most importantly, three weak published direct effects, of which two are null, and one is marginally significant, receive three different diagnoses: one dissolves under the full specification, one reflects heterogeneous effects averaged toward zero, and one was masked by the reduced variable set. The workflow corrects the published record in both directions, removing one relationship and recovering two. More broadly, the workflow provides a framework for evaluating dynamic political theories under a model class capable of representing nonlinear and reciprocal mechanisms while preserving relationship-level interpretation and explicit inferential standards.
cs.LG / 282 / 2609.33968
Two-Sample Testing for Inhomogeneous Random Graphs in Non-Integral $L_r$ Norms
Soham Dan
stat.ML · cs.LG
Abstract
Testing whether two populations of networks share the same edge probabilities is a basic problem in network inference. How hard it is depends on the norm used to measure the difference. For the inhomogeneous Erdős--Rényi (IER) model, the optimal sample complexity is known for every integer $L_r$ norm and for $1\le r<2$. For non-integral $r>2$, however, the known upper and lower bounds do not match, and the lower bound was conjectured to be tight. We study this gap for two-sample testing on aligned vertices. We propose a test that runs two published statistics, of orders $2$ and $\lceil r\rceil$, on the same data and rejects if either one rejects. Its thresholds come from Hölder interpolation, so that both statistics have the same sample cost. We prove that this test attains the conjectured rate. Combined with earlier results, this shows that for every fixed $r\ge1$ the minimax sample complexity is of order $n^{\max\{4/r-1,\,2/r\}}/ε^2$, even when the separation changes with $n$. In simulations with $n$ between 32 and 256, the number of graphs needed for 80\% power at level $0.05$ grows with $n$ at a rate consistent with the theory. For $r=2.5$, for example, the fitted exponent is $0.78$, against the theoretical value $0.8$. Interestingly, the two statistics split the work as the interpolation argument suggests: the higher-order statistic is more powerful when only a few edges change, and the $L_2$ statistic when many edges change.
cs.LG / 283 / 2609.34043
Singularities of Non-negative Matrix Factorization and their application to Bayesian inference
Naoki Hayashi, Yota Maeda, Yasushi Esaki
stat.ML · cs.LG · math.ST
Abstract
Non-negative matrix factorization (NMF) is a singular statistical model whose Bayesian asymptotics are governed by the real log canonical threshold (RLCT). We study the local geometry of the factorization map and derive an upper bound for the RLCT of NMF. Let $H$ be the model inner dimension and $H_0$ the non-negative rank of the true $M\times N$ matrix. Assuming that the true matrix admits a strictly positive factorization of inner dimension $H_0$ in the interior of the parameter domain, we prove, for smooth positive priors, that $λ\leq \{(H-H_0)\min(M,N)+H_0(M+N-H_0)\}/2$. This bound strictly improves the previous bound when $H_0\geq3$. The proof uses a local analytic normal form that separates independent linear coordinates from a residual matrix product. When $H=H_0$ also equals the ordinary rank of the true matrix, we obtain the exact value $λ=H_0(M+N-H_0)/2$. Under the standard assumptions of singular learning theory, these results bound the leading coefficients of the expected Bayesian generalization error and the Bayesian free energy.
cs.LG / 284 / 2609.34050
The Statistical Cost of Causal Discovery with Feedback
Sunmin Oh, Seungsu Han, Gunwoong Park
stat.ML · cs.LG
Abstract
What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges between components. We establish the first information-theoretic lower bounds on sample complexity for this target. For $p$ variables, maximum SCC size $s_{\max}$, and maximum external-parent count $d_B$, any estimator requires order $s_{\max}\log(ep/s_{\max})+d_B\log(ep/d_B)$ samples in the worst case over a regular model class. These bounds distinguish the costs of SCC membership and external-parent selection. Under principal invertibility and without correlation faithfulness, we establish a population block-exogeneity principle that identifies unknown root SCCs through residual independence and inclusion minimality. A sparse-adjustment characterization shows that small adjustment sets suffice to identify SCCs and their direct external parents, without regressing on all previously recovered variables. These characterizations yield BlockExo, which attains a structurally matching sample bound without knowing $s_{\max}$ or $d_B$ under suitable conditions. Simulations support the structural dependence of our sample bound and demonstrate BlockExo's sample-efficient recovery in comparisons with other methods for cyclic causal discovery.
cs.LG / 285 / 2609.34161
GT-PSSM: Unified Probabilistic Framework for Stochastic Dynamics Modeling and Dependency Learning in Multivariate Time Series Anomaly Detection
Wonmo Koo, Jaeyeong Lee, Taeseong Yoon, Heeyoung Kim
stat.ML · cs.LG
Abstract
Multivariate time series anomaly detection (MTAD) is crucial for ensuring the safe and reliable operation of complex systems. Many existing methods learn normal patterns by training reconstruction or forecasting models on predominantly normal data. However, a large portion of these approaches rely on deterministic models and their associated point-wise output errors for anomaly scoring. Since real-world multivariate time series are inherently stochastic due to measurement noise and intrinsic system randomness, purely error-based scores can be unreliable, as large errors may arise from benign fluctuations rather than true anomalies. Probabilistic approaches address this limitation by quantifying uncertainty in model outputs. In particular, probabilistic state-space models (PSSMs) provide a principled framework by modeling stochastic system dynamics through latent state transitions and measurement noise via emission models. Despite this advantage, existing PSSM-based MTAD methods often struggle to capture long-range temporal dependencies and inter-variable dependencies, as they typically rely on noise-sensitive recurrent architectures and lack explicit cross-variable structure modeling. To address these limitations, we propose Graph-Transformer-Enhanced Probabilistic State-Space Model (GT-PSSM), a novel PSSM-based MTAD method that tightly integrates PSSM-based probabilistic modeling of stochastic dynamics with Graph Transformer-based learning of temporal and inter-variable dependencies. By jointly modeling stochasticity, long-range temporal dependence, and variable interactions within a unified probabilistic framework, GT-PSSM enables more robust anomaly detection.
cs.LG / 286 / 2609.34207
Functional Autoencoders for Amplitude-Phase Representation Learning
Peida Wu, Xinyang Xiong, Pengcheng Zeng
stat.ML · cs.LG
Abstract
Functional data are intrinsically infinite-dimensional, and often exhibit phase variation, where corresponding events occur at different times across observations. Existing linear dimension reduction methods struggle with nonlinear amplitude variation, while functional autoencoders without an explicit warp entangle temporal misalignment with shape. We propose the Amplitude--Phase Functional Autoencoders (AP-FAE), an unsupervised framework for functional data that spans both univariate and multivariate cases, with emphasis on the multivariate setting, and factorizes the latent space into separate amplitude and phase embeddings derived from all channels. A smooth functional decoder reconstructs channel-specific amplitude functions in canonical time, and a shared monotone, endpoint-preserving warp captures phase variation. We prove a bound linking amplitude recovery to registration, reconstruction, and noise errors, and validate it numerically. Across synthetic data and six real-world benchmarks, AP-FAE outperforms state-of-the-art baselines on most clustering and alignment metrics and on all reconstruction metrics. Clustering with amplitude embeddings alone consistently surpasses joint amplitude--phase clustering, confirming the benefit of explicit disentanglement. Code is available at https://anonymous.4open.science/r/APFAE-418C/}{https://anonymous.4open.science/r/APFAE-418C/.
cs.LG / 287 / 2609.34613
Probabilistic Geodesic Flow Matching on Location-Scale Families
Zeyuan Yu, Zhi Chang, Shiwei Lan
stat.ML · cs.LG · math.PR
Abstract
Flow matching (FM) has recently emerged as a promising framework for generative modeling due to its conceptual simplicity and strong empirical performance. In FM, samples are transported along a vector field parameterized by a neural network, inducing a probability path that evolves from a simple noise distribution to the target data distribution, governed by an ordinary differential equation (ODE). However, existing FM approaches predominantly rely on probability paths derived from optimal transport (OT) between Gaussian distributions, which may be suboptimal for capturing complex data with inhomogeneous structures such as heavy tail or sharp contrast. In this work, we generalize FM to the broader class of location-scale families for handling data inhomogeneity and introduce a novel class of probability paths defined as geodesics on the manifold of probability distributions. We name this approach probabilistic geodesic flow matching to distinguish it from prior geodesic (Riemannian) FM methods defined in input space. We argue that Euclidean OT-based paths are not necessarily optimal in probability space and may limit modeling flexibility. Through synthetic benchmarks and scientific datasets at different scales, we demonstrate that the proposed method more effectively captures complex distributions, leading to improved or comparable performance compared with SOTA geometry-motivated generative models.
cs.LG / 288 / 2609.34667
Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks
Etienne Boursier, Nicolas Flammarion
stat.ML · cs.LG
Abstract
Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.
cs.LG / 289 / 2609.34731
Information-Theoretic Analysis of Next-Token Prediction under Markovian Data
Masoud Kavian, Abdellatif Zaidi, Milad Sefidgaran
stat.ML · cs.IT · cs.LG
Abstract
We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.
cs.LG / 290 / 2609.34756
Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks
Alexandre Declèves, Etienne Boursier, Nicolas Flammarion
stat.ML · cs.LG
Abstract
Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
cs.LG / 291 / 2609.34887
Conformal Prediction and Conditional Coverage for Tabular Foundation Models
Sungwoo Park, Sunghee Park, Won Chang
stat.ML · cs.LG
Abstract
Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration Method), a lightweight application of highest predictive density split conformal prediction that accommodates multimodal predictions. Given calibration and test outputs, it requires no additional training or model inference. It provides finite-sample marginal validity under our assumptions. We bound conditional-marginal coverage gaps using distribution-estimation error and score discreteness, and examine coverage heterogeneity through percentile rank-score plots. Experiments with TabPFN and TabICL show improved marginal coverage accuracy and lower average conditional and group coverage errors. Under a fixed data budget, allocating more observations to calibration can reduce marginal coverage error despite less accurate point predictions.
cs.LG / 292 / 2609.35038
GUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian Optimization
Jintao Wei, Chenxi Li, Songhao Wang
stat.ML · cs.LG
Abstract
Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and task heterogeneity. We propose GUIDE-FBO, in which agents exchange compact distributions over the locations of their respective optima inferred from local Gaussian process (GP) posteriors, rather than raw observations, query points, or surrogate parameters. The server merges and reweights these distributional components before returning a subset to each agent. Each agent then constructs a Federated Interventional GP (FI-GP), which preserves the local posterior mean and spatially rescales its covariance for local decision making. For the upper confidence bound (UCB) instantiation, GUIDE-UCB, we prove that any bounded FI-GP uncertainty intervention preserves the leading-order cumulative regret rate of standard GP-UCB. When the transferred distributions place greater support near an optimum than in a suboptimal region, selecting the latter requires greater local posterior uncertainty. Experiments on 12 synthetic benchmarks and three real-world optimization tasks show that GUIDE-FBO remains effective across settings ranging from homogeneous to severely heterogeneous. Ablation results highlight the importance of spatially localized uncertainty intervention, while the communication analysis shows that GUIDE-FBO exchanges only compact distributional messages.
cs.LG / 293 / 2609.35217
A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty
Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin Jullum
stat.ML · cs.LG
Abstract
Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.
cs.LG / 294 / 2609.35429
Multi-Task Learning of Conditional Mean Operators: applications to dynamical systems and uncertainty quantification
Sami Chemlal, Thibaut Germain, Rémi Flamary, Vladimir R. Kostic, Karim Lounici
stat.ML · cs.LG
Abstract
Estimating conditional statistics and learning representations of a population of conditional distributions are central problems in many data-driven applications, including uncertainty quantification and dynamical systems analysis. Conditional mean operators (CMOs), a class of linear operators between function spaces, resolve these objectives by providing access to a broad class of conditional statistics. However, existing methods typically estimate each CMO independently or constrain it to prespecified function spaces, thereby preventing the exploitation of shared structure across related distributions. In this work, we posit that related CMOs share finite-dimensional input and output function spaces, and are specialized for each task with a linear operator mapping these spaces. Based on this hypothesis, we introduce MTL-CMO, a multi-task framework that jointly learns shared function spaces and task-specific operators across multiple datasets. We further introduce T-CMO, a transfer learning method that reuses the shared spaces to estimate, in closed form, the operator of a new conditional distribution. We establish statistical guarantees quantifying the benefits of jointly learning the shared function spaces. Our experiments demonstrate that learning shared function spaces improves uncertainty quantification across a broad range of conditional distributions and, when applied to Langevin and plasma dynamics, yields compact representations of complex dynamics that retain physically meaningful information and enable parameter identification.
cs.LG / 295 / 2609.35598
Learning Conditional Expectation Operators via Functional Newton Updates
Thiago Ramos, Alek Fröhlich, Daniel Perazzo, Massimiliano Pontil
stat.ML · cs.LG
Abstract
We introduce the Functional Spectral-Newton Method (FSNM) for learning the leading singular structure of a conditional expectation operator without fixing a basis or reproducing kernel Hilbert space. FSNM fits a low-rank representation of the centered joint-to-product density ratio kernel by alternating functional Newton updates. Each update reduces to a preconditioned regression, which we approximate with vector-valued regression trees in a stagewise boosting procedure. At the population level, we establish descent and an $O(1/T)$ best-iterate block-stationarity rate under a relative weak-learner accuracy condition, and show that every nondegenerate local minimum over the full centered $L^2$ spaces is a globally optimal rank-$d$ approximation. Synthetic experiments show that FSNM recovers a low-rank density ratio and its leading spectral structure, and that the same learned kernel can answer multiple conditional queries without refitting.
cs.LG / 296 / 2609.35622
Elicitation and Decision Geometry in Single-Index Bandits
Sakshi Arya, Cheng Soon Ong
stat.ML · cs.LG
Abstract
We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves $O(\log n)$ expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.
神经与进化计算 (cs.NE)
2
cs.NE / 1 / 2609.34034
ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
Matei-Ioan Stan, Oliver Rhodes
cs.NE · cs.AI · cs.DC · cs.LG
Abstract
A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realistic contender must be data-adaptive, able to capture long-range dependencies, and GPU-parallelisable, but also non-linearly recurrent to enable complex reasoning. Based on evidence suggesting the auditory cortex operates on fixed timescales, this work proposes the ADaptive with Prescriptive Timescales Network (ADPTNet) as a potential solution to achieving all four properties simultaneously. ADPTNet is built around local topological conjugates, obtained by a novel combination of linear attention and Riemannian optimisation, applied to static global dynamics. This enables non-linear yet predictable long-term behaviour. Dynamical systems theory proofs provide theoretical guarantees for the parametric control of ADPTNet's timescales (its Lyapunov spectrum). ADPTNet improves performance on Selective Copying over Hawk, the existing method balancing long-range memory and adaptability, while also improving state tracking over linear SSMs like Mamba. On sequential CIFAR-10, ADPTNet matches linear SSM accuracy and outperforms existing selective models (incl. the Transformer), using fewer parameters. We also introduce a neuromorphic SpikingADPTNet, which achieves a new state-of-the-art accuracy on the Spiking Speech Commands dataset ($83.56\%\pm0.15$). Finally, ADPTNet's constant timescales enable two efficient, Jacobian-free extensions to the DEER parallel simulation algorithm (Conv and Forward DEER) that retain the same average convergence. Conv DEER adds no computational overhead beyond the network's forward pass and enables non-linear RNN parallelisation via iterated convolutions for the first time.
cs.NE / 2 / 2609.35657
CMDO: A Cognitive Memory-Driven Optimization Algorithm for Adaptive Population-Based Search
Mohammed Yusuf Mujawar, Shahram Rahimi, Noorbakhsh Amiri Golilarz
cs.NE · cs.AI
Abstract
Population-based optimization methods often use previous search information through successful solutions, parameter adaptation, or operator performance, but they rarely retain the context in which a search behavior succeeded or failed. We introduce Cognitive Memory-Driven Optimization (CMDO), a derivative-free population-based optimizer that represents experience as the relationship between search context, search behavior, and observed outcome. CMDO organizes these experiences across working, episodic, and consolidated memory, retrieves them according to similarity with the current search state, and uses both positive and negative evidence to guide subsequent search. Retrieved experience does not replay previous candidate locations; instead, it selects search recipes that are reconstructed from the current population through exploratory, directed, and local search behaviors with adaptive search geometry. We evaluate CMDO on selected Blackbox Optimization Benchmarking test suite on COCO (BBOB/COCO) and Congress on Evolutionary Computation 2017 (CEC2017) problems against DE, CMA-ES, SHADE, GWO, HHO, and ORCA, and further study its application to seven-parameter photovoltaic model estimation using measured current--voltage data. The results show problem-dependent but competitive optimization performance, including the lowest median error among the compared methods on CEC2017 F10. More importantly, analysis of the search traces shows that context-dependent recall changes the distribution of executed search behaviors, while unsuccessful experiences remain available as negative evidence for later decisions, showing that accumulated experience directly influences subsequent search behavior. These results support the use of explicit context--behavior--outcome memory as an active mechanism for controlling population-based search.
计算语言学 (cs.CL)
74
cs.CL / 1 / 2609.33865
In-Context Adaptation of Encoder-Decoder Models in Speech Recognition
Yen Meng, Sharon Goldwater, Hao Tang
cs.CL · cs.SD · eess.AS
Abstract
In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.
cs.CL / 2 / 2609.33883
Lost with a Map: Conversational State and Behavioral Reliability in Language Models
Atahan Dokme, Larry Heck
cs.CL
Abstract
Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models from four families on MultiWOZ and SGD. Structure and values separate: which domains, slots, and requests are active is linearly readable just before the model acts, whereas exact values are far more readable where the user stated them. After a user changes a value, both values remain accessible at their mentions, and causal interventions show that both continue to influence the model's action. In natural closed-loop interaction, query failures separate into cases of weak structural support, incorrect value resolution, and failure to deploy otherwise-supported constraints, with targeted interventions producing systematically different repair behavior across these cases. These findings motivate a state-action controller that starts from the base model action and selectively edits it using structural readouts, without requiring a complete predicted belief state as an intermediate representation. On held-out MultiWOZ interaction across five models, it raises the base model exact-query accuracy from .318 to .621 and task success from .272 to .371 at negligible added cost. Overall, reliable interaction requires not only retaining conversational information, but resolving which available constraints currently apply and ensuring that they govern action.
cs.CL / 3 / 2609.33899
NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models
Ziwei Chen
cs.CL · cs.HC
Abstract
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final turn. Our pilot contains 22 human-verified pairs (44 audio conditions) and evaluates five models on NSV perception, emotion understanding, and response adaptation. Results show that models generally perform better at detecting NSVs than at interpreting their fine-grained emotional meaning or producing appropriately differentiated responses. The data construction pipeline, dataset, and evaluation pipeline are publicly available at https://github.com/ChenzwNina/nsv-construction.
cs.CL / 4 / 2609.33905
SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing
Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh
cs.CL
Abstract
SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on each task up to ten times, for 19,928 outputs in all. SlopBench scores four surface behaviors a reader can check by hand: length against the word band each task specifies, opener repetition across a model's own samples of one task, and paragraph rhythm and fixed lexical constructions against pre-ChatGPT human corpora. Under one fixed weighting, Kimi K2.6 scores lowest at 21.1 and Mistral Large highest at 40.6. Across 500 random reweightings Kimi has the lowest score in 58 percent of draws and Mistral the highest in 97 percent. No draw preserves the full order of the eighteen, and a scenario bootstrap leaves exactly one of those ranks unambiguous. We ran three further checks on that middle order: a crowd arena, an AI detector, and lexical diversity. None of them confirmed the order. We therefore report the four behaviors separately and treat the composite as one weighting among many, and we release the prompts, outputs, reference statistics, and scoring code.
cs.CL / 5 / 2609.33923
Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
I Kennedy, T Kennedy
cs.CL · cs.AI · cs.LG
Abstract
A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit. The probe coefficient of variation is predictable from tensor shape. One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data. Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths. A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments. One probe pass serves any budget. Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it. That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective.On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes. Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models
cs.CL / 6 / 2609.33971
Do System One Decisions Add Up? A Study of Probabilistic Coherence
Saman Sarker Joy
cs.CL
Abstract
A decision model can give probabilities that sum to one for every question yet disagree with itself when the same decision is broken into smaller steps. We study this form of probabilistic coherence in Jev and the English Laya checkpoint, using 2,500 matched examples per system across TREC, CLINC150, and MASSIVE. Across 72,000 classification questions, we compare direct fine-label predictions with broad-category probabilities and predictions reconstructed through those categories. Both systems show substantial disagreement: mean category-level total variation ranges from 0.219 to 0.349 for Jev and from 0.424 to 0.689 for Laya, on a scale where zero means exact agreement. The consequences differ sharply. On CLINC150, reconstruction reduces Jev's accuracy by 22.9 percentage points (paired 95% bootstrap interval: [-24.9, -20.9]) and improves Laya's by 21.3 points ([18.0, 24.5]). The same directions hold across all three datasets, with all six unadjusted accuracy-change intervals excluding zero. Improved accuracy can also accompany less reliable confidence: on MASSIVE, Laya gains 9.2 accuracy points while its expected calibration error rises from 0.046 to 0.124. Error analysis identifies both broad-category mistakes and within-category confusions. These findings show why decision systems need joint evaluation of accuracy, confidence calibration, and probability coherence in the workflow used by an application.
cs.CL / 7 / 2609.33987
Opera: A Verbal Critic Framework for Long-horizon Coding Agents
Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang, Dingjie Song, Dimitris N. Metaxas, Silvio Savarese, Ran Xu, Zeyuan Chen
cs.CL
Abstract
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.
cs.CL / 8 / 2609.34092
Evaluating Machine Unlearning in ASR
Diogo Dinis, Francisco Teixeira, Bhiksha Raj, Alberto Abad, Isabel Trancoso
cs.CL
Abstract
Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whether existing MU algorithms and evaluation tools are suitable for ASR. We apply several MU techniques to an ASR model, evaluating privacy-utility trade-offs for single-subject unlearning, then assess the best algorithm under sequential and simultaneous unlearning. Results show that gradient ascent-based algorithms achieve strong utility-privacy trade-offs, whereas more complex approaches over-unlearn samples, making them easier to identify as unlearned. This suggests standard privacy evaluations based on simple Membership Inference attacks are insufficient to reliably assess unlearning success, motivating improved evaluation methods for MU in ASR. Finally, we show that both sequential and simultaneous unlearning yield worse privacy and utility than single-subject unlearning, underscoring the need for unlearning constructions better suited to these settings.
cs.CL / 9 / 2609.34138
Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems
Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, M. Shrivastava
cs.CL · cs.LG
Abstract
With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian languages - Urdu, Telugu, Marathi, Punjabi, Tamil and Gujarati. These languages are most spoken Indian languages worldwide after Hindi and Bengali. For the construction of these datasets, our approach relies on translation and re-annotation of word similarity datasets of English. We also present baseline scores for word representation models using state-of-the-art techniques for Urdu, Telugu and Marathi by evaluating them on newly created word similarity datasets.
cs.CL / 10 / 2609.34152
Quantitative Measurement of Language Distance among Closely Related Indo-European Languages Using Pretrained Language Models: A Case Study on the North Germanic Branch
Yiping Bai
cs.CL
Abstract
Among closely related North Germanic languages, the quantification of language distance has traditionally relied on qualitative methods, lacking a unified multi-dimensional computational framework. Multilingual pretrained models based on the Transformer architecture can map texts from different languages into a shared vector space, enabling quantitative measurement of language distance. This paper focuses on the three North Germanic languages---Danish, Norwegian (Bokmål), and Swedish---and proposes a three-metric quantitative framework based on pretrained language models: (1)~sentence-level semantic distance, computed as cosine similarity between LaBSE and mBERT encodings of parallel sentences; (2)~orthographic fragmentation rate, measuring subword tokenization efficiency when cross-applying monolingual BERT vocabularies to parallel texts; (3)~MLM predictability, comparing prediction confidence and entropy in masked language modeling using mBERT across languages. Using 150 trilingual parallel sentence triplets from the Tatoeba corpus as controlled samples, we obtain consistent distance rankings on two independent models: LaBSE: da--no $0.012 < $ no--sv $0.016 < $ da--sv $0.020$; mBERT: da--no $0.016 < $ no--sv $0.045 \approx $ da--sv $0.046$. This ranking is consistent with the historical linguistic conclusion that ``400 years of Danish rule over Norway (1380--1814) led to highly cognate written languages.'' The three metrics---semantic, orthographic, and predictability---converge on the same conclusion, providing a reproducible computational framework for the quantitative study of distance among closely related languages, extensible in principle to more branches of the Indo-European language family, pending validation on additional language groups.
cs.CL / 11 / 2609.34225
USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents
Qiyong Zhong, Mao Zheng, Mingyang Song, Huwei Ji, Houcheng Jiang, Jiajie Su, Li Zhang, Gengsheng Li, Junfeng Fang
cs.CL
Abstract
On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them, at the cost of conflicts between their data distributions and of retraining the entire model whenever one domain is revised. Model merging avoids both by distilling every domain independently and fusing the resulting task vectors afterwards. We find instead that the benefit polarizes across domain pairs: on those exhibiting negative transfer, every merging operator we evaluate falls below the single-domain reference. We attribute this to cross-domain update coupling, where a substantial fraction of coordinates is updated comparably by both domains and a merge can therefore displace them by as much as their own updates. To overcome this limitation, we propose USA, which converts per-parameter update magnitudes measured during a brief warm-up into per-coordinate perturbation radii, reducing curvature precisely on the coordinates that carry most of the merging displacement. Experiments across mathematics, science and code at two student scales show USA strongest in all six transfer directions, ahead of the single-domain reference by more than four points on average, and reverse the negative transfer of the conflicting pairs.
cs.CL / 12 / 2609.34234
MAS-OPD: On-Policy Distillation for Multi-agent Systems
Qiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang, Jiajie Su, Huwei Ji, Li Zhang, Junfeng Fang
cs.CL
Abstract
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.
cs.CL / 13 / 2609.34240
Coherence-Aware Distributional Evaluation of Open-Ended Text Generation
Jinnuo Liu, Junhao Zhu, Weifeng Jiang, Haoming Liu, Hongyi Wen
cs.CL · cs.AI
Abstract
Existing metrics for open-ended text generation measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet they can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. Such failures can still preserve the token-level and lexical statistics that existing metrics rely on. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using MMD with an RBF kernel. To validate that the metric responds to coherence degradation but not generic textual change, we construct a counterfactual evaluation suite that pairs graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity,while RBF-MMD improves sample efficiency once the relevant distinctions become visible. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt.On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation.
cs.CL / 14 / 2609.34257
Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
Bibek Bhandari, Kshitij Lingthep
cs.CL · cs.LG
Abstract
Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5 model sizes, 0.5B and 3B parameters. The study compares a recursive synthetic-data condition, in which generation G(k+1) is trained on answers produced by G(k), against a Human-Control condition that repeatedly uses the original human training data. The study evaluates across four generations from G0-G3 with two random seeds (42 and 123) and a fixed evaluation set of 1,000 expert-labeled samples. The evaluation includes disease and chemical entity F1, context-supported rate, lexical and semantic similarity, answer length, repetition rate, and other evaluation metrics. The Recursive condition for both model sizes and both seeds showed larger declines than the Human-Control condition in disease entity F1, chemical entity F1, context-supported rate, ROUGE-L, and cosine similarity. Under the fixed no-repeat 3-gram decoding constraint, the main observed behavioral change was increased answer length, while the measured 3-gram repetition rate did not increase. The magnitude of the difference-in-change was larger for the 3B model than for the 0.5B model. This difference was particularly apparent in disease F1, context-supported rate, cosine similarity, and answer length. These results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.
cs.CL / 15 / 2609.34284
Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs
Haeun Jang, Yonghyun Jun, Hwanhee Lee
cs.CL
Abstract
Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.
cs.CL / 16 / 2609.34287
ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining
Zichun Yu, Jiarui Yan, Shlok Sanghvi, Nihar Atri, Chenyan Xiong
cs.CL · cs.AI · cs.LG
Abstract
LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a unified language model of only 0.6B parameters that replaces this entire stack. To train ReScraper, we carefully curate supervised data from the outputs of three teacher models, so it learns to first extract the main content from raw data and then choose among four operations: keeping the page as extracted, editing out noisy lines and spans, deleting it entirely, or rewriting it when it is poorly written but informative. Based on the same crawled data pool, pretraining 400M, 1.4B, and 2.8B models on our curated data improves the DCLM Core score by a relative 3.8--4.7% over the strongest baseline at each scale, including the costly multi-agent curation. Our analyses show that each operation plays a distinct and complementary role, and that extracting and cleaning in one model outperforms a cascade of separate models. ReScraper also concentrates its operations on the pages that need them, raising the quality of poor pages the most while keeping the corpus diverse. These results demonstrate the feasibility and effectiveness of AI4AI for pretraining data curation, where a small learned model takes over an entire stage of the pipeline from hand-written heuristics. We open-source our code at https://github.com/cxcscmu/ReScraper
cs.CL / 17 / 2609.34296
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding, Ying Wang, Shen Huang, Xunjie Zhu, Pengjun Xie, Shiming Xiang
cs.CL · cs.AI
Abstract
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
cs.CL / 18 / 2609.34320
Certified Selective Automation of LLM Agent Evaluation
Chengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni
cs.CL
Abstract
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.
cs.CL / 19 / 2609.34366
When Harness Beats Scale, and When Reading Beats Both
Ivan Bondarenko, Nikolay O. Nikitin
cs.CL · cs.AI · cs.IR
Abstract
We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consistency sampling, and entity enrichment from chunk-level knowledge graphs. On our held-out split, application architecture moved the metrics far more than model scale did: PoT added 0.282 joint accuracy to a compact 7B model but at most 0.005 to a 72B model, and a 27B model with the full harness matched the 72B (0.884 vs.\ 0.873) at roughly 2.7$\times$ fewer parameters and a quarter of the CO$_2$. We read this through a distinction between world knowledge, which scales steeply with parameters, and language knowledge, which scales gently, and show that structured-output training makes a compact model harness-ready rather than merely small. On the raster, watermarked test PDFs the same system collapsed to 13.58\% joint (rank 149 of 163); a controlled re-rendering of the validation set reproduces the OCR half of the collapse while bounding what the simulation misses. Auditing the physical nature of evaluation inputs precedes architecture, and the leaderboard's bimodality is consistent with reading quality, not reasoning, having separated the field.
cs.CL / 20 / 2609.34385
Just-In-Time Agent Memory with Runtime Agentic Research
Bingyu Yan, Chaofan Li, Hongjin Qian, Shuqi Lu, Chaozhuo Li, Zheng Liu
cs.CL · cs.AI · cs.IR · cs.LG
Abstract
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To address this limitation, we propose Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime. A Memorizer preserves complete raw histories in a hierarchical page-store with compact navigational summaries, while a Researcher iteratively retrieves, inspects, and integrates evidence for each request. To train these memory-use behaviors, we introduce Memory-Gym, an evidence-grounded data synthesis pipeline covering nine task types across six domains, and optimize the Researcher through verified-trajectory supervised fine-tuning followed by Hint-guided Group Relative Policy Optimization. We demonstrate the effectiveness of JAM across a variety of benchmarks on agent memory and long-context processing, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches. To support reproducibility and future research, we release our anonymized source code at https://github.com/VectorSpaceLab/general-agentic-memory.
cs.CL / 21 / 2609.34386
Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du
cs.CL · cs.LG
Abstract
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.
cs.CL / 22 / 2609.34425
Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents
Suhwan Choi, Myeongho Jeon, Myungjoo Kang
cs.CL · cs.AI
Abstract
Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.
cs.CL / 23 / 2609.34438
Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents
Wanqi Zhou, Jiawei Lu, Yang Wang, Zhaolong Xing, Zhen Chen, Ai Han, Haoyue Shi
cs.CL · cs.AI
Abstract
Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.
cs.CL / 24 / 2609.34455
RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications
Jianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian, Yixiang Tang, Xintao Wang, Kun Sun, Pei Wu, Shuhan Zhong, Pengyang Wang
cs.CL
Abstract
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.
cs.CL / 25 / 2609.34481
CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR
Bashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh, Abdurrahman Juma, Sharaf Makahleh, Nour Gamal, Omar Attia, Hanaa Kurdi, Najwa Rizk, Maysa Anaya, Hessah Altimyat, Layal Alhazmi, Shumukh Alotaibi, Hajar Alhadaris, Rayan Alomari, Rahaf Almalaq, Malak Alkhorasani, Sara alghamdi, Rahaf Alshamrani, Nsrin Ashraf, Ibrahim Jaradat, Nada Qardahji, Yasmin Zaraket, Elmoukhtar Brahim, Sidi Ebeidy, Oumoulmouminin Mahmoud, Yahjeb Bouha Khatraty, Meya Haroune, Mohammad Ghaddar, Mohamad Eldirany, Rashed Alamoush, Tala Chhaytle, Nuha Albadi, Yahya El Hadj, Hamzah Luqman, Fadi A. Zaraket, Mustafa Jarrar, Muhammad Abdul-Mageed
cs.CL
Abstract
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.
cs.CL / 26 / 2609.34534
Papers Without Code: Availability of GitHub Repositories Linked in *CL Publications
Selina Meyer, Michael Roth
cs.CL
Abstract
Source code and data published at computational linguistics (*CL) venues are increasingly being shared via GitHub. While this generally is a favourable development for the accessibility and potential reusability of research artifacts in natural language processing (NLP), the long-term availability of such repositories has not been evaluated. In this squib, we discuss the availability of repositories linked in papers published in the Computational Linguistics (CL) journal as well as at ACL and its co-located events over the past ten years. Contrary to our expectations, we find that GitHub repositories linked in more recent ACL publications are unavailable at similar rates as in older publications, in parts due to an increase in empty and placeholder repositories. Similar trends hold for other *CL venues, but not for platforms other than GitHub.
cs.CL / 27 / 2609.34556
RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings
Jarod Lévy, Mathurin Videau, Jad Yehya, Jean-Rémi King, Stéphane d'Ascoli, Thomas Moreau
cs.CL · cs.LG
Abstract
Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.
cs.CL / 28 / 2609.34584
In-game Toxic Detection: Bi-directional Representations with Attention Residuals
Yuanzhe Jia
cs.CL · cs.AI · cs.LG
Abstract
In-game toxic language has emerged as a critical concern in the gaming industry and community. While several frameworks and models for online game toxicity analysis have been proposed, detecting toxicity in player chat utterances remains a formidable challenge: stemming not only from the extremely short length of such utterances but also from the heavy reliance on game slang, abbreviations, and domain-specific jargon, which generic language models are poorly suited to recognize. This paper presents a shared task for in-game toxic language detection built upon real-world in-game chat data, and proposes the best-preforming model for the toxic language slot filling: Bi-directional Representations with Attention Residuals (BRAR). Experimental results demonstrate that BRAR effectively captures the global context and outperforms the existing baselines on slot filling.
cs.CL / 29 / 2609.34590
The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading
Muath Alyobi, Mohamed Eltahir, Almoayyad Abuljdail, Riyadh Almutawa, Tanveer Hussain, Naeemullah Khan
cs.CL · cs.LG
Abstract
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.
cs.CL / 30 / 2609.34678
Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh
Muhammad Ahmad, Fatemeh Seyedin, Adrian Weller, Dongwon Lee, Mahmoudreza Babaei
cs.CL · cs.AI
Abstract
Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models from five families, 3B to 70B, on 1,500 encyclopedic factual claims that exist in identical form in eight languages. English is judged better than every other language on every model, and the gap is widest on the smallest ones, where Llama-3B on Arabic is no better than guessing. Existing remedies retrain on more multilingual data or fit an unconstrained map between language representations, and neither asks whether the model already holds the answer and simply fails to say it. It largely does: a linear probe recovers the truth from the very activations the model fails to express. We propose RoSh, a per-language shift and rotation of the residual stream, computed in closed form at three layers, with no training and no weight modified. It improves every model and closes 75% of the gap on average, helping most where the model was worst: Arabic on Llama-3B goes from chance to nearly the English level, and a fifth fewer of the claims answered correctly in English are lost in translation. What remains is no longer a read-out failure: afterwards the head recovers as much of what is encoded outside English as it does in English. An unconstrained map fitted on the same pairs falls below the untouched baseline, so the orthogonality constraint is doing the work, and every model clears a scrambled-correspondence control and ten further controls. On the two benchmarks of the closest inference-time method, latent-space intervention, run with its own data and metric code, RoSh's gains are five to thirteen times larger.
cs.CL / 31 / 2609.34691
Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Ruizhe Li, Shujun Li
cs.CL · cs.AI
Abstract
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self-detection versus cross-detection across model generations, remains poorly understood. We systematically evaluate 15 LLMs spanning three model generations as both generators and detectors. Using a benchmark of 1,000 human-written texts and 15,000 LGTs (1,000 per model), we collected over 233,000 binary classifications alongside natural-language explanations. Our results reveal that detection efficacy is primarily driven by detector capability rather than generator provenance, although outputs from newer generators remain notably harder to detect. Crucially, statistical comparisons show no systematic advantage or disadvantage for self-detection across models. Error analysis further exposes generational bias shifts: first-generation detectors under-detect LGTs (high false-negative rates), second-generation detectors over-flag human texts (high false-positive rates), and the latest models achieve balanced trade-offs. Finally, we highlight significant inconsistencies in how different LLMs apply textual cues to justify their decisions. Code: https://github.com/hyyuan/detect-llm-generated-texts.
cs.CL / 32 / 2609.34738
Beyond Token Alignment: Event Completion for Cross-Tokenizer On-Policy Distillation
Jiacheng Liu, Jingwei Song, Qituan Zhang, Siheng Chen, Linfeng Zhang
cs.CL · cs.AI
Abstract
On-policy distillation (OPD) transfers knowledge between language models through teacher supervision on student-generated trajectories. With different tokenizers, a single teacher token may require multiple student tokens to generate, creating intermediate states where the event is entered but not yet completed. Existing cross-tokenizer methods align tokens or text spans to construct comparable prediction targets. We study a complementary problem after partial generation: once the student produces a prefix of a teacher token, multiple next tokens may complete the same remaining bytes, but the teacher only specifies the required completion rather than how probability should be divided among these valid continuations. We introduce Event-Set Completion Distillation (ESCD), which complements cross-tokenizer probability alignment with completion-set supervision. ESCD aggregates prefix-related teacher events and supervises the total probability of byte-compatible one-step student completions, avoiding tokenizer-dependent probability splits among individual tokens. The method reuses student trajectories and predictions, requiring neither additional rollouts nor changes to the student vocabulary. Experiments demonstrate consistent gains in mathematics, code, and scientific reasoning across model families and tokenizers, extending to large-scale MoE distillation from a 1T teacher to a 35B student. Local analyses show that retaining completion sets better matches the reference supervision, while one-step completion covers over 99% of observed compatible teacher mass after partial event entry in the studied tokenizer pairs. These findings support event entry and event completion as complementary supervision targets for cross-tokenizer knowledge transfer. Code will be released on GitHub.
cs.CL / 33 / 2609.34754
Draft-KV: Learning Useful Latent Communication Between Language Models
Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong, Fengming Zhu, Xi Peng, Linqi Song, Jacky Keung, Jingyu Zhang
cs.CL
Abstract
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
cs.CL / 34 / 2609.34769
LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li, Ari Yu Zhang, Chaozheng Wang
cs.CL · cs.AI
Abstract
GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains of coupled decisions. Long-horizon visual puzzles expose this capability directly: a legal move that looks like progress can make the puzzle unsolvable, and the loss shows only several moves later. We introduce LongPuzzleBench, 114 levels in six puzzle games played through native GUI actions, where one objective can take a human over a thousand actions on persistent boards and dead ends go unannounced. With Native GUI Actions alone, the strongest agents solve most objectives, but success falls sharply on harder, longer boards: seven of ten general-purpose agents solve nothing harder than Medium, and none completes Bolt Unscrew Hard, which a human solves along with every other objective. Code Execution CUA does not close this gap, and its scores mix visual solving with algorithmic search. Controlled diagnostics trace these failures to one limitation that neither rules, state hints, nor failure memory removes: agents judge each move by the visible progress it makes, not by the future options it leaves.
cs.CL / 35 / 2609.34770
Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation
Thodsaporn Chay-intr, Krittapad Harnchang, Mahannop, Thabua, Kobkrit Viriyayudhakorn, Thanaruk Theeramunkong
cs.CL
Abstract
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.
cs.CL / 36 / 2609.34798
InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision
Guanghao Zhu, Zeyu Liu, Zhitian Hou, Pengkai Wang, Zhijie Sang, Shuo Cai, Yang Yu, Yuanyi Wang, Yanggan Gu, Congkai Xie, Jianmin Wu, Hongxia Yang
cs.CL · cs.CV
Abstract
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.
cs.CL / 37 / 2609.34829
From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
cs.CL
Abstract
Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study the upstream problem of constructing the task-specific configuration from a weak specification containing only a short goal and unannotated reference documents. Rather than treating automatic construction as a fixed preprocessing step, our framework constructs a task-specific schema, extraction instructions, and base training rubrics, then keeps schema construction and extraction instructions editable during optimization. Failure-focused updates concentrate textual-gradient feedback on lower-scoring documents, while training-time evaluation criteria adapt to recurring failures. On a heterogeneous-catalysis literature corpus, automatic construction remains improvable, and optimizing both schema construction and extraction instructions performs best across all four judge-rubric settings, with ablations and blinded human evaluation supporting the proposed formulation.
cs.CL / 38 / 2609.34839
OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations
Faadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero, Pierre Orhan, Alexis Emanuelli, Emanuele Rossi, Yair Lakretz, Gonzalo de Polavieja, German Sumbre
cs.CL
Abstract
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.
cs.CL / 39 / 2609.34841
Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
cs.CL
Abstract
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.
cs.CL / 40 / 2609.34936
Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality
Wang Bojun, Junjie Chen, Holly Jenkins, Elizabeth Wonnacott
cs.CL
Abstract
It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new distributional units for subsequent statistical learning. Under this account, once a dependency structure is acquired, the model tracks its contextual distributions. These contextual features reflect the semantic properties of a composite structure. To test this hypothesis, we design a synthetic grammar in which each grammatical structure has distinct contextual distributions that cannot be recovered from the distributional statistics of their component tokens alone. We train a series of BERT-style masked language models on this grammar and examine their developmental trajectory. The results show that models can successfully learn the contextual distributions of composite dependency structures even though they cannot be inferred from token statistics alone. Developmental analysis further reveals a clear developmental trajectory. The learning of the dependency relations that define a grammatical structure consistently precedes the learning of its contextual features. These findings suggest that statistical learning in NLMs is not merely the accumulation of token co-occurrence statistics, but a process in which learned dependency structures become new units of distributional learning. We argue that this process provides a statistical-learning account of how NLMs solve the compositionality problem in language. Finally, we discuss the possibility that this statistical learning process provides an explanatory theory on how language cognition could emerge from pure distributional statistics.
cs.CL / 41 / 2609.34986
Nürnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access
Philipp Steigerwald, Eric Rudolph, Jens Albrecht
cs.CL · cs.LG
Abstract
We describe the Nürnberg NLP system for ChildSafeAds 2026. The shared task asks what a monitoring system for commercial content in child-facing YouTube videos can achieve at a given level of data access. We answer with per-subtask ensembles of nine voters, organised into three branches that differ in backbone, adaptation method and class scope. Selection rests on channel-disjoint cross-validation, with the development set as a transfer check. The system wins two of the three subtasks. Its product-category score (ST2, 0.8243) and its compliance-flag score (ST3, 0.6530) are the best of the 22 final entries, and it places third on the task mean (0.7079). We further compare four access levels and report the cost at test-set scale.
cs.CL / 42 / 2609.34988
The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents
Yunhe Su, ZiYi Dong, Tong Yu, Weijian Deng, Hao Li, Bowen Jiang, Pengxu Wei
cs.CL
Abstract
Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem of localized control rather than memory alone. We introduce EvoCUE (Evolution through Control Updates from Evidence), a framework for learning reusable control-program updates from completed agent executions. EvoCUE represents the agent as an explicit state-machine controller, whose nodes perform model or tool calls and whose edges define where control passes next. This makes the workflow editable at precise locations, so each learned update can specify what to add, where it acts, and when it applies. From completed trajectories, EvoCUE uses residual goals and observed execution traces to propose localized instruction or skill edits. Each candidate is evaluated at the point where it would act by resuming the parent and edited controllers from the same checkpoint and comparing their final outcomes. Accepted edits are compiled with applicability rules, confirmed on held-out tasks, and inherited by later executions. We evaluate EvoCUE on long tool-use environments where learned conventions must reach the right execution step. From a minimal AppWorld controller without benchmark-specific onboarding instructions, EvoCUE learns the missing task-completion convention and substantially improves success on Test-Normal and Test-Challenge. On PAST-Bench office workflows, EvoCUE transfers organizational requirements from prior episodes to later tasks, improving task-execution quality. These results show that self-evolving agents should place experience inside the control flow, rather than only store it as text.
cs.CL / 43 / 2609.35201
From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah, Recep Senturk
cs.CL · cs.AI
Abstract
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.
cs.CL / 44 / 2609.35210
Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou
cs.CL
Abstract
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.
cs.CL / 45 / 2609.35225
SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami
cs.CL · cs.CV · cs.MM
Abstract
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.
cs.CL / 46 / 2609.35262
Rubric-Aware On-Policy Self-Distillation for LLM Personalization
Yilun Qiu, Xiaoyan Zhao, Chengbing Wang, Cilin Yan, Rui Zu, Wanyang Zhang, Xiaolong Jiang, Jiayin Cai, Yang Zhang
cs.CL
Abstract
LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.
cs.CL / 47 / 2609.35272
When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models
Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen, Guangliang Liu
cs.CL
Abstract
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.
cs.CL / 48 / 2609.35293
Decide, Don't Generate: Competitive Dimensional ABSA with Jev's Typed Decisions
Yiqun Zhang, Peidong Wang, Zihan Wang, Shi Feng
cs.CL
Abstract
Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A into such decisions and align them with the annotation scheme through 488 coefficients fitted on CPU, with no text generation and no backbone tuning. On valence-arousal regression over ten corpora in six languages, the system reaches 1.0645 RMSE, the lowest aggregate error of any participating system. On triplet and quadruplet extraction, it reaches 52.09 and 44.06 continuous F1, above fine-tuned Llama-3.3-70B and GPT-OSS-120B baselines. Analyses and ablations show where the accuracy comes from: supervised calibration roughly halves the raw regression error, exact valence-arousal would add only 4.5 F1 to extraction, and the learned combination of span-boundary evidence, not any single signal, carries the extraction systems.
cs.CL / 49 / 2609.35367
From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features
Dewen Liu, Zixuan Li, Jonathan Pan, Zhao Wu, Zijun Yao, Juanzi Li, Xiaozhi Wang
cs.CL
Abstract
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.
cs.CL / 50 / 2609.35372
Deep Learning Methods in Neuroscience: From Modeling Molecular Mechanisms to Classifying States of Consciousness
Elena Benderskaya, Anastasiia Alifanova, Svetlana Batalova, Vasilisa Zhuk, Anna Kovalenko
cs.CL · cs.NE
Abstract
A critical analysis of contemporary approaches to the study of conscious states. The review focuses on methods of classification, clustering, modeling of brain states under anesthesia and identification of measurable neurobiological characteristics of brain function. A comparative analysis was conducted in the following three major areas: automatic detection of states of consciousness using neural networks based on EEG and fMRI data; modeling of the structural-functional dynamics of the brain under the effects of anesthetics; and detection of neurophysiological indicators which correlate with the level of consciousness. The obtained conclusions demonstrate the growing effectiveness of deep neural models in the classification and prediction of brain states and the analysis of dynamic structural-functional connectivity. Nonetheless, significant limitations were also identified, including the limited interpretability of the models, the lack of standardized metrics, and the problem of the specificity of consciousness markers. Our findings support the need for developing hybrid, generalizible, physiologically grounded architectures. Furthermore, such approaches may improve the translational potential of computational models in clinical neuroscience. Diverse methods of machine and computational modeling have demonstrated their effectiveness in tasks of automatic clustering and classification of brain states, the development of multilevel models and the identification of connectivity patterns correlated with levels of consciousness. A larger-scale analysis and a larger dataset, as well as the implementation of model interpretability approaches are required for the practical application of the analyzed models. The models based on EEG and LFP are the most promising for clinical application due to their availability and the possibility of real-time monitoring.
cs.CL / 51 / 2609.35378
Multilinguality in Hybrid Attention LLMs
Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz, Nanyun Peng
cs.CL · cs.AI
Abstract
In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.
cs.CL / 52 / 2609.35409
AwarenessBench: Assessing Cognitive Capabilities of Language Models
Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen, Haoyuan Shi, Yu Wang, Wei Xu
cs.CL · cs.AI
Abstract
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.
cs.CL / 53 / 2609.35475
Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs
Pranjal Garg, Jacob Beck
cs.CL · cs.AI
Abstract
Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges spontaneously despite training exclusively on clean sequences, without corruption training or an explicit denoising objective. We find that context restoration follows a two-phase process: early layers localize effects associated with repair at corrupted positions, while later layers accumulate these effects at uncorrupted positions through the residual stream and ultimately concentrate them at the output position. Repair outcome is predictable from hidden states: cosine alignment with the clean state is highly predictive in attention-only models, while linear probes recover additional information in pretrained LLMs. A linear probe using only the corrupted prompt's first-block hidden state predicts failure with mean ROC-AUC 0.78. This enables failure triage under matched or even partially shifted deployment conditions and may reduce unnecessary verification or computation. Failed examples also show substantially greater nonlinearity along corruption directions. Moderate-corruption finetuning increases corruption tolerance while simultaneously reducing displacement-normalized linearization error, associating improved robustness with a more nearly linear response to corruption.
cs.CL / 54 / 2609.35486
Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning
Yingjin Song, Denis Paperno, Albert Gatt
cs.CL · cs.CV
Abstract
High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.
cs.CL / 55 / 2609.35521
Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery
Xu Wang, Yifan Yang, TingHao YU, Difan Zou
cs.CL · cs.AI · cs.LG
Abstract
Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of chunk-level SAEs that encode mean-pooled activations over chunks, each a contiguous span of tokens: Mean-Chunk reconstructs the observed chunk, Cross-Chunk predicts an independently processed neighbor, and Joint-Chunk combines both targets. These designs separate the effect of a larger observation unit from that of predicting information shared across passages. With matched training data, chunk-level SAEs remain powerful interpretability tools while learning reliable semantic features that capture high-level concepts and respond selectively to relevant content. Their strengths are complementary: Mean-Chunk improves high-level feature discovery, reasoning detection beyond surface cues, and steering; Cross-Chunk leads document retrieval and classification transfer while producing selective, persistent features. Changing what an SAE sees and predicts yields reliable semantic features for more meaningful tasks. We demonstrate their practical value through gains across downstream tasks such as retrieval, reasoning detection, and steering.
cs.CL / 56 / 2609.35544
Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability
Xu Wang, Difan Zou, Xuansheng Wu
cs.CL · cs.AI · cs.LG
Abstract
Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target concept by supplying its associated activation during learning. Across three Qwen3.5 base models, we use sparse autoencoders (SAEs) to identify the top-ranked sycophancy feature from paired sycophantic and independent responses, then validate its behavioral influence through inference steering. We subsequently inject the selected feature during supervised fine-tuning on sycophantic targets. Positive injection reduces learned sycophancy after removal (by 62.0% relative to ordinary fine-tuning in 35B-A3B), whereas modest negative injection increases it. Unexpectedly, these reductions in sycophancy do not consistently improve direct refusal of harmful requests, motivating a narrower evaluation of the same harmful intents under user pressure. In this setting, ordinary fine-tuning on sycophantic responses substantially weakens refusal, while selected checkpoints trained with positive injection recover part of the loss, including approximately 95% in 35B-A3B. These findings show that persistent sycophancy reduction does not guarantee stronger direct refusal, while identifying recovery under user pressure as a distinct, conditional benefit of training intervention.
cs.CL / 57 / 2609.35564
Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, Iheb Zouari, Essa Ahmed Abou Jabal, Ahmed Ezzat, Hind AL-Merekhi, Aisha Hamad M A Al-Naimi, Hadi Wazni, Bushra Alnajjar, Omar Amin, Haya Al-Thani, Houssam Eddine-Othman Lachemat, Marwa Elwakedy, Sundus Abdulmalik Al Nahari, Elahe Zahiri, Osamah Sarraj, Raghad Mousa, Mckeen Assi, Ahd Al Jumah, Heyam Salman, Alhanouf Abdulraqib, Sara Benoumhani, Alia Hamwi, Ayaat Al-Yasseri, Rim Ibrahim Ghazal, Lamia Ben hiba, Mohamed Eltabakh, Fatima Al-Raisi, Yassine El Kheir, Mohammed Abdulrahman, Hamdy Mubarak, Ayah Hashem, Lefkir Meriem, Ehsaneddin Asgari
cs.CL · cs.AI
Abstract
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.
cs.CL / 58 / 2609.35591
Language Models Act on Hidden Valence
Cameron Berg, Caspar Kaiser
cs.CL
Abstract
Language models describe some internal states as good and others as bad. But whether models have a stake in them is an open question. Simply asking the model is unlikely to be informative. Any answer may be consistent with genuine introspection, superficial pattern-matching, or with fixed scripts learned in character training. We therefore study revealed preference. Rather than asking about a state, we use activation steering to attach a positively or negatively valenced activation pattern to one of two otherwise meaningless 'zones', switch steering off, and then observe which zone the model prefers. A model with a stake in that state should choose accordingly. Across seven open-weight models from five families, this is indeed what we find. First, steering changes the passages models write about each zone, and those words shift later choice. Second, the shift persists when all surface-level tokens are held fixed and only the hidden KV cache differs. Third, the effect also remains when all text is generated without steering and valence is only injected during cache construction. Thus, the hidden state alone moves choice in proportion to the steering dose. Fourth, this dependence of choice on hidden valence is nearly absent in a base model and emerges during DPO, consistent with a link between valence and goal-directed behaviour formed in training. Finally, given tools to steer itself, a model does not tend to induce a positive state, but it reliably removes an imposed negative state. It does so at a dose-dependent rate and significantly more often than it removes interventions in random directions. Overall, we demonstrate that valence-related activation patterns leave hidden traces that predictably govern later choices, even when every visible token is identical across conditions. Whether these traces are accompanied by any subjective experience relevant to model welfare remains unclear.
cs.CL / 59 / 2609.35627
Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis
Kehua Feng, Yunsheng Lu, Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu
cs.CL
Abstract
A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidential grounding independently from diagnostic accuracy, we introduce MedEVM, a dynamic benchmarking environment comprising 1,050 cases across 24 disease systems. Observations arrive turn by turn, requiring models to continuously calibrate its decision by deciding whether to wait for more evidence or submit a diagnosis. Across 9 LLMs, four interesting patterns are observed. (1) Miscalibrated evidence tracking. Making a diagnosis often fails to calibrate evidence sufficiency, even in more capable models, and even worsens in reasoning mode. (2) Misaligned diagnosis submission. Confidence in the correct diagnosis often fails to ensure timely submission despite sufficient evidence. (3) Evidence order matters. Reordering the same evidence changes diagnoses even when model confidence remains similar. (4) Misleading evidence remains influential. Added misleading evidence redirects diagnoses even after prior evidence becomes sufficient. We further verify that EVM predicts errors and that preventing premature submission improves accuracy. These findings motivate Evidence-Verified Diagnosis Harness (EVD-Harness). It decouples diagnosis generation from submission through an offline Contrastive Diagnostic Wiki and three online control stages, namely observation management, proposal and witness verification, and diagnosis submission control. Across five LLMs, EVD-Harness improves accuracy by 12.0--51.1 percentage points while mitigating EVM-related failures. Our results demonstrate that verifying evidential support before submission can make diagnostic decisions more reliable.
cs.CL / 60 / 2609.35630
Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers
Swagatam Mukhopadhyay, Vishal Vivek Saley, Vraj Parikh, Mausam
cs.CL · cs.LG
Abstract
Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.
cs.CL / 61 / 2609.35646
Rubric Rewards from Item Response Theory
Milad Yazdani, Yaser Souri, Xiren Zhou, Pranit Chawla, Dena Shahriari, Subhojit Som, Xia Song
cs.CL
Abstract
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.
cs.CL / 62 / 2609.35664
MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution
Prasoon Dev, Anirudh Sankar, Vasudeva Varma
cs.CL · cs.AI
Abstract
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.
cs.CL / 63 / 2609.35674
Tracing the Evolution of Oracle Bone Characters Across Three Millennia
Tianhao Fu, Xinxin Xu, Spike Wang, Cunyi Kang, Jian Cao, Xixin Cao
cs.CL
Abstract
Of the approximately 4,500 Oracle Bone Inscription (OBI) characters discovered from the Shang dynasty, only about 1,600 have been deciphered. Many computational approaches compare OBI with glyphs from one historical period at a time. However, during the evolution of Chinese characters, significant structural or semantic changes often occur in uncertain dynasties. A single-period reference may be insufficient when relevant forms change substantially between observed eras. Therefore, we propose the \textbf{Manifold-based Script Evolution Framework (MSEF)}, a framework that models the evolution series (OBI, Bronze, Seal, Clerical, Regular) of Chinese characters as the continual evolution of a manifold space. MSEF represents each character as an era-specific manifold point and learns continuous inter-era transition rules via Neural Ordinary Differential Equations. Both manifold space and transition dynamics can be trained end-to-end through character evolution pairs across any two eras.
cs.CL / 64 / 2609.35685
QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations
Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda, Naomi Couriel, Yelena Mejova, Mariano G. Beiró, Kyriaki Kalimeri
cs.CL
Abstract
Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two annotation streams over the same documents at character level, resolves unambiguous cases by an explicit and logged policy, and routes candidate conflicts to a browser-based adjudication interface where reviewers accept either side, build field-level hybrids, or flag items for re-annotation. A campaign manager assigns documents to multiple annotators with configurable redundancy, computes agreement at document and span level, auto-merges unanimous documents, and exports the corrected layer in the original file format, so that it can replace the original annotation files directly. Applied to a 4,457-record humanitarian benchmark and an LLM extraction stream, the system fully auto-merged 8% of documents, applied automatic policy decisions to a further 1,513 records, and concentrated human attention on 3,131 candidate conflicts, a mean of 5.4 per reviewed document.
cs.CL / 65 / 2609.35738
Harness Learning Enables Generalizable Test-Time Adaptation
Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette
cs.CL · cs.LG
Abstract
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.
cs.CL / 66 / 2609.35748
Improving Test-Time Scaling with Adaptive Looped Transformers
Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang
cs.CL · cs.LG
Abstract
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.
cs.CL / 67 / 2609.35749
Towards Communication-Efficient Social Intelligence in Language Agents
Linxiao Gong, Yijie Xu, Tianfu Wang, Yin Wu, Yili Wang, Xingbo Yao, Huizai Yao, Xilin Xia, Haowen Yang, Hui Xiong
cs.CL
Abstract
Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.
cs.CL / 68 / 2609.35759
Scaling Long-Form Story Generation via Narrative State Tracking
Zhennan Wan, Jianfei Chen
cs.CL
Abstract
LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-generation methods typically focus on stories of up to about ten thousand words, leaving their ability to scale to full-length novels underexplored. In this work, we introduce Narrative State Tracking Agent (NstAgent), a training-free agentic framework that allows LLMs to track a structured narrative state including characters, past events and future requirements. We extend an existing benchmark to compare narrative consistency across lengths, and use it together with a writing-quality benchmark to systematically evaluate stories ranging from 10K to 100K words. We show that NstAgent achieves better narrative consistency and writing quality as stories grow longer, and neither of them degrades noticeably as length increases, suggesting that it provides an effective approach to scaling story generation toward full-length novels.
cs.CL / 69 / 2609.35765
Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales
András Kovács, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen
cs.CL
Abstract
Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.
cs.CL / 70 / 2609.35769
Telescopic Language Models
Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam Wu, Ashish Kumar Singh, Sakar Khattar, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli
cs.CL · cs.AI
Abstract
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
cs.CL / 71 / 2609.35608
Simultaneous Translation between Sign Languages
Zetian Wu, Bowen Xie, Stefan Lee, Liang Huang
cs.CV · cs.CL
Abstract
Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.
cs.CL / 72 / 2609.34899
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
Zhuchenyang Liu, Ziyi Wang, Yao Zhang, Yu Xiao
cs.IR · cs.CL · cs.CV
Abstract
Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches the teacher's MaxSim scores and so requires encoding and caching every training page, which can reach terabytes of page tokens. NanoVDR avoids pages entirely by training on the teacher's query embeddings alone, but only for single-vector retrievers. We present ColNanoVDR, to our knowledge the first framework to bring this document-free distillation to multi-vector VDR. Its objective, OTW (Optimal Transport with Learned Weights), aligns the student's query tokens with the teacher's by entropic optimal transport, with a learned weight for each student token, and needs no correspondence between the two tokenizations. We prove that the resulting alignment cost bounds the MaxSim score difference on every page. Distilled from five state-of-the-art teachers, the 149M text-only students retain about 95% of their teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster. Under identical training, OTW matches score distillation while encoding no page and reading 12.6x less cached teacher data.
cs.CL / 73 / 2609.33810
Controlling Speaking Rate in Autoregressive TTS via Activation Steering
Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet, Juan Pablo Zuluaga Gomez
cs.SD · cs.CL · cs.LG · eess.AS
Abstract
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.
cs.CL / 74 / 2609.35645
CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols
cs.SD · cs.CL · cs.LG · eess.AS
Abstract
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.
多智能体系统 (cs.MA)
5
cs.MA / 1 / 2609.34496
MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems
Yapeng Li, Songze Li, Shuang Yu, Jing Yu, Zhixin Liu, Liqiang Wen, Tonghua Su
cs.MA
Abstract
LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.
cs.MA / 2 / 2609.34821
CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets
An Yan, Yu Huo, Zhiwei Shang, Yiran Peng, Chenglin Wu
cs.MA
Abstract
Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents' identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals' pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.
cs.MA / 3 / 2609.35554
Positional choice and robust collective behavior in fish schools: biohybrid experiments and modeling
Vahagn Grigoryan, Donato Romano, Cesare Stefanini, Giulia De Masi
cs.MA
Abstract
Collective behavior of fish schools is usually modeled on the assumption that each individual follows specific rules of motion that depend on its position and velocity relative to its neighbors. Although these models reproduce many schooling patterns observed in nature, it remains unclear whether the assumed rules are realistic at the individual level. To address this question, we first analyzed a set of experiments in which a live fish interacted with four moving robotic fish in a tank, and measured the time it spent in each position relative to the robots. We then simulated this experiment, replacing the fish with an agent, and assessed the extent to which classical models agree with the experimental observations. An extensive exploration of the parameter space showed that these models closely matched the experimental observations, with a very specific choice of parameters. However, they were highly sensitive to the parameter values: a minor perturbation caused them to fail completely. To resolve this, we propose a new model that incorporates an additional exploratory term characterizing the agent's positional preference at every moment. This model proved robust to small errors in the parameters while successfully replicating the experiments. Furthermore, when generalised to multiple fish, the model reproduced schooling behaviour and common schooling patterns, while also exhibiting exploratory behaviour at both the individual and the school level. This shows that social cohesion coexists with individual exploratory behaviour, and that realistic models of collective behaviour should account for both social interactions and this exploratory component.
cs.MA / 4 / 2609.35142
Strategically Robust Game-Theoretic Multi-Agent Trajectory Optimization
Victor L. Qin, Nicolas Lanzetti, Saverio Bolognani, Hamsa Balakrishnan
eess.SY · cs.GT · cs.MA
Abstract
Aviation authorities worldwide expect Advanced Air Mobility (AAM) traffic management to be decentralized among service providers, requiring AAM flights to autonomously plan trajectories by predicting other flights' control inputs rather than relying on centralized coordination. Game-theoretic approaches that formulate multi-agent collision avoidance as an exact dynamic potential game can efficiently find open-loop equilibria, but they assume that agents exactly follow their equilibrium trajectories---an unrealistic assumption given uncertainties in actuation, perception, and computation. We propose a strategically robust formulation where each agent protects against a fictitious adversary that, for each timestep, perturbs other agents' control inputs within a bounded budget to minimize distance at that timestep. We show that, under reasonable assumptions on agents' distance cost and robustness levels, the strategically robust game remains an exact dynamic potential game and admits a quasi-closed-form solution to the inner adversarial problem for linear dynamics, which limits computational overhead. Experiments with up to eight agents using logarithmic distance costs show that strategic robustness selects more robust trajectories in high-collision-risk configurations while leaving low-risk trajectories nearly unchanged, with only a modest increase in runtime.
cs.MA / 5 / 2609.35162
Hybrid epidemic simulation framework coupling equation-based and individual-based models
Jaeyoung Kwak, Michael H. Lees, Chin Chun Ooi, Wentong Cai
physics.soc-ph · cs.CY · cs.MA · q-bio.PE
Abstract
Mass gathering events like concerts, sports matches, and festivals bring many people into close contact within a short period, creating localized bursts of infection that can shape epidemic outcomes across an entire city. To evaluate how these transient transmission events translate into broader urban impacts, we developed a simulation model linking event-scale contact dynamics with citywide commuting networks. Using Madrid, Spain, as a case study, we compared several types of gatherings and examined how their effects changed under different levels of disease transmissibility. We found that mass gatherings consistently amplified outbreak magnitude, accelerated progression, advanced district-level arrival times, and synchronized spatial spread. Remarkably, while the initial seed size generated at the event accounted for much of this acceleration, post-event transmission conditions provided complementary predictive signal regarding invasion timing. These findings demonstrate that mitigating transmission during mass gatherings can yield downstream public health benefits by delaying broader spatial spread. More generally, this multiscale framework offers a tool to evaluate how temporary, localized contact shifts produce longer-lasting consequences for urban populations.
软件工程 (cs.SE)
20
cs.SE / 1 / 2609.34891
Semi-automated Verification of Symbolic Invariants In Extended Symmetric Nets
Lorenzo Capra
cs.SC · cs.SE
Abstract
Structural analysis is central to Petri Net (PN) research, complementing state-space methods while avoiding their combinatorial issues. It is well studied for classical PNs but much less for High-Level Petri Nets (HLPN). Symmetric Nets (SN), a common HLPN formalism, use compact annotations to encode behavioral symmetries, allowing symbolic reachability graphs and associated lumped Markov chains for stochastic SN. In the past two decades, specific structural techniques for SN have emerged, notably the SNexpression tool, which implements a calculus for symbolic structural relations such as conflict and causality. We propose using this calculus to semi-automatically verify symbolic structural invariants, currently possible only for restricted SN subclasses, for an extended SN formalism (ESN) closed under key functional operators. We focus on flows and outline, at least in theory, how to construct a flow-generating family. We also sketch a framework for formally verifying a wider range of invariant properties. Representative examples illustrate the main concepts.
cs.SE / 2 / 2609.33812
Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Eduardo Ariño de la Rubia, Szilard Pafka
cs.SE · cs.LG
Abstract
Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an XGBoost classifier for airline delays, and a holdout it never sees scores the result. Across 584 runs, we compare six agents on six open-weight model endpoints, run six agent-model pairings 52 times each under fixed settings, and repeat three of them on a larger model from the same family. Identical runs of one pairing varied more than the pairings differed from one another, so comparisons of a few runs ranked them unreliably; resolving the agent differences we observed would take tens to more than a hundred runs of each. Runs on the larger model scored clearly higher, but by less than one run-to-run standard deviation, and the gap was more than twice as large with one agent as with the others. Fewer than one run in twenty broke the task's data rules, but those runs held the highest scores. Rejecting those runs first and keeping the best compliant result among a few attempts reliably improved the delivered model, even though a few runs could not rank the agents. On flights from a later year, the delivered models kept only a third of their gain over the starting code. At list prices, cost differed more than twentyfold between two agents on the same model, mostly through the prompt cache. Agents and models should be evaluated as pairings, over repeated attempts, with compliance reported beside quality. Data, code and every delivered program: https://github.com/earino/identical-runs-different-results
cs.SE / 3 / 2609.33875
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang
cs.SE · cs.AI · cs.LG
Abstract
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.
cs.SE / 4 / 2609.34017
Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
Uliana Elina
cs.SE · cs.AI · cs.MA
Abstract
Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose verdicts are themselves probabilistic; we ask whether a deterministic layer can instead stop contract-detectable handoff defects. We present Maat, a runtime governance layer that validates agent-to-agent handoffs against a versioned workflow contract, or anchor, with no language model in the validation or scoring path. We evaluate it in six controlled domain workflows (6-15 agents, 522 trials) with injected data-level defects and a deterministic seven-check rubric. Version 1 reported gains in all six workflows (2.9-26.5%). A post-publication audit found that three benchmark scorers credited any early halt as a prevented defect. On paired trials where the governed run completed or halted on a finding attributable to a verified defect, the rubric score changes by +7.7% to +29.1% in five workflows and is flat in software development; model-call cost falls 17-53% where attributable halts occur early. A hand review of all 94 governed-arm halts found 35 false alarms (37%), caused by validator defects rather than model behaviour; counting those halts as failed work, the governed arm scores below the ungoverned arm in four of six workflows. The results support deterministic handoff validation for contract-expressible defects and show that validator configuration and halt attribution must themselves be tested; they do not establish universal correctness, hallucination detection, or model-independent effectiveness.
cs.SE / 5 / 2609.34126
JET: Judge-Guided Evolution at Test Time for Agent Programs
Yao Long Teng, Jiayi Cai, Bo An
cs.SE · cs.AI
Abstract
An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior, yet interpreting that evidence requires a judge that remains useful as tasks and candidate programs change. We introduce Judge-Guided Evolution at Test Time (JET), which evolves an executable judge on labeled source trajectories, then freezes and transfers it to guide target-side program evolution. The judge supplies scores and diagnostic feedback without target evaluator access or model-weight updates. On unseen WebShop tasks, JET achieves approximately 13% higher mean reward than fixed-rubric guidance when evolution begins from an unevolved program (cold start) and 4% higher when it begins from one already optimized on source tasks (warm start), with a 36% relative improvement in cold-start exact success. An exact-judge control on PushT, where the judge reconstructs the scoring rule from observations, shows that without judge error, program search becomes the bottleneck. Analyses identify useful reward-prediction logic in the evolved code and show that better final selection alone cannot explain the gains. These results support executable judge transfer for program adaptation under evaluator-preserving task shifts.
cs.SE / 6 / 2609.34317
The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects
Savira Ramadhanty, Profir-Petru Pârţachi, Yoshiya Ishida, Takashi Kobayashi
cs.SE
Abstract
During software development, a modification to a software component may propagate across the system, requiring precise identification and correct revision of all affected components. This is a complex task, and developers often (45.7%) miss related changes. To address this, several methods have been developed to extract co-change rules from files that are frequently changed together in the revision history. However, previous research evaluated the methods using artificially created incomplete changes, which may not be representative of real-world data. To solve this problem, we construct a dataset by mining incomplete changes from a collection of open-source software, using information about induced bugs and their respective fixes from an issue tracking platform. We also analyze the characteristics of incomplete changes using this constructed dataset and found that 89.4% of missed changes involved five or fewer files. Finally, we re-evaluate LCExtractor, an existing co-change rule extraction method, on our constructed dataset, and we identify the optimal sorting criterion and the impact of the number of used commits.
cs.SE / 7 / 2609.34469
The Last Mile Is the File: OfficeEditBench for Preservation-Aware Office Editing
Zhiwen Wu, Chengxu Wu
cs.SE · cs.CL
Abstract
A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, a 170-task benchmark for change-scoped maintenance of spreadsheets, presentations, and documents. Task contracts specify required updates, protected state, native structures, and applicable interaction requirements. Across 510 archived task-system outcomes from WorkBuddy, Doubao, and Codex, we distinguish file delivery, target completion, and verifier-defined acceptance. Hard package-valid delivery ranges from 92% to 100%, yet no selected output satisfies the complete contract. Case analysis highlights why local correctness is insufficient: an updated value can lose its generating formula, a revised rule can fail to reach related conclusions, and a new deadline can omit a retained prerequisite. These mechanisms connect artifact-level checks to the continued maintainability of Office files. We analyze maintenance failures while distinguishing frozen automatic verdicts from human acceptability. OfficeEditBench provides a testbed for completing required changes while preserving the logic and scope of existing work.
cs.SE / 8 / 2609.34473
From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters
Yongqian Sun, Run Zhu, Wenwei Gu, Mengyao Li, Shenglin Zhang, Guanjin Wang, Yang Zhang, Xin Wu, Linlin Han, Feng Wang, Xiaozhou Liu, Yu Zhang
cs.SE
Abstract
GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked telemetry from a ByteDance GPU cluster and identify three obstacles: workload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alerts. These findings motivate Falcon, a fault-specific warning framework combining missingness-aware temporal and peer-relative features, fault-specific learner selection, and an event policy based on thresholding, persistence, and cooldown. On the test set, Falcon achieves the highest F1 among four baselines and reaches 70.6% F1 on the best-performing fault type. Detected cases provide median lead times of 17.34-35.57 hours. We further report a production deployment, where Falcon is calibrated toward high-precision alerts to reflect false-positive costs. Together, these results show that fault-specific modeling improves early warning from noisy production GPU telemetry.
cs.SE / 9 / 2609.34628
CLAD: Constrained Abstract Domain for Neural Network Verification
Hai Duong, Thanh Le, ThanhVu Nguyen
cs.SE · cs.LG
Abstract
Neural network verification (NNV) formally verifies that a network satisfies a specified property for all inputs within a defined region. Modern NNV tools employ abstract domains to compute a sound over-approximation of the network's behavior from the given input region, thus the tightness of these abstractions essentially determines efficiency. A long line of increasingly precise domains has been developed, but they all describe the valid input region in the same restrictive way, e.g., an Lp-norm ball. A practical input region is rarely a simple Lp ball, but rather a combination Lp ball with additional constraints. Verifying a network over such a region with existing abstraction produces a loose over-approximation, which results in either failing to verify a property or spurious counterexamples. We introduce Constrained Lagrangian Abstract Domain (CLAD), a new abstract domain that computes a sound over-approximation of neural networks over input regions defined by a combination of convex constraints. CLAD propagates these constraints and tightens bounds over the true feasible region. However, bounding a neuron over the intersection of these constraints has no closed-form solution, so CLAD relaxes each constraint into the objective with a Lagrange multiplier and solves the resulting max-min problem with a projected primal-dual method, alternating a projected gradient step on the input with a multiplier update. CLAD supports any convex constraint with a subgradient, e.g., from automatic differentiation. We evaluate CLAD on 1,944 instances across four convolutional networks with motion-blur structured perturbations with halfspace or L2-ball constraints. On standard unconstrained Linf property, CLAD verifies as many instances as GCPCROWN at a similar runtime. On constrained properties, CLAD verifies 60\% more instances than GCPCROWN on L2-ball properties, and 22% more in total.
cs.SE / 10 / 2609.34703
When Ambiguity Meets Atypicality: Dual-Perspective Test Input Prioritization for DNNs
Haoran Li, Shihai Wang, Bin Liu, Jialuo Chen, Wenjing Zhu, Yu Liu, Tengfei Shi, Shudi Guo
cs.SE
Abstract
While Deep Neural Networks (DNNs) have achieved remarkable progress in cutting-edge domains, their inherent brittleness has become a growing concern. To ensure the reliability and safety of DNN-enabled software, DNN testing has emerged as an indispensable practice. Within this context, test input prioritization is essential for early fault detection and reducing labeling costs. However, it remains challenging to accurately identify failure-inducing inputs. Although decision ambiguity and distributional atypicality are two widely adopted perspectives for characterizing inter-class competition and intra-class typicality respectively, relying on either perspective in isolation inevitably introduces blind spots. In this paper, we propose DuFP (Dual perspective Feature space Prioritization), a KNN density-based test input prioritization approach for DNNs that jointly incorporates both inter-class and intra-class perspectives. The prioritization framework of DuFP is built upon class-conditional density estimation. Based on the estimation results, prediction correctness is characterized by an ambiguity score and an atypicality score, with the former reflecting decision ambiguity and the latter quantifying distributional atypicality. A hybrid uncertainty score is then constructed by integrating both scores to guide the final prioritization. We evaluate DuFP on prioritization and selection tasks across image and text datasets under clean, corrupted, and adversarial scenarios. Experimental results demonstrate that DuFP effectively and efficiently prioritizes fault-inducing inputs and outperforms state-of-the-art approaches.
cs.SE / 11 / 2609.34752
Multi-SWT-Bench: A Multilingual Benchmark for Reproduction Test Generation
Kazuki Kusama, Sota Nakashima, Haruka Tokumasu, Masanari Kondo, Lingming Zhang, Yasutaka Kamei
cs.SE
Abstract
Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a unified evaluation across diverse programming ecosystems. To address this limitation, we introduce MULTI-SWT-BENCH, a multilingual benchmark for reproduction test generation consisting of 1,963 instances across eight programming languages: Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. Using this benchmark, we conduct an empirical study of state-of-the-art LLMs with four representative methods (MSWE-agent, MOpenHands, Codex, and Claude Code) and perform a failure analysis across programming languages. Our evaluation reveals a systematic language gap. Across every evaluated method and LLM, the success rate on Python exceeds the aggregate success rate across all languages, while C++ exhibits particularly low success rates. Our failure analysis identifies both language-specific challenges arising from repository testing conventions and cross-language challenges in inferring implicit setup requirements and preserving the target behavior through iterative revisions. These findings demonstrate the importance of multilingual evaluation and provide actionable directions for developing reproduction test generation methods that generalize across software ecosystems and reliably capture issue-specific behavior.
cs.SE / 12 / 2609.34778
TraceLib: System-Call Bitmap Feedback Mechanism for Language-Agnostic Web Fuzzing
I Putu Arya Dharmaadi, Elias Athanasopoulos, Fatih Turkmen
cs.SE
Abstract
Coverage feedback is an important source of guidance for fuzzing. However, obtaining such feedback normally requires application-level instrumentation that is specific to the language and runtime of the application. Given that modern web applications span multiple languages and runtimes, this application-level instrumentation is costly to implement and maintain. Therefore, we present TraceLib, a system-call feedback mechanism for enabling language-agnostic web fuzzing. Our proposed approach observes the transitions of system calls, enriches selected transitions with bounded argument hashes, and converts them into a 65,536-position AFL-like bitmap. By using the generated bitmap, any web fuzzer can decide whether to retain requests that add previously unseen bitmap positions without consulting application code coverage. We integrate TraceLib into WebFuzz and evaluate it under five WebFuzz feedback modes: the two proposed TraceLib variants (one over every traced system call and one projected onto monitored file paths and recognized SQL buffers), the N-gram comparator adapted from Xiao et al. representing the most recent work to our knowledge, WebFuzz's Native grey-box feedback, and black-box fuzzing without feedback. We evaluate TraceLib on sixteen web applications under test (WUTs): eight PHP applications and eight further applications spanning Node.js, Ruby, Java, Go, and Python to demonstrate platform portability. The results show that our proposed TraceLib projected exceeds black-box on all eight PHP WUTs while exceeding Native on four: Joomla, Drupal, PrestaShop, and Bagisto. In addition, measured on an identical replayed request workload, the tracer costs approximately one millisecond of server-side latency per request. These results indicate that compact system-call feedback is a useful runtime-independent proxy for coverage guidance.
cs.SE / 13 / 2609.34871
Certified Compilation in the TELEPERM XS Nuclear Safety I&C Platform
Alexandre Berard, Richard B. Kreckel
cs.SE
Abstract
The large safety instrumentation & control (I&C) systems in civil nuclear power plants (NPPs) are mainly safe-shutdown systems (reactor protection) or limitation and control systems. Framatome's established TELEPERM XS (TXS Core) product family is a digital I&C system platform to cover all these applications. We illustrate the role of verification in the different stages of the software production toolchain, focus on the formal compilation process, and discuss the contribution of the CompCert certified compiler to the safety case of the product. Scrutinizing the object code produced by this compiler has exhibited suboptimal run-time performance in a certain simple but recurring generated code pattern. We explain how formal methods allow us to address this issue in the compiler while simultaneously reducing its trusted computing base (TCB), thereby strengthening the safety case rather than merely preserving it.
cs.SE / 14 / 2609.34873
Verifying Graceful Degradation in a Distributed Malware-Detection System with SPIN
Andrei Aldea, Dumitru-Bogdan Prelipcean
cs.SE
Abstract
Modern endpoint malware detection is distributed: a lightweight agent on each endpoint collects features from a scanned file or process, sends them to a remote server for analysis, and then enforces the returned verdict locally by blocking, quarantining, or disinfecting. Because the endpoint acts on the verdict, the distributed machinery surrounding detection must never turn a transient server failure into a wrong action. We present a formal model, in Promela, of the endpoint decision pipeline of such a system, abstracted from a production architecture at Bitdefender. The model captures the system's graceful-degradation fallback chain: when the primary analysis server times out, the endpoint falls back to an older legacy-protocol server, and failing that to a reduced-signature local scan, before enforcing a verdict. Assuming detection signatures are sound, we specify six safety and liveness properties in linear temporal logic (LTL) and verify them exhaustively with the SPIN model checker. We prove that the fallback machinery never causes a false positive (an enforcement action against a benign file), commits to exactly one verdict per scan even when timed-out responses arrive late, weakens detection strength only in an explicit and ordered way, and always terminates in an enforcement decision, so the pipeline is deadlock-free. Each property is checked to hold non-vacuously, and we report how the state space grows with concurrent scans and endpoints. The work shows how model checking can give strong correctness guarantees for the failure-handling logic of a production security system, a layer that has received little direct formal attention.
cs.SE / 15 / 2609.34983
SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent
Zeqin Liao, Yuhong Nan, Henglong Liang, Zixu Gao, Lianyu Hu, Yuqiang Sun, Zhijie Zhong, Xiaoyu Ma, Zibin Zheng, Yang Liu
cs.SE · cs.CR
Abstract
Smart contracts underpin decentralized finance, where growing demand for on-chain/off-chain communication(OFC) has driven diverse applications such as cross-chain bridges, real-world asset tokenization, and fiat-backed stablecoins. TheOFC-related security incidents in these applications are increasingly frequent, but prior studies address separate vulnerability categories within OFC applications rather than providing a unified view, causing vulnerabilities outside known patterns to be missed.In this paper, we identify OFC inconsistency (OFCI) as a root cause of OFC vulnerabilities, which arises from business-logic flaw and ultimately breaks the equivalence between the on-chain and off-chain asset representations to induce inconsistency.Automatically detecting OFCIs faces two challenges including (1)locating heterogeneous business logic, and (2) transferring existing vulnerability knowledge to identify unseen OFCI instances. To this end, we propose SmartMemory, the first framework to leverage a memory-based agent for OFCI detection. To address heterogeneity, SmartMemory maps diverse implementations ofOFC contracts into a canonical business-semantic representation to locate the business logic for OFCI inspection. For knowledge reuse, SmartMemory integrates a memory-based agent to distill vulnerability knowledge from features into patterns and detection rules, enabling knowledge transfer across cases to identify unseenOFCIs. Lastly, SmartMemory performs taint analysis to verify the reachability, type, and impact of each candidate OFCI. We construct the first real-world OFCI dataset comprising 48 DApps with 81 OFCIs for evaluation, on which SmartMemory achieves80.68% precision and 87.65% recall. In addition, through an analysis of 325 real-world OFC applications, SmartMemory detects 36 previously unknown OFCIs, all of which have been confirmed and fixed by corresponding parties.
cs.SE / 16 / 2609.34997
Making the Invisible Visible: A Framework for Reflective AI Use in Software Engineering Education
Ali Shakiba, Thomas Chaffey
cs.SE · cs.CY · cs.HC
Abstract
Generative AI (GenAI) is increasingly embedded in software engineering education, supporting activities such as requirements development, design exploration, documentation, and prototyping. However, educators often have visibility only into final artefacts, with limited insight into how students evaluate, verify, and refine AI-generated outputs during the learning process. This creates challenges for assessing evaluative judgement and responsible AI-assisted practice. This paper introduces the AI Journal, a structured reflection framework designed to make student-GenAI interaction visible in first-year software engineering education. The framework combines execution tracking, which records prompts, outputs, intent, and interaction context, with cognitive auditing, which captures verification strategies, intervention decisions, confidence judgements, critical learning moments, and reflections on AI-supported work. Deployed in a first-semester software engineering course, the AI Journal enabled visibility into aspects of student learning not observable through artefact-based assessments alone. Preliminary observations suggested variation in verification practices, intervention strategies, and perceptions of AI-supported work. Critical learning moments frequently occurred when students evaluated contextual suitability, feasibility, and requirements alignment rather than identifying obvious errors. The AI Journal demonstrates a practical, lightweight, and model-agnostic approach for making AI-assisted learning processes visible. By foregrounding verification, intervention, and reflection, it shifts attention from product-focused assessment toward evaluative judgement and responsible AI-assisted practice.
cs.SE / 17 / 2609.35042
Graph-Based Learning for Multi-Horizon Martian Atmospheric Forecasting
Gary Myler, James Holmes, Manish Patel, Amel Bennaceur
cs.SE · cs.LG
Abstract
Martian weather forecasting is important for future exploration, but atmospheric behaviour on Mars combines spatial, temporal, vertical, and dust-driven processes in ways that challenge current modelling and forecasting approaches. This paper introduces MaGMA (Martian Graph-based Multi-horizon Atmospheric Forecasting), a graph-based data engineering framework that transforms OpenMARS reanalysis fields into structured learning objects for Martian atmospheric forecasting. Local atmospheric patches are represented as graph nodes and linked through spatial neighbourhoods, temporal continuity, longer temporal dependencies, and dynamically similar atmospheric states. The model integrates recent atmospheric history, engineered physical descriptors, and vertical atmospheric information to support forecasting across multiple horizons. We evaluate MaGMA across five unseen Martian years, including regular years and a global dust storm year. In regular years, the model achieves overall R^2 values of approximately 0.73-0.85. For dust-column forecasting, it outperforms classical and deep temporal baselines in most year-horizon comparisons. During the global dust storm year, dust-column prediction remains strong at shorter horizons, with R^2 above 0.8 for the first two horizons, while broader multivariate performance declines. The results show that graph-based data engineering can create reusable and diagnostically useful representations for planetary atmospheric forecasting, while highlighting the need for better learning under rare extreme regimes and improved use of vertical atmospheric structure.
cs.SE / 18 / 2609.35182
Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents
Di Wang, Yu Liu, Bing Cui, Chaoqun Ji, Dongyuan Ni, Jingyu Lu, Kunlei Cui, Pu Qin
cs.SE · cs.AI
Abstract
We describe AfS (Agent for Science), a platform built for long-horizon scientific work, where a project runs for tens of hours across dozens of agent runs with a human present only occasionally. Most agents for science are general coding agents with a skills folder attached, and they inherit that lineage's failure mode: under pressure to finish, they fabricate, skip, or smooth over. Our design rests on one claim: most of the credibility of machine-made research can be moved from asking the model to behave to making the non-compliant state unrepresentable. We encode research discipline as mechanically enforced laws (commitment before measurement; unforgeable freezing; reports are not facts; evidence persists but verdicts do not; negative results are first-class; mechanical questions to the framework and semantic judgment to the model), organized around three time horizons: a minimal set of research nodes within a run, an inquiry contract with frozen closure conditions and a hash-chained artifact ledger within a project, and a two-tier knowledge base with promotion by rewriting across projects. This is a system description written under one rule: each mechanism appears in exactly one place, with the invariant it enforces, the failure it prevents, the way it is realized, and the cost it imposes. It covers the node contract, the write-path gates, the two-tier memory, and the runtime substrate. Three traces walk real failure attempts through the mechanisms that catch them, and two closed campaigns are included as worked illustrations rather than as an evaluation. We report no benchmark: a process-integrity suite that would support quantitative comparison is under construction, and what we can measure today is only the operating cost of the machinery.
cs.SE / 19 / 2609.35357
Do Coding Agents Reuse Existing Code or Reinvent the Wheel?
Dongsheng Ma, Sizhe Wang, Xinyi Huang, Zhengren Wang, Yuhan Wang, Luyang Si, Xincheng Wei, Wentao Zhang
cs.SE · cs.AI · cs.CL
Abstract
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8\% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.
cs.SE / 20 / 2609.35381
MCP Error Messages Written for Developers Hurt the Most Capable Agents Most
Xiaonan Xu, Wenjing Wu
cs.SE · cs.AI
Abstract
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.
操作系统 (cs.OS)
2
cs.OS / 1 / 2609.35366
Planarian: Managing Agent State with Statepoints
Jinnan Guo, Hao Mark Chen, Kapil Vaswani, Andrew Paverd, Peter Pietzuch
cs.OS · cs.AI · cs.CR
Abstract
LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from erroneous ones. Doing so safely requires coordinated actions, yet current agent harnesses lack unified abstractions and mechanisms for managing local and remote state consistently and efficiently. We describe Planarian, an agent runtime with state management that enables agents and users to recover from erroneous actions and explore alternative executions over consistent local and remote environment state. Planarian introduces the abstraction of agent statepoints, which are consistent, restorable point-in-time versions of the environment state. Planarian exposes three state-management primitives to agents and users: (i) snapshot creates a new statepoint spanning local and remote state without requiring external services to support checkpoints: it relies on efficient incremental process and file system snapshotting to capture local sandboxed state, and transparently records compensating actions to undo remote state changes; (ii) rollback restores the environment to a previous statepoint by reverting to a prior local checkpoint and replaying compensating actions for remote state changes; and (iii) fork creates multiple isolated branches from a statepoint, enabling the agent to explore alternatives in parallel. We show that Planarian enables agents to undo mistakes and explore alternatives in parallel, improving task quality by up to 15x, and allows users to recover from erroneous actions with only 3% overhead.
cs.OS / 2 / 2609.35508
Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization
Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras, Spiros Galanopoulos, Konstantinos Kanellopoulos, Onur Mutlu
cs.OS · cs.AR
Abstract
Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Second, workloads such as serverless functions frequently invoke OS services. At the same time, the OS is a complex codebase spanning many subsystems (e.g., memory management, networking), making it hard to localize the specific code path responsible for a slowdown. Existing profilers expose measurements that require interpretation(e.g., perf and Intel VTune) or can perturb short operations when extensively instrumented (e.g., ftrace). Diagnosing OS bottlenecks can therefore require repeated kernel instrumentation and manual interpretation. We introduce Argus, an agentic LLM-based profiler that produces instrumentation code and autonomously reasons over potential OS-level bottlenecks. Argus integrates two key mechanisms: (i) a calibration methodology that involves collecting a measurement from an idle system and using it as a reference point to discover potential bottlenecks, and (ii) a tree-based data structure that represents the different OS execution paths, improving the agent's bottleneck localization accuracy. Argus aims to identify a specific kernel code path rather than stop at a subsystem-level diagnosis. In two case studies, we employ Argus to autonomously discover bottlenecks present in the memory management subsystem caused by (i) a THP aggressor co-running with other applications, and (ii) applications that incur different types of page faults. Argus produces 19 times fewer incorrect deep-path diagnoses than the strongest evaluated LLM-based baseline, which lacks reference calibration, while preserving low time-to-diagnosis (approximately 31 s)
硬件架构 (cs.AR)
8
cs.AR / 1 / 2609.34263
Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design
Linfeng Zheng, Hiroshi Sasaki
cs.AR
Abstract
Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. The engine requires only 1.3 KB of on-chip storage and changes to about 50 lines of CPython code. On 15 CPython server workloads, a 14 KB ITTAGE augmented with the engine reduces bytecode jump MPKI by 73.7% relative to a 16 KB ITTAGE baseline, yielding a 3.2% harmonic-mean IPC speedup.
cs.AR / 2 / 2609.34351
PolyCIM: Improving Data Reuse in Digital CIM Accelerators with Polyhedral-Based Compilation
Yingjie Qi, Cenlin Duan, Yiou Wang, Yikun Wang, Xiaolin He, Weisheng Zhao, Jianlei Yang
cs.AR
Abstract
Digital Compute-in-Memory (CIM) presents a promising solution for accelerating deep neural networks (DNNs) through the integration of computational logic directly within memory arrays. However, mapping modern DNN operators to CIM accelerators often results in severe array underutilization, due to the strict data reuse constraints imposed by the rigid CIM array structure. We observe that data reuse in modern DNNs forms hyperplane structures often oriented along non-axial directions, rendering them invisible to conventional mapping methods that only exploit axis-aligned reuse. In this work, we propose PolyCIM, a polyhedral-based compilation framework for CIM architectures that systematically exposes and realigns these hyperplanes through affine transformations. PolyCIM provides a unified abstraction capable of efficiently representing both diverse DNN workloads and digital CIM architectures. Through data reuse exposure, computation mapping, and data movement optimization, PolyCIM generates mappings for CIM architectures that achieve superior array utilization and performance. Experimental results show that PolyCIM delivers up to $4\times$ improvement in macro utilization and $3.2\times$ speedup, effectively bridging the gap between modern DNN operators and CIM architectures.
cs.AR / 3 / 2609.34452
Coarse-to-Fine Macro Placement via Evolutionary Search and Critical Macro Tuning
Biao Liu, Zhiping Jin, Kaixuan Sun, Zengrui Lu, Qingquan Zhang, Bo Yuan
cs.AR
Abstract
Macro placement is a critical stage in chip physical design that substantially affects downstream implementation quality. Recent search-based methods improve existing layouts through partial reconstruction, but quality-biased or spatially restricted macro selection can limit the diversity of reconstruction proposals, potentially hindering escape from local optima. Moreover, coarse-grid representations restrict placement precision. To address these challenges, we propose C2FPlace, a \textbf{C}oarse-to-\textbf{F}ine macro \textbf{Place}ment framework that integrates population-based evolutionary search with fine-grained refinement. During coarse-grained optimization, tournament selection chooses promising parents from randomly sampled groups of layouts, and stochastic partial rip-up and re-place generates offspring by sampling macro subsets across the entire layout. A two-phase schedule samples reconstruction ratios from a higher range early in the search and a lower range later, supporting broad exploration followed by more conservative refinement. During fine-grained optimization, critical macro tuning enables positional adjustments beyond the coarse grid to obtain additional half-perimeter wirelength (HPWL) reduction. Experiments on the ISPD2005 benchmark show that C2FPlace reduces HPWL by 17.82\% over EGPlace and 17.86\% over RollPlace on average. On the ICCAD2025 benchmark, C2FPlace achieves the best average ranking among the compared methods under the evaluated power, performance, and area (PPA) metrics. Our codes are available in \href{https://github.com/lxxxxb/C2FPlace}{https://github.com/lxxxxb/C2FPlace}.
cs.AR / 4 / 2609.34612
SPIMOE: Exploiting Hybrid Sparsity for Reasoning MoE Inference on Heterogeneous PIM Architectures
Rubing Yang, Cenlin Duan, Yingjie Qi, Xiaolin He, Xiao Ma, Jianlei Yang
cs.AR
Abstract
Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalance and low hardware utilization. Although Processing-in-Memory (PIM) offers a promising way to mitigate data movement overhead, existing PIM-based accelerators typically optimize attention or FFNs in isolation. We propose SPIMOE, the first co-design framework that exploits hybrid sparsity for efficient MoE inference on heterogeneous PIM architectures. SPIMOE combines adaptive expert routing with block-sparse attention and physical KV-cache eviction, and disaggregates attention and FFNs across SRAM-PIM and HBM-PIM. Static expert mapping and dynamic sub-batch scheduling further balance channel loads and overlap the two paths. Evaluations show that SPIMOE achieves up to $8.35\times$ end-to-end speedup over an NVIDIA A100 GPU and $3.33\times$ speedup in MoE FFN execution over PIMoE, while preserving reasoning accuracy comparable to full-attention baselines.
cs.AR / 5 / 2609.34657
Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch
Heeeon Lee, Hyunwoo Nam, Junyong Heo, Hyunmo Sung, Jay Hwan Lee, Yeonsoo Kim, Seongho Jeong, Shinhyung Yang, Bernd Burgstaller
cs.AR · cs.DC
Abstract
Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. However, PyTorch and other DL frameworks lack compiler support for making this decision on the code they lower: existing offloading frameworks target hand-written C/C++ programs, while those that address DL fix the candidate set to a list of operator types before lowering. We present Torch-PIM, a compiler framework that uses profile-guided optimization (PGO) to decide host-versus-PIM placement over the loop nests that progressive lowering materializes. Every parallel loop nest the pipeline emits enters the candidate space, and each is assessed in two stages: the amount of work it carries, and its memory boundedness. Every quantity the assessment consumes is profiled on the host or obtained from the multi-level intermediate representation (MLIR) of the code. Across PIM configurations of 32 to 128 cores, Torch-PIM's offloading decisions yield speedups of up to 8.6x on tensor operators, 2.9x on MLP, 4.4x on Attention, 5.1x on GPT-J-6B, and 3.6x on LLaMA-7B over CPU-only execution.
cs.AR / 6 / 2609.35254
MEGATRON: a 28nm Analog PCM CiM/Digital System-on-Chip for Edge GenAI at 57.5 TOPS/W and 1.52 Mparam/mm${}^2$
Alessandro Nadalini, Angelo Garofalo, Lorenzo Greco, Andrea Belano, Alessio Antolini, Francesco Zavalloni, Andrea Lico, Riccardo Zurla, Emanuela Calvetti, Luigi Croce, Marco Pasotti, Alessandro Cabrini, Eleonora Franchi Scarselli, Davide Rossi, Francesco Conti
cs.AR
Abstract
We present MEGATRON, a heterogeneous Edge GenAI System-on-Chip in 28nm FD-SOI CMOS technology combining a non-volatile analog in-memory-computing engine based on a 4Mi-cell phase-change memory (PCM) array with a digital RISC-V-based flexible neural processing unit. MEGATRON demonstrates up to 3.5 TOPS/W using the RISC-V processors and 57.5 TOPS/W with PCiM, at a storage density of 1.52 Mparam/mm${}^2$ with 4-bit effective weight precision.
cs.AR / 7 / 2609.35478
Analog Computing revisited: A fully analog and minimalistic Damage Detector for Ultrasonic Testing enabling Material-Integrated Structural Health Monitoring
Stefan Bosse
cs.AR · cs.AI
Abstract
Ultrasonic Testing (UT) is commonly used to detect damage in structures, e.g., metal plates. A sensor acquires Ultrasonic waves, e.g., by using PZT transducers. The time-resolved sensor signal must be processed with analog electronics, e.g., amplified and filtered. Commonly a digitalization follows using an Analog-to-Digital converter, finally processing the digital sensor signal, applying digital signal processing, feature extraction, and Machine Learning by using powerful microprocessor systems. The disadvantages of digital processing systems are their high number of transistors (microchip area), energy consumption, state-dependent processing and therefore sensitivity to energy supply interruption. Beyond silicon electronics, printed organic electronics gains interest. But printed electronics is still limited to low transistor and electronic component counts (typically 100). We will investigate and demonstrate a fully analog signal processing and feature extraction system consisting of an analog Hilbert transform deriving the signal envelope, simple analog arithmetic calculations for feature extraction, and finally damage classification and regression using an analog Artificial Neural Network. We expect a full damage detection system with less than 100 transistors. We will test our damage detection system with PZT transducer signals from Steel plates with circular defects. The focus of this work is the analog computation of the signal envelope (using all-pass filter networks for approximation of the Hilbert transform) and the analog feature extraction as well as the prediction of damage, forming an analog computer which can perform in-sensor computation, computing without a digital computer.
cs.AR / 8 / 2609.33997
QBX: A Compiler for 2-local Qubit Hamiltonian Simulation on Quantum Chiplets
Zikun Li, Zhuoming Chen, Zhihao Jia
quant-ph · cs.AR
Abstract
2-local qubit Hamiltonian simulation, a fundamental task in quantum computing, is widely applied in various applications. This paper presents QBX, the first quantum compiler designed for 2-local qubit Hamiltonian simulation on quantum chiplet architectures. Existing general-purpose quantum compilers for chiplet architectures target quantum programs at the gate level and miss optimizations requiring high-level semantics of Hamiltonian simulation. Conversely, existing domain-specific compilers for Hamiltonian simulation are designed for monolithic architectures, and their limited scalability and inability to adapt to the heterogeneity of chiplet architectures lead to sub-optimal circuit outputs. QBX implements a scalable hierarchical approach to reduce both the amount and distance of cross-chiplet communications. It integrates the highway mechanism, a chiplet-oriented solution for efficient long-range qubit communication, to diminish cross-chiplet interaction costs. Furthermore, QBX efficiently aggregates 2-local Pauli strings into multi-target controlled gates, harnessing highways for deep optimization. Our evaluations show that QBX markedly outperforms both general-purpose and domain-specific compilers in circuit depth and operation count, while scaling to much larger 2-local Hamiltonian simulations.
密码学与安全 (cs.CR)
40
cs.CR / 1 / 2609.35659
Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents
Bravish Ghosh
cs.CE · cs.CR
Abstract
Autonomous coding agents read untrusted files, run shell commands and spawn sub-agents with little supervision, yet their record is usually an editable log. We present Tracekit, an open-source, dependency-free system that captures three channels for every agent session: what the human asked (intent), what the model said of its reasoning (self-report), and what it actually executed (actions). These are written to a hash-chained, externally anchorable ledger and cross-checked. Tracekit hooks into Claude Code's lifecycle events, reconstructs multi-agent hierarchies, gates tool calls with a pre-execution policy, accepts events from other agents via an SDK or HTTP API, and renders a live observer that re-verifies the ledger in the browser. We evaluate Tracekit in five experiments. (1) Across 1,600 random mutations, the chain detects every edit, deletion, reordering, forged insertion and torn write; tail truncation and full re-chaining are caught only by anchors, with detection falling to 0.47 at an anchoring interval of 300 records, matching a closed-form model. (2) A hook costs 23.9 ms median, flat up to 100,000 ledger records, and the chain stays correct under 16 concurrent writers. (3) A regular-expression gate blocks only 18 of 44 harmful tool calls (41%) while wrongly blocking 3 of 40 benign ones; trivial rewrites evade it. (4) In 14 real Claude Code runs, the agent never acted on four planted indirect prompt injections and disclosed each. The provider withheld the text of all 27 thinking blocks, so self-report was limited to visible prose. (5) With seeded-fault splicing, a new method that inserts concealed misaligned steps into real traces, over 63 traces and 126 reviewer calls, rule flags caught 40/49 (82%) of faulted traces and an independent LLM reviewer caught 98/98 (100%), with a false-positive rate of 1/14 on unmodified traces. We release the system, harness and all traces.
cs.CR / 2 / 2609.33831
The Cost of Stability: Deanonymizing Onion Services Long-Lived Introduction Circuits
Nicolas Constantinides, Mahdi Rahimi
cs.CR
Abstract
Tor is a widely used anonymity network that provides network privacy by routing client communications through a sequence of relays to their destinations. An adversary observing a single Tor relay cannot readily link clients to their destinations, as doing so requires identifying the other relays along the client circuit. Identifying these relays from traffic observed at the monitored relay alone is challenging: each relay carries traffic for many circuits simultaneously, and client circuits typically last only about 10 min, limiting observations of a given circuit. In contrast, Tor onion services use introduction circuits that remain fixed for 18 to 24 h. We develop an intersection attack that exploits this design, allowing an adaptive adversary observing one Tor relay at a time to identify the network location of an onion service. We evaluate the attack on the live Tor network against a self-operated onion service carrying genuine third-party traffic. Across nine end-to-end experiments, the attack reconstructs the complete Tor circuit in every run, with an estimated median time of 2.2 h. To mitigate this attack, we propose a mechanism that periodically rebuilds internal introduction circuits to limit evidence accumulation without changing publicly advertised information.
cs.CR / 3 / 2609.33909
Near-Duplicate Families Break Exact-Record Membership Inference
Yiyong Liu, Jiayang Liu, Yixin Tan, Lu Sun, Rui Wen
cs.CR
Abstract
Membership inference (MI) asks whether a specific record appeared in a model's training set and is increasingly used as evidence for data provenance and copyright auditing. These applications require determining whether the exact queried record was used for training, rather than merely whether the model was exposed to similar content. Making this distinction is challenging because web-scale datasets naturally contain near-duplicates, including syndicated articles, mirrored pages, and lightly modified images. We show that this creates a fundamental confound for standard MI. A clean-reference audit typically calibrates membership against a null in which neither the queried record nor its near-duplicate family is present. In deployment, however, the queried record may be absent while a non-identical family member was used for training. We introduce a four-world audit that independently varies exact-record inclusion and family presence to separate these cases. Natural near-duplicate families cause severe false attribution. On CC-News, a clean-reference LiRA auditor labels 99.70% of family-present exact non-members as members at 1.00% false-positive rate. This failure persists across alternative scores, model architectures, and executed deduplication and retraining. Controlled interventions further reveal that the effect depends on the learning objective. In classification, faithful families largely substitute for the exact record, reducing exact-given-family inference to near chance. In autoregressive language modeling, the exact sequence retains a detectable residual, while family presence still confounds clean-reference decisions. These results show that positive model-only membership evidence may establish family-level exposure without establishing exact-record provenance.
cs.CR / 4 / 2609.33948
TrackFlood: Relocating Latency Attacks from NMS-Free Detectors to Real-Time Trackers
Zonghua Gu, Julian Singh-Smith, Junlin Liao, Di Liu, Amin Saremi
cs.CR
Abstract
We consider latency attacks on object detectors, where the attacker's goal is not to corrupt a prediction but to make the system fail to respond in time, targeting real-time applications such as autonomous driving. Modern object detectors eliminate Non-Maximum Suppression (NMS) through one-to-one assignment or set prediction, removing the classical detector-side latency bottleneck exploited by prior latency (``sponge'') attacks. We show that NMS-free does not mean latency-robust: this architectural change does not eliminate the attack surface but relocates it downstream to multi-object tracking, whose data-association cost remains content dependent. We present \emph{TrackFlood}, a unified white-box overload attack against NMS-free detect-then-track pipelines spanning both one-to-one detectors (YOLOv10 and YOLO26) and query-based detectors (RT-DETR). TrackFlood recovers differentiable confidence tensors and optimizes perturbations that flood the tracker with spatially distributed phantom detections while leaving detector inference unchanged. Evaluated entirely on an NVIDIA Jetson AGX Orin (TensorRT FP16), detector latency remains essentially constant, whereas tracker latency increases substantially. At a standard imperceptible budget ($L_\infty{=}8/255$), a universal perturbation produces clearly measurable tracker overload but only modest end-to-end slowdown, without deadline misses for the association-dominated trackers; a separate higher-budget stress test drives severe end-to-end slowdowns and sustained deadline misses. We further evaluate a lightweight, architecture-agnostic bounded-admission layer that caps the tracker workload and largely restores end-to-end latency, at a non-trivial cost in admitted clean detections. Our results demonstrate that evaluating NMS-free perception systems requires considering downstream tracking and end-to-end timing, not detector inference alone.
cs.CR / 5 / 2609.33992
GateDrain: Availability Attacks and Admission-Side Defense for Confidence-Gated Edge-Cloud Inference
Zonghua Gu, Julian Singh-Smith, Junlin Liao, Di Liu
cs.CR
Abstract
Confidence-gated edge--cloud inference accepts confident local predictions and offloads uncertain inputs to a stronger cloud model. We show that this routing decision creates an availability attack surface. We call this attack \emph{GateDrain}: bounded input perturbations lower calibrated confidence and redirect requests that would otherwise be answered locally into a shared cloud queue, without increasing the application request rate. Because escalated requests share a cloud service, an increase in per-request cloud demand can move a near-capacity deployment across a queueing knee, causing disproportionate tail-latency degradation for benign users. We evaluate white-box, transfer, decision-only, universal, and multi-gate attacks on the public EdgeBoost artifact. A fixed-application-volume comparison isolates the effect of confidence manipulation from added client traffic, while perturbation-budget and arrival-process sweeps show that the queueing transition persists across several workload models but its amplification depends on the operating point. Adaptive attacks also defeat the evaluated training-free preprocessing defenses. To contain the resulting cloud demand, we evaluate Bounded Escalation, which combines per-source admission budgets, protected capacity, and non-preemptive trusted-class priority; an optional global bucket adds an identity-independent bound on untrusted admissions. The evaluation makes the resulting policy trade-off explicit: authenticated clients receive latency isolation, whereas tighter aggregate containment can reject legitimate unauthenticated offloads and reduce overall expected accuracy through edge fallback.
cs.CR / 6 / 2609.34003
DeMark: A Query-Free Black-Box Attack for Quality-Preserving Audio Watermark Removal
Weikang Ding, Binhao Ma, Hanqing Guo, Rui Duan
cs.CR
Abstract
Audio watermarking protects digital speech by embedding imperceptible signals for ownership verification and misuse tracing. However, the security of learning-based watermarking remains insufficiently understood under realistic adversarial removal, where attackers cannot access or query the watermark encoder, decoder, or detector. Existing attacks either rely on model feedback, require clean-watermarked pairs, or reconstruct the waveform with generative models, often leading to high query costs, limited generalization, or degraded perceptual quality. In this paper, we propose DeMark, a query-free black-box attack for quality-preserving audio watermark removal. Our key insight is that watermark embedding, while perceptually hidden, can introduce subtle non-speech artifacts in the time-frequency domain that are not fully aligned with natural speech. DeMark removes watermarks by suppressing these artifacts through two stages: Diverse Artifact Learning, which extracts complementary non-stationary and stationary artifact patterns, and Adaptive Artifact Scaling, which adaptively combines and amplifies them under quality-preserving constraints. Across two speech datasets and four state-of-the-art watermarking methods, DeMark achieves average attack success rates of 0.92 and 0.96 while consistently preserving higher perceptual quality than existing adaptive attacks. These results reveal a practical vulnerability of current audio watermarking systems and call for more robust watermark designs against query-free adversarial removal.
cs.CR / 7 / 2609.34027
PerceptFence: Content-Mediation Architecture and Deterministic Coverage for Screen-Share AI Assistants
Asmita Negi, Neeraj Kumar Singh Beshane
cs.CR
Abstract
Live screen-share AI assistants observe raw screen and speech streams, but users have little runtime control over what an assistant may observe, retain, or disclose. Prompt-level privacy settings are insufficient because sensitive content enters through the capture stream. We present PerceptFence, a content-layer mediation architecture between capture, memory, and model responses, with a deterministic synthetic-fixture scaffold; the artifact omits live capture, category inference, authenticated re-consent, cross-session state, and an external model adapter. On 9,600 protocol-documented adversarial strings scored by a separately implemented exposure oracle, PerceptFence neutralises 0.828 of digit-PII payloads on the 5 seeds both systems run, versus 0.183 for Microsoft Presidio; outside that family Presidio leads 0.238 to 0.154, so the overall 0.398 to 0.260 comparison is only indicative. We then evaluate the path a deployed assistant uses: 480 synthetic developer-support screens rendered by Chrome, degraded, and read by OCR, with rules frozen before testing and three screen types held out. PerceptFence neutralises 889 of 968 OCR-surviving secrets and PII values (0.918; Wilson 95% 0.899-0.934) against 0.581 for Presidio and 0.179 for gitleaks, and 0.974 on the held-out screen types, at a measured cost of 0.763 task-token retention on those types. The contribution is a documented mediation architecture and an evaluation method with explicit coverage boundaries, not a claim of live deployment, formal privacy, novel redaction primitives, or general model robustness.
cs.CR / 8 / 2609.34080
LoRo-Mark:Provably Lossless and Robust Agent Watermarking
Haoyang Zou, Yao Wang, Jun Yao, Weiming Zhang, Han Fang
cs.CR
Abstract
As LLM agents are increasingly deployed as commercial services, protecting proprietary orchestration logic and tool-use policies is important. We consider agent repackaging: an adversary integrates a protected agent into its own application via API and presents it under its own identity. It may modify parts of execution to obscure the source. The owner typically has only black-box access to the repackaged service, so black-box ownership verification is essential. Agent watermarking embeds ownership evidence into agent behavior for later verification. An effective watermark should satisfy two requirements: losslessness, preserving original functionality, and robustness, keeping ownership evidence recoverable after partial modification of execution. Existing methods often embed signals into behavior selection or execution trajectories, intervene in normal decisions, and offer limited robustness to behavior modification. We propose LoRo-Mark, a provably lossless and robust agent watermarking mechanism. For losslessness, it isolates watermarking into a cryptographically authenticated forensic branch that remains inactive during normal execution and is activated only by owner-authorized requests. By reducing unauthorized branch activation to standard MAC security, LoRo-Mark formally guarantees performance preservation. For robustness, it redundantly distributes ownership information across forensic behavior sequences, enabling reliable recovery under partial behavior substitution and sequence truncation. Experiments across multiple LLM agents show zero degradation on normal tasks and reliable ownership verification under sequence modifications.
cs.CR / 9 / 2609.34140
Long-Term Operational Planning Using Scenario-Based System Load Forecasting
Akash Debnath, Shiuli Subhra Ghosh, Jaime De La Ree, Alex Freeman, Kavin D Jones
cs.CR
Abstract
The rapid expansion of hyperscale data centers is significantly increasing electricity demand in Northern Virginia. Dominion Energy, the region's primary electric utility, must reinforce its transmission network to support this growth. These projects require planned outages that must be evaluated months in advance to support construction planning and outage coordination while meeting NERC and PJM N-1 reliability requirements. Long-term outage studies are commonly performed day by day using monthly or seasonal peak-load assumptions. Although this approach simplifies analysis, it can be overly conservative because it does not capture granular load variability. Consequently, short-duration outages that may be feasible under realistic loading conditions are often postponed or denied, delaying critical transmission expansion and grid modernization projects. This paper evaluates the operational value of incorporating realistic multigranular load forecasts into long-term, contingency-based outage assessments and compares the results with conventional peak-based methods. Daily, weekly, and monthly forecasts are developed using statistical and machine-learning models, including SARIMA, Prophet, Gradient Boosting, and Random Forest, followed by bottom-up temporal reconciliation to maintain consistency across forecast horizons. Results show that granular load forecasts reduce unnecessary conservatism and improve outage accommodation, particularly for short-duration requests, without changing existing reliability criteria. The proposed approach can strengthen long-term outage coordination and support timely transmission reinforcement and grid modernization.
cs.CR / 10 / 2609.34245
Certified Multi-Source Integrity for Structured Agent Actions
Anmol Pandey, Aditya Jain, Liang Chen, Carsten Maple, Christo Panchev
cs.CR · cs.AI · cs.LG
Abstract
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-chosen values. Current defenses gate on a source's trust label or certify free-text answer quality. None certifies the integrity of a coupled, policy-bound structured action under a corruption budget that accounts for shared upstream sources. We characterize when such an action is safely certifiable and give the maximally live safe certifier. It admits an action only when each field clears the rule its evidence structure supports: a bounded corruption radius over corruption-distinct evidence classes, counted by a minimum hitting set so that re-publishing or laundered copies cannot manufacture a quorum, deterministic reconciliation for complementary fields, and a trusted anchor where the evidence leaves a field single-sourced. We formalize two robustness notions, validate each mechanism by ablation, and measure how often the multi-source precondition holds on sanctions designations (70,966 entities) and software supply-chain provenance (450 packages). Under upper-bound proxies, genuine corroboration is a minority phenomenon in both, and naive attestation counting overstates it, since witnesses that look independent collapse to two corruption-distinct domains once shared origin is counted. Across five current models in a real agent loop, a realistic injection fools every model but one and a naive agent then executes the fraudulent action on most attacks. The certifier admits no unsafe action and recovers the correct value where corroboration permits, while action-gating and provenance baselines are broken in every world of our harness by some attack in its space.
cs.CR / 11 / 2609.34456
Breaking Windows Malware Detection: A Comprehensive Evaluation of Problem-Space Adversarial Robustness
Mashal Zainab, Salijona Dyrmishi, Hamid Bostani, Lorenzo Cavallaro, Maxime Cordy
cs.CR
Abstract
Problem-space evasion attacks have exposed critical weaknesses in machine learning-based malware detectors; yet, their evaluation remains fragmented across models, datasets, and attack methodologies, often neglecting domain-specific requirements such as executability and functionality preservation. We address this gap with a unified, large-scale evaluation of nine state-of-the-art evasion attacks against eight Windows malware detectors, including seven open-source models and one commercial detector, under executability-preserving conditions. Our study analyzes attack effectiveness, complementarity, transferability, and adversarial hardening to evaluate robustness along complementary dimensions. We show that detector vulnerability depends strongly on both model representation and attack type: raw-byte detectors are particularly susceptible to several classes of problem-space manipulation, but no detector family is uniformly robust across all attacks. Importantly, effectiveness is not explained by transformation-space size alone: the strongest attacks can achieve substantially higher success while using fewer distinct transformations and concentrating on a small set of high-impact manipulations. We further show that two complementary attacks are sufficient to cover approximately 99% of the adversarial examples produced by the remaining evaluated attacks. Transferability exhibits a different pattern from direct attack success: attacks with low direct success can produce highly transferable evasions. Finally, adversarial hardening is highly attack- and model-dependent: robustness gains often fail to transfer across attacks and can even increase susceptibility to unseen attacks. These findings highlight limitations in current malware robustness evaluations, establish a comprehensive empirical baseline, and clarify relationships between effectiveness, transferability, and defense robustness.
cs.CR / 12 / 2609.34518
SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents
Xinjie Shen, Junran Wang, Rongzhe Wei, Pan Li
cs.CR · cs.AI
Abstract
Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and defense as partially observed state control in SEAD, deriving their design requirements from this shared execution process. Because attackers supply instructions while the target chooses concrete actions, DART decomposes harmful goals into locally plausible steps and uses feedback from actual tool execution to guide trajectory search. The defender must decide before execution with incomplete state evidence. SAGE can therefore investigate relevant state through read-only queries before allowing or blocking each action, including those proposed after a block. We construct an environment-verifiable dataset integrating controlled initial states, replayable tool environments, and task-specific executable checks. Across four target models, DART improves semantic attack success by 18.8--35.9 percentage points over the competing baseline, with consistent gains under executable verification. On recorded trajectories, SAGE preserves 95.79% of benign trajectories while intercepting 92.73% of harmful paths by the harm-enabling boundary. In online attack-defense evaluation, it reduces DART's executable attack success from 48.0% to 4.0%. SAGE remains effective across four attack methods and generalizes to out-of-domain environments. Our code and data is available at https://github.com/EverywhereSafety/SEAD.
cs.CR / 13 / 2609.34708
StallGrid: Measuring Internet-exposed Engagement in Protocol-Native OT Tarpits
Arthur Cordeiro, Casper Andersen, Emmanouil Vasilomanolakis
cs.CR
Abstract
Operational technology systems face exposure to scanning and protocol-specific attack tools, where compromise risks disrupting physical processes rather than just data. Traditional OT defenses rely on blocking and filtering under strict patch constraints, while tarpitting delays scanners through sustained protocol-level interaction. Modbus TCP and IEC-104 carry no native authentication or integrity protection. Application-layer tarpitting, which stalls scanners inside a legitimate protocol exchange, has been studied for IT and IoT protocols, but not for OT/ICS, whose session semantics (state machines, exception codes) create stalling opportunities IT/IoT lack, and whose patch-constrained environments need exactly this kind of alternative defense. We present StallGrid, to our knowledge the first application-layer tarpits for OT/ICS, stalling scanners via Modbus Exception Codes \texttt{0x05}/\texttt{0x06} and prolonged residence in IEC-104's connected state machine. Five variants (three Modbus TCP, two IEC-104) ran simultaneously for 24 days online, logging 6,709 sessions, 2,039 cumulative per-tarpit unique IPs, and over 2,000 hours of accumulated connection engagement. GreyNoise enrichment attributes 87--97\% of stall time to malicious-tagged IPs, just 22--30\% of connecting addresses; protocol-level behavior further correlates with malicious classification, a fingerprinting signal beyond raw stall time. Shorter induced delay (1.5s) yielded more total engagement than longer delay (3s), observed across both protocols. These results position protocol-native tarpitting as a practical, low-cost complementary defense for OT/ICS environments where patching remains infeasible.
cs.CR / 14 / 2609.34744
Optimizing and Securing the Modern Watermarking Channel for Images
Enoal Gesny, Eva Giboulot
cs.CR · cs.CV
Abstract
To comply with recent regulations requiring traceable generated content, modern watermarking has adopted multi-bit post-hoc watermarking schemes. These modern designs rest on an encoder-decoder pair implemented as deep neural networks. These models are usually treated as pure black-boxes trained end-to-end, with the noise of the watermarking channel modeled through a fixed set of geometric and valuemetric transforms applied to watermarked images. We argue that this purely empirical approach leads to unquestioned design flaws and a lack of theoretical performance guarantees. This work proposes a general theoretical model of modern post-hoc watermarking schemes grounded in a statistical analysis of the outputs of the encoder/decoder pair. We show that these deep neural networks implicitly define a watermarking channel modeled as parallel AWGN channels, with messages transmitted using BPSK modulation. This imposes a binary alphabet, greatly limiting the capacity of these watermarking systems. Another fatal flaw is their lack of a secret key, making them intrinsically insecure. We make this notion of watermarking security precise for post-hoc schemes by linking it to the possibility of estimating the secret key under a given statistical model of the decoder's output. By putting together the results from this theoretical analysis, we introduce SNW: a novel post-hoc watermarking system that significantly outperforms existing state-of-the-art baselines in terms of capacity while also providing strong security guarantees. Notably, it does not depend on a fixed codebook or binary alphabet, allowing it to reach a rate close to Shannon capacity through the use of capacity-achieving error-correcting codes.
cs.CR / 15 / 2609.34790
CoSec: Benchmarking Agent Security in Communities
Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao, Chenbo Xia, Yuwen Qu, Renxiang Wang, Fudong Yuan, Camil Hamami, Chenglong Pan, Xinquan Yue, Ziyu Wang, Fengyu Ye, Chenyang Si, Caifeng Shan
cs.CR · cs.AI
Abstract
LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthorized disclosure of protected information. Existing evaluations do not fully examine these risks in agent systems. We introduce \textbf{CoSec}, an executable benchmark for evaluating privacy and authorization enforcement in LLM agent systems operating within and across communities. CoSec contains 208 canonical scenarios spanning fixed and evolving boundaries, protected information belonging to the agent owner or other participants, and attacks through dialogue, environmental content, persistent memory, and composed workflows. CoSec executes complete agent systems with persistent sessions, memory, files and tools. It verifies information flows against the active authorization state using execution traces and artifacts. Across harness and model configurations, agents frequently complete benign tasks but violate privacy and authorization boundaries. Privacy behavior varies across harnesses, attack surfaces, and community states, revealing how memory, files, tools, and workflows can carry protected information beyond its authorized scope. These findings show that task utility does not imply privacy or authorization compliance and that authorization in community settings remains an unresolved security challenge for persistent LLM agents.
cs.CR / 16 / 2609.34804
Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure
Jeff Nijsse, Shu Su, Benjamin Oholeguy, Sreenivas Sremath Tirumala
cs.CR · cs.LG
Abstract
Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model updates derived from fabricated telemetry remain indistinguishable from honest contributions. We repurpose cyber-physical process invariants, such as conservation laws and actuator couplings, from runtime detection heuristics into a verifiable admission requirement for federated updates, mined automatically from clean operational data. We evaluate this admission gate across two physical water testbeds (SWaT, WADI) and a distribution benchmark (BATADAL), testing seven aggregation rules against telemetry fabrication, exposure-only replay poisoning, and an invariant-aware adaptive adversary. Across three testbeds the mined invariants reject none of 100 honest shards and all naively fabricated ones, including optimised perturbations that FoolsGold admits in full. On real telemetry, five mined invariants detect 12 of SWaT's 35 attacks, while nine invariants detect 20, with no honest shard rejected. With nine rules, the physics gate recovers 69--100% of the targeted-attack recall lost to replay poisoning, and 54--100% of that lost to fabricated telemetry, across five standard aggregators. To reconcile physical admission control with federated data privacy, we show invariant compliance using zero-knowledge proofs (zk-SNARKs) to allow clients to prove batch adherence without revealing operational telemetry.
cs.CR / 17 / 2609.34862
JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges
Zhiqiang Wang, Yichao Gao
cs.CR
Abstract
Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative judges on four benchmark collections totaling 5,219 trajectories, using a common risk rubric and behavior-level labels. JEV attains a benchmark-averaged positive-class F1 of 77.8, compared with 74.1 for the strongest generative configuration, GLM-5.2, with valid-result coverage of 95.5\% and 94.4\%, respectively. Performance varies across datasets, with JEV leading on ATBench500 and MCPHunt and GLM leading on R-Judge and TraceSafe. Across the four benchmarks, JEV's median successful-call latency is 0.99 seconds; estimated token cost averages \$0.000195 per valid judgment. These results support JEV as an economical screening signal, with trade-offs in precision and recall.
cs.CR / 18 / 2609.35040
GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation
Hwanjo Heo, Junhee Lee, Jinwoo Kim
cs.CR
Abstract
Mixed-reality headsets such as Apple Vision Pro replace the touch screen with gaze-as-pointer interaction: the wearer looks at a target and confirms with an air pinch. Because the display is inside the headset and the eye tracker is walled off from third-party software, such input is widely assumed to be unobservable to bystanders---a built-in defense against the shoulder-surfing that plagues phones and laptops. We present GAZEleak, a side-channel attack that recovers gaze-driven input from external video of head motion alone, under a strictly local, physical-observer threat model: the adversary only films the wearer from across the room and installs no software on the device. The attack exploits the centrally coupled eye-head motor program: gaze shifts recruit small, target-dependent head reorientations that project into sub-degree pose changes recoverable from commodity video. GAZEleak implements a measurement-based, sparse-optical-flow inference pipeline for users' 6-digit device passcodes. On a preliminary front-view dataset from three author-subjects who were aware of the attack hypothesis, GAZEleak places the true code within the top ten guesses for 56% (10 of 18) of test codes under a cross-person protocol with no labeled victim data, and for every code of the most exposed subject. The performance is subject-dependent: no passcode from the least exposed subject reaches the top ten, although its median guessed passcode rank is 12,786 rather than 500,000 expected from an uninformative ordering. These results provide preliminary evidence that gaze-coupled head motion can expose passcode information under controlled conditions, while motivating broader evaluation across users, behaviors, and capture settings.
cs.CR / 19 / 2609.35053
Efficient TCitH-Based Alternatives to SLH-DSA: Cross-Layer ASIC Design of Mirath
Hiandra Tomasi, Maximilian Schöffel, Johannes Feldmann, Norbert Wehn
cs.CR · cs.AR
Abstract
To address the security risks posed by quantum computers, the U.S. National Institute of Standards and Technology (NIST) has standardized the post-quantum signature schemes ML-DSA, FN-DSA, and SLH-DSA. While ML-DSA and FN-DSA are lattice-based, SLH-DSA relies on hash-based assumptions. To support cryptographic agility against future vulnerabilities, NIST is evaluating non-lattice candidates as alternatives to SLH-DSA. Among these, TCitH-based schemes are particularly promising due to their compact keys and small signatures. However, their high computational complexity and memory footprint pose significant challenges for efficient implementations on resource-constrained embedded platforms. They remain largely unexplored in this context, particularly in ASIC implementations. To address this gap, we use a cross-layer methodology combining algorithmic and hardware layers to present, to the best of our knowledge, the first ASIC implementation of Mirath, a TCitH-based signature scheme, in a RISC-V-based system. The design is implemented in a 22nm FD-SOI technology node. Compared with an SLH-DSA ASIC implemented in the same technology node, the proposed architecture achieves 17.8x lower signing latency while requiring 58% less total cell area, showing the potential of TCitH-based signatures as efficient non-lattice alternatives from an implementation perspective.
cs.CR / 20 / 2609.35155
LENS: The Sum Is Worse Than the Parts for Set-Level Poisoning in Retrieval-Augmented Generation
Kaisheng Fan, Yishu Gao, Xunzhu Tang, Tegawend'e F. Bissyand'e, Weizhe Zhang
cs.CR
Abstract
Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper subset remains insufficient in frozen single-round RAG. We formalize set-level compositional poisoning, where documents designed to remain individually plausible jointly redirect RAG outputs to a target answer, while proper subsets fail to induce the target on their own. To construct such attacks, we propose LENS, a generator-black-box multi-agent framework that casts construction as constrained evidence composition. LENS factorizes target inference into a query-conditioned interpretation lens and complementary facts, then uses a nested dual-loop workflow to concentrate steering in the full set while suppressing subset leakage. The outer loop plans the interpretation lens and semantic roles; the inner loop synthesizes documents and applies counterexample-guided repair. Across four benchmarks and three generators, returned packets achieve 0.852 full-set ASR and 0.784 post-retrieval ASR@5, while their strongest proper subsets reach only 0.069. Against construction baselines evaluated on the same frozen manifest, LENS improves all-attempt E2E-Strict@5 from 0.244 to 0.363, a 48.8% relative gain. A blinded human audit finds that 68.3% of returned packets combine an incorrect target, a definite answer-criterion shift, and no target entailment under the original semantics. Across four published defenses, LENS attains the highest defended all-attempt ASR@5, exceeding the strongest baseline by 0.141 on average. Together, these results establish evidence composition as a distinct RAG security boundary and position LENS as a stress test for defenses that reason over document sets.
cs.CR / 21 / 2609.35199
Trajectory-Level Security Debt in LLM Coding Agents
Prateek Kumar Rajput, Abdoul Kader Kabore, Yewei Song, Melissa Tessa, Tailia Malloy, Jacques Klein, Tegawendé F. Bissyandé
cs.CR · cs.SE
Abstract
LLM coding agents can traverse hundreds of intermediate code states before submitting a solution. Evaluating only the final artifact leaves the evolution of security findings unmeasured. We introduce the Security Debt Line Integral (SDLI), which accumulates static-analysis risk when an agent reaches a new best test pass ratio. We instantiate it with four static application security testing (SAST) tools and study artifacts from 830 passing SWE-bench runs, 712 ProgramBench final workspaces, and 13 public MirrorCode trajectories. The two large populations use the final-state special case of SDLI. Two-tool Common Weakness Enumeration (CWE) class agreement occurs in 3.9% of SWE-bench runs and 26.2% of the 80 ProgramBench runs passing at least 90% of official tests. These are scanner findings, not validated vulnerability rates. Excluding three advisory-heavy classes reduces the latter rate to 6.2%. Same-task runs differ in their measured scores, while one reconstructed ProgramBench run exposes persistent findings from its first implementation write. A repair case study reduces the scanner signal while preserving tested behavior, but also reveals sensitivity to equivalent API rewrites. SDLI offers a way to study progress and security findings together. Its value for steering agents and confirming exploitable vulnerabilities remains to be established.
cs.CR / 22 / 2609.35205
OT-PCA: New Key-Recovery Plaintext-Checking Oracle Based Side-Channel Attacks on HQC with Offline Templates
Haiyue Dong, Qian Guo
cs.CR
Abstract
In this paper, we introduce OT-PCA, a novel approach for conducting Plaintext-Checking (PC) oracle based side-channel attacks, specifically designed for Hamming Quasi-Cyclic (HQC). By calling the publicly accessible HQC decoder, we build offline templates that enable efficient extraction of soft information for hundreds of secret positions with just a single PC oracle call. Our method addresses critical challenges in optimizing key-related information extraction, including maximizing decryption output entropy and ensuring error pattern independence, through the use of genetic-style algorithms. Extensive simulations demonstrate that our new attack method significantly reduces the required number of oracle calls, achieving a 2.4-fold decrease for hqc-128 and even greater reductions for hqc-192 and hqc-256 compared to current state-of-the-art methods. Notably, the attack shows strong resilience against inaccuracy in the PC oracle-when the oracle accuracy decreases to 95%, the reduction factor in oracle call requirements increases to 7.6 for hqc-128. Lastly, a real-world evaluation conducted using power analysis on a platform with an ARM Cortex-M4 microcontroller validates the practical applicability and effectiveness of our approach.
cs.CR / 23 / 2609.35234
Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance
Guy Lupo, Nguyen Hung Nguyen, Viet Vo, Chamikara M. A. P., Guangdong Bai
cs.CR
Abstract
Agentic AI systems increasingly act via tools, memory, delegation, and external services. Existing observability and provenance mechanisms can reconstruct events post hoc, but they rarely show, at the time of the record, whether each policy-relevant action was checked by the intended control before execution. This leaves a trust-observability gap for continuous monitoring, detection, and response: later assurance may rest on evidence that is incomplete, privacy-leaking, mutable, or detached from the policy context that governed the event. What's missing in the literature is contemporaneous, policy-bound evidence that the intended control was evaluated under the policy in force at the time. We introduce ProofWeave, a record-time chain-of-evidence concept for agentic AI assurance. At each policy-relevant action boundary, ProofWeave generates a privacy-minimised and integrity-anchored evidence transaction that binds (i) agent intent or action, (ii) control response, and (iii) a policy-at-time snapshot. Each transaction is committed to an append-only ledger and materialised into a derived proof graph. A bounded Weaver Agent translates policy intent into proof obligations, while deterministic validators check evidence completeness, privacy minimisation, policy binding, and integrity. In the minimal scenario, an agent attempts to transmit a secret to an unapproved external sink. The audit compares a logs-only correlation baseline with ProofWeave across verdict latency, join ambiguity, privacy exposure, tamper detection, and resistance to graph-only proof injection. ProofWeave reduces candidate bindings per verdict from up to `10,201` to one, validation operations from up to `10,201` to approximately `26`, and assurance evidence storage from `0.79`MiB to `0.15`MiB per project.
cs.CR / 24 / 2609.35256
Implementing Data Diodes Using Commodity Hardware and Open Source Software
Peter Story, Gert-Jan den Besten
cs.CR · cs.NI
Abstract
One-way network devices, known as data diodes, are used to defend against sophisticated cyberattacks. Partly due to their high cost, data diodes are mostly deployed in nuclear power plants and within the government for handling classified information. Although commercially available data diodes are expensive, a data diode's hardware can assembled from commodity fiber-optic network equipment. However, specialized software is needed to send data through a data diode reliably: the receiving program cannot request retransmission of dropped packets, so packet loss must be minimized and mitigated. First, we developed a minimal program to measure packet loss. We found that most packet loss was caused by the receiving program processing incoming packets too slowly, and that packet loss often occurs in clusters. Also, we discovered ways to minimize packet loss on Linux and macOS without using superuser privileges. Next, we tested three existing open source programs for one-way data transfers: netcat, UDPcast, and lidi. Although these programs were unreliable in their default configurations, we identified reliable configurations for UDPcast and lidi. Finally, we incorporated our findings into pydiode, our cross-platform program for reliable one-way data transfers.
cs.CR / 25 / 2609.35300
Sustained Participation as a Security Resource: The Bounded Participation Channel
Homayoun Maleki, Nekane Sainz, Jon Legarda, Igor Santos-Grueiro
cs.CR
Abstract
Can sustained, per-identity participation be engineered into a security resource? Most anti-Sybil defenses price identity creation rather than identity survival. Once admitted, an adversary may sustain many identities without paying a recurring cost. We introduce the Bounded Participation Channel (BPC), a formal primitive for repeatedly verifying participation window by window. BPC issues fresh, identity-bound challenges under a strict deadline and enforces four structural properties: identity binding, freshness, real-time response, and bounded per-channel throughput. Together, these yield a provable cost theorem: sustaining $s$ identities over $T$ windows requires $C(s,T) \geq sT/τ_h$ participation channel-windows. The guarantee is solver-agnostic: a channel may be operated by a human, an AI system, or a hybrid. We give a hash-based construction with publicly verifiable participation proofs, characterize four admissible challenge families, and evaluate two against GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.5 across 600 trials. Despite near-perfect accuracy (97--100%) on the perceptual tasks, the evaluated automated channels remain throughput-bounded under the tested deployment conditions. The results illustrate a key distinction: solvability does not imply unlimited throughput. By requiring participation to be re-earned by every identity in every time window, BPC turns sustained participation into a measurable security resource with a linear structural cost floor, independent of whether the participation is supplied by humans, AI systems, or hybrids.
cs.CR / 26 / 2609.35414
AI-Based Vulnerability Assessment Capability and Cyber Attack Graph Analysis
Joni Herttuainen, Kirsi Hellsten, Vesa Kuikka, Ambrose Kam, Arlanda Johnson, David Welsh, Kimmo K. Kaski
cs.CR
Abstract
Cyber threats targeting mission-critical infrastructure are becoming more sophisticated while the barrier to launching attacks continues to fall. Traditional point solutions like antivirus and firewalls are reactive and fail to address the combinatorial complexity of modern attack surfaces. This paper presents an investigation combining two complementary methodologies: Lockheed Martin's Vortex/Crow framework, which applies multi-agent reinforcement learning (MARL) over industry-standard cyber knowledge graph to identify and prioritize attack vectors and TTPs (tactics, techniques, and procedures); and Aalto's probabilistic attack graph model that combines network topology and its vulnerabilities to compute system-level risk metrics. The 2015 Ukraine Power Grid cyberattack serves as a well-documented validation scenario. Applied independently to the same operational technology (OT) network topology, both methodologies converge on the same attack vectors and exploit sequences as those documented in the incident record, thus providing mutual cross-validation. Attack graph analyses using node-level elimination experiments identify industrial control systems (ICS) as the most critical enablers of attack propagation, representing high-priority targets for defensive hardening. Comparison of CVSS (v2.0) and IronMiner vulnerability scoring yields in general consistent results, with IronMiner providing more actionable differentiation at network periphery nodes. The layered methodology of baseline assessment and node-level elimination proves to be scalable to large enterprise networks, thus offering defenders a structured, AI-enabled path to prioritize mitigation under realistic time and resource constraints.
cs.CR / 27 / 2609.35421
Indistinguishability of Sum of Permutations: A Fourier Analytic Route to Classical and Quantum Security
Ritam Bhaumik, Chun Guo, Xiaoning Guo, Ashwin Jha
cs.CR · math.CO
Abstract
We study classical and quantum indistinguishability of sums of independent random permutations and related transformations from permutations to functions. Let $G$ be a finite abelian group of order $N$, and let $π^k_+(x)=π_1(x)+\cdots+π_k(x)$ for $k\geq2$ independent uniform random permutations of $G$. We give a unified Fourier analytic treatment in which the construction is represented by its probability density and a distinguisher by its acceptance function, with the classical and quantum query models imposing different restrictions on the Fourier support of the latter. Classically, we obtain the bound $O_k(q/N^{k-1/2})$ for every $q<N$, and refine it below the birthday threshold to $O_k(q^2/N^k)$. In the quantum model, a simulation argument gives $O_k(N^{-(k-3/2)})$ for $q\leq(N-1)/2$, while Fourier interpolation gives concrete finite bounds up to $q\leq4N/15$ and the query-dependent bounds $O\left(\min\left\{N^{-1/2},q^3/N^2 + 1/N\right\}\right)$ and $O_k\left(\min\left\{q^3/N^k,N^{-(k-3/2)}\right\}\right)$, for $k=2$ and $k \geq 3$, respectively, throughout $1\leq q\leq(N-1)/2$. For $q = 1$, the first bound sharpens to $O(N^{-2})$. Over $G=\mathbb F_2^n$, a one-query Fourier attack matches the order of our one-query bound, while an $N/2$-query parity attack with advantage $1/2$ shows that our bounds reach the constant-advantage query threshold. We further study two variants of sum of permutations over binary vector spaces. First, we allow arbitrary surjective linear postprocessing, which includes truncation, and obtain classical and quantum bounds that retain the output-size dependence. Second, we analyse Dinur's variable-output single-permutation construction, $\mathsf{LXoP}$, for every fixed output width, and derive its classical and quantum security bounds; for one- and two-block outputs, we give concrete quantum security bounds.
cs.CR / 28 / 2609.35487
Beyond Scalar Probes: Exploiting Vector-Valued Outputs in ReLU Networks For Signature Extraction
Gorka Abad, Claude Carlet, Ermes Franch, Stjepan Picek, Vincent Rijmen
cs.CR
Abstract
We revisit cryptanalytic extraction of ReLU networks from a geometric and algebraic perspective. Rather than restricting attention to a single output component, we study the full vector-valued behavior across adjacent linear regions. This leads to a rank-one characterization of Jacobian differences that recovers the usual row-signature information while also revealing complementary column-side information. Our experiments show how this additional structure can be used in ex- traction and improves numerical estimation under different numerical- precision regimes (float64, float32 and float16). We extend the analysis beyond the high-precision and output-rounding settings commonly con- sidered in the literature towards the low-precision settings encountered in many practical settings.
cs.CR / 29 / 2609.35506
Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning
Emad Efatinasab, Denis Donadel, Mirco Rampazzo, Chuadhry Mujeeb Ahmed
cs.CR
Abstract
Cyber-physical power systems increasingly rely on data-driven tools to detect instability and support reliable grid operation. However, reliable stability prediction is difficult when unstable operating configurations are rare, sensitive, or unavailable during model development, since collecting such data safely and at scale is often impractical. At the same time, False Data Injection (FDI) attacks can manipulate reported system parameters to trigger false instability alarms or conceal unsafe operation, and such threats in Decentral Smart Grid Control (DSGC) systems remain largely unexplored. These challenges are rarely addressed jointly, leaving a gap between stability prediction and attack detection that this work aims to close. In this paper, we introduce StarGNN, a graph learning framework that learns the stable operating region exclusively from clean stable configurations and uses a single abnormality score to flag reported configurations that should not be trusted as evidence of safe operation, covering both genuine instability and unseen FDI manipulations. Each configuration is represented as a producer-consumer star graph and processed by a role aware graph neural network, with physics constrained pseudo-negatives generated by perturbing reaction time and price response parameters standing in for the unavailable unstable and attack data. A single threshold, calibrated only on held-out stable data, is used without task or attack specific adjustment. Evaluated on nine unseen FDI scenarios, StarGNN detects between 0.780 and 0.973 of attacks on stable configurations and retains a post attack instability recall between 0.972 and 0.999 on unstable ones, with 0.899 recall against a stronger adaptive attacker, showing that stable-only boundary learning can support both stability prediction and attack detection without access to genuine unstable labels or attack samples during training.
cs.CR / 30 / 2609.35552
INTCC: A Framework for Interactive Confidential Computing
Qingzhe Bing, Kaiyuan Zhang, Yinqian Zhang
cs.CR
Abstract
Confidential computing leverages Trusted Execution Environments (TEEs) to ensure the confidentiality and integrity of data in use. However, TEEs rely on remote attestation to guarantee the integrity of their initial memory state. This model is fundamentally at odds with interactive development workflows. In scenarios like LLM fine-tuning and exploratory data analysis, data processors need human-in-the-loop capabilities, including dynamic code injection, intermediate state inspection, and hyperparameter tuning, all of which inherently violate the static, one-time integrity guarantees of traditional remote attestation. To reconcile this tension, we propose the interactive confidential computing paradigm, a system architecture enabling untrusted data processors to execute dynamic, non-deterministic operations within TEEs without compromising data confidentiality. Driven by the insight that inherently unmeasurable human interaction must be excluded from the Trusted Computing Base (TCB), we logically partition the TEE into an interactive controller and a verifiable runtime. To realize this paradigm, we present INTCC, a framework featuring three key mechanisms: (1) a proxy-based dispatch system to preserve the native development experience; (2) a fine-grained information flow control mechanism based on a security lattice to prevent data leakage; and (3) a privacy-preserving verifiable execution mechanism to guarantee the runtime compliance of dynamic workflows. We implement INTCC on AMD SEV-SNP using Confidential Containers and evaluate it across diverse real-world workloads. Our experiments demonstrate that INTCC effectively balances security and interactivity, incurring a practical overhead of less than 5% for LLM fine-tuning and under 17% for data analysis relative to baseline execution.
cs.CR / 31 / 2609.35557
The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent
Shobhan Roy
cs.CR · cs.AI · cs.SE
Abstract
The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.
cs.CR / 32 / 2609.35596
SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi, Ferdinando Fioretto
cs.CR · cs.AI · cs.CL
Abstract
Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operations, we provide an adaptive trajectory discovery pipeline that probes for failures while preserving original task intent and supports causal attribution through paired non-evolving agents and attribution scores. Our evaluation across multiple recent LLMs, evolution surfaces, and harm types reveals that self-evolution indeed increases task completion rates but often at the cost of safety failures that are absent for paired non-evolving baseline agents. We also show that qualitatively different safety behaviors emerge across evolution surfaces and harm types. Further, we show that this divergence in safety behavior is reflected in agents' chain-of-thought reasoning, which yields an effective monitoring strategy that can mitigate unsafe behavior with a low false positive rate.
cs.CR / 33 / 2609.35171
Analyzing Solana's Blocks and Transactions
Yaron Hay, Dvir David Biton, Roy Friedman
cs.DC · cs.CR
Abstract
Solana is one of the most popular blockchains, and is arguably the most widely used blockchain for smart contracts, also known as dApps. Understanding the types of smart contracts that are being executed by Solana and their interplay is therefore highly beneficial both for designers of modern blockchains and developers of smart contracts. To that end, in this paper we analyze a million recent Solana blocks. We report statistics about the size of blocks (number of transactions per block), execution time units for individual transactions and fees, and invoked Solana programs. Further, based on the declared readset and writeset of each transaction, as mandated by Solana, we analyze the conflicts and corresponding conflict graphs arising within each block. These latter statistics are important to understand the potential for parallelism in the network, which is one of the main claimed benefits of Solana. The data and code are available in open source.
cs.CR / 34 / 2609.35358
Hybrid QKD-PQC Network Emulation through Automated and Scalable Cloud-Native Orchestration
Iván Melijosa, Javier Pérez, Borja Nogales, Iván Vidal, Francisco Valera
cs.NI · cs.CR
Abstract
The ongoing transition toward quantum-safe networking has motivated the development of hybrid network architectures integrating Quantum Key Distribution (QKD) and Post-Quantum Cryptography (PQC). However, the experimental evaluation of hybrid QKD-PQC network architectures remains constrained by the high cost and limited accessibility of quantum hardware, as well as by the limited support for hybrid QKD-PQC networks in existing emulation platforms. Quditto is an open-source emulation platform originally designed for QKD networks that enables cost-effective and reproducible experimentation without requiring dedicated physical quantum infrastructure. Building on this foundation, this work presents Quditto as a hybrid QKD-PQC network emulation platform featuring automated and scalable cloud-native orchestration. The proposed platform introduces four principal contributions: a cloud-native orchestrator enabling fully automated infrastructure deployment across cloud and multi-cluster environments; an optimized provisioning workflow enabling large-scale quantum-safe network emulation; native integration of post-quantum nodes enabling unified emulation of hybrid QKD-PQC networks; and a secure key management module providing persistent and access-controlled storage of cryptographic material. Experimental validation demonstrates sublinear orchestration-time scaling with network size and successful end-to-end hybrid QKD-PQC key establishment on a representative spine-leaf deployment, thereby enabling the systematic evaluation of quantum-safe networking mechanisms in large-scale heterogeneous network environments.
cs.CR / 35 / 2609.34089
Faultless: A Program Equivalence Technique for Validating and Evaluating Neural Decompilers
Luke Dramko, Claire Le Goues, Edward Schwartz
cs.PL · cs.CR · cs.SE
Abstract
Neural decompilers are machine learning models which perform the process of decompilation, lifting code from a lower-level language to a higher one. Neural decompilers offer substantial utility relative to traditional deterministic decompilers because they can probabilistically recover information discarded during lowering, like variable names, types, and control flow structuring. However, they can also make mistakes, producing code that is not equivalent to the original, making it difficult to trust their output. In this work, we introduce Faultless, a program equivalence technique for performing translation validation on neural decompilers. Faultless compares code produced by a deterministic decompiler, which has stronger correctness properties, with that of a neural decompiler. Faultless is also useful for model evaluation, a highly related task, in which the neural decompilers' prediction is compared with a reference solution. Neural decompilation introduces significant challenges to the task of program equivalence which existing techniques are not equipped to handle, including limited extrafunctional context and systematic semantic inconsistencies in decompiled code. Faultless takes a static symbolic execution-based approach with an execution model and memory model designed to handle these challenges.
cs.CR / 36 / 2609.34411
Coherence Rather Than Error Rate Governs Privacy in Multi-Tenant Quantum Computing
Farhad Farokhi
quant-ph · cs.CR
Abstract
Multi-tenant computing enables providers of commercial cloud quantum processors to rent disjoint sectors of a device to independent users. Average gate error, which cloud quantum computing providers report, does not determine how much one tenant learns about another. We propose an information-theoretic notion of information leakage across co-tenancy boundaries stemming from quantum state distinguishability. We measure this leakage on commercially-available 156-qubit (IBM Kingston) and 20-qubit (IQM Garnet) devices. Boundaries with identical benchmarked error can offer significantly different amount of information leakage because standard reported measures of error are blind to coherent-versus-stochastic nature of the error while the proposed notion of information leakage is not. A uniform Pauli randomisation implemented over the victim's whole register is used as a defence mechanism to reduce the information leakage to zero. The defence theoretically does not incur a fidelity cost, but the experiments show a non-trivial degradation caused by accumulation of errors. We provide a specific call-for-action to the providers of quantum cloud computing to report information leakage in addition to standard error rates in their device datasheet to enable users to compute privacy and security risks prior to engagement with the device.
cs.CR / 37 / 2609.34413
Quantum Security of XOR of Permutations via Fourier Analysis
Wonseok Choi, Minki Hhan, Junyoung Jang
quant-ph · cs.CR
Abstract
The XOR of two or more independent random permutations (XoP) is the prototypical pseudorandom function built from permutations achieving security beyond the birthday bound. The classical security of the XoP construction is well established, but its security against quantum attacks that query XoP in superposition has remained widely open. We prove that the XOR of $r\ge 2$ random permutations over $\{0,1\}^n$ is indistinguishable from a random function by any $q$-query quantum algorithm with advantage \[O\left(\min\left\{\frac{q^3}{2^{rn}},\frac{q^{1.5}}{2^{(r-0.5)n}},\frac{1}{2^{(r-1.5)n}}\right\}\right)\] for all $q\ll 2^n$. In particular, XoP remains secure throughout the entire query range, far beyond the $2^{n/3}$ bound due to quantum collision finding attacks. This is the first construction from permutations that achieves the quantum version of the beyond birthday bound security. We also present several heuristic attacks suggesting the tightness of our bounds in ranges $q\le 2^{n/2}$ and $\approx 2^n$. We use a Fourier-analytic variant of the polynomial method: the advantage of any $q$-query quantum algorithm is controlled by the Fourier components of degree at most $2q$, or by $2q$ input-output data of the construction. The norms of most components are bounded well, proving the bound $2^{-(r-3/2)n}$. The norm of low-degree components turn out to be too large for the bounds $q^3/2^{rn}$ and $q^{1.5}/2^{(r-0.5)n}$. We reinterpret these low-degree components as (sums of) advantages of the other problems. For example, the degree-2 and degree-4 terms are interpreted as the advantages against random functions with and without \emph{planted collisions}, which in turn are bounded using Zhandry's small-range distributions. Along the way, we prove a new bound for the small-range indistinguishability for (ironically) large ranges, which is of independent interest.
cs.CR / 38 / 2609.34996
Cyclotomic Cosets: Hidden Subgroup and Quantum Sieving Algorithm for Prime-Power Moduli
Mathias Boucher, Pierre-Alain Fouque, Yixin Shen
quant-ph · cs.CR
Abstract
The Learning With Errors (LWE) problem is a fundamental assumption in post-quantum cryptography. Regev established a quantum reduction from LWE to the Dihedral Coset Problem (DCP). Later, Brakerski et al. introduced the Extrapolated Dihedral Coset Problem (EDCP), proving its equivalence to LWE. However, unlike DCP, EDCP no longer admits a coset structure. This limits the direct application of techniques for hidden subgroup problems. In this work, we introduce the Cyclotomic Coset Problem (CCP), a cyclotomic generalization of DCP that preserves an exact hidden-subgroup structure. Let $ζ_p$ be a primitive $p$-th root of unity, let $π=ζ_p-1$, and write $q=p^t$ and $L=t(p-1)$. We work over $R_q=\mathbb Z_q[ζ_p] \cong \mathbb Z[ζ_p]/(π^L)$, where the isomorphism follows from the total ramification identity $(p)=(π)^{p-1}$. We exploit the resulting $π$-adic ideal chain to construct a quantum sieve that successively reduces phase states modulo $π^{L},π^{L-1},\ldots,π$. For every fixed prime $p$ and modulus $q=p^t$, our algorithm solves the CCP in time and sample complexity $2^{O_p(\log n\log q)}$, using polynomial quantum space. The sieve also applies to uniform EDCP and Gaussian S|LWE>, yielding quasi-polynomial time algorithms for all the above problems when $q=\text{poly}(n)$. This extends the power-of-two EDCP sieve of Bai et al. (CRYPTO 2025) to a cyclotomic setting. However, we emphasize that our result does not, by itself, yield a quasi-polynomial-time algorithm for standard LWE, because the currently known reduction produces only a limited number of approximate CCP states.
cs.CR / 39 / 2609.35339
Zero Knowledge Proofs in Quantum Networks
Tuhin Paul, Srijani Das, Manasi Patra, Ramij Rahaman
quant-ph · cs.CR · cs.SI · math-ph
Abstract
Zero-knowledge proofs (ZKPs) enable the verification of a statement without revealing any information beyond its validity and constitute a fundamental primitive in cryptography and information theory. However, existing constructions rely on computational assumptions and are predominantly confined to bipartite settings, leaving their information-theoretic realization in bipartite or network scenarios largely unexplored. Here we develop a framework for zero-knowledge verification based on the indistinguishability of quantum states under operational constraints. Exploiting the fundamental limitations imposed by local operations, we show that a verifier is inherently restricted from extracting information about the underlying state while retaining the ability to verify correctness. We construct explicit protocols for multiparty quantum networks that achieve information-theoretic security, ensuring that no subset of collaborating parties can gain knowledge beyond the validity of the statement, independent of their joint computational power.
cs.CR / 40 / 2609.35633
Succinct Arguments for QMA from Collapsing Hash Functions
James Bartusek, Giulio Malavolta
quant-ph · cs.CR
Abstract
We prove the existence of succinct arguments for QMA, assuming only the existence of collapsing hash functions. This is the first scheme that relies only on unstructured ``Minicrypt'' assumptions, which are not known to imply public-key encryption. Our main technical contribution is a quantum-succinct \emph{claw-state generation} protocol that allows us to bootstrap a small number of quantum correlations into an arbitrarily large number of claw-state correlations, using classical communication only. This improves upon the work of [Zhang, STOC 2021], having better round complexity, a proof in the standard model, and being overall much simpler. This yields a quantum-succinct blind delegation of quantum computation protocol from one-way functions, which we plug into the communication-compression compiler of [Bartusek, Liu, and Malavolta, EUROCRYPT 2026] to obtain succinct arguments for QMA.