← Back to Index
Daily Research Digest

arXiv Papers

2026-10-05
529
Papers
9
Categories
128
Translated
收藏清单 0
精选 · Favorites
128
cs.AI / 1 / 2610.02330
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
行动前选择:面向长时程工具使用智能体的比较价值估计
Yu Li, Zheng Zhang, Xin Liu, Shengtian Yang, Guangfeng Cai, Lei Feng
cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.
Chinese Translation
大语言模型(LLMs)依赖长时程工具调用序列来完成复杂任务,其中每次调用都可能改变任务状态并影响后续决策。在长时程工具使用中,最终结果奖励在长交互轨迹上提供的信用分配较弱。步骤级奖励可以提供更有针对性的反馈,但获得可靠的步骤监督通常需要人工或 LLM 判断,或额外的 rollout 来估计中间决策的下游影响。在本文中,我们认为有效的工具使用智能体应在执行下一个可能的工具调用之前,估计其长时程价值。这一目标需要对同一上下文下的备选调用进行比较监督,而记录的轨迹只包含实际采取的调用。因此,我们提出用于工具使用智能体的比较推理(Comparative Inference for Tool-use Agents,CITA)。CITA 从配对信号中训练一个比较推理模型(Comparative Inference Model,CIM),这些信号结合了观察到的工具行为、来自贝叶斯工具图模拟器的可扩展监督,以及来自基于 LLM 的比较的语义判断。训练得到的 CIM 学习估计在当前上下文下,一个可能的下一个工具调用支持最终任务成功的可能性有多大。在三个工具使用基准和多个骨干 LLM 上,CITA 持续提升 Tool F1 和任务成功率。额外分析表明,CIM 学习到了针对比较性工具选择的准确步骤级价值估计。
cs.AI / 2 / 2610.02372
Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion
在文本到图像扩散中穿越满足度-多样性前沿
Kevin Zhai, Siva Rajesh Kasa, Soumya Roy, Sumit Negi, Mubarak Shah
cs.AI
diffusion
扩散模型相关
Abstract
Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user's preferences, measured by a learned reward, and differ visually from the others to maintain diversity. Existing methods are limited: they either address reward and diversity separately or combine them in one aggregate score, enabling high diversity to offset low rewards. In this paper, we address these limitations by formulating generation as satisficing: every image (candidate) must satisfy a reward floor and the batch of images must satisfy a diversity cutoff. The reward floor controls the balance between worst-candidate reward and batch diversity; we show that varying this floor defines a Pareto frontier. To traverse this frontier, we introduce SatisDive, a training-free inference-time method. SatisDive uses a batch-relative reward cutoff to distinguish lower- from higher-reward candidates, emphasizing reward improvement for candidates below the cutoff and diversity among candidates above it. On Pick-a-Pic, at matched DreamSim, SatisDive improves worst-candidate reward over FK steering by up to 0.43 with FLUX.1-dev as the base model and HPSv3 as the reward, and by up to 0.70 with SANA-1.6B as the base model and ImageReward as the reward. More broadly, across their overlapping DreamSim ranges, SatisDive's satisfaction-diversity curve Pareto-dominates FK steering's curve in each setting.
Chinese Translation
文本到图像生成使用户能够探索由同一提示词生成的若干图像。为了使这些生成的图像有用,每一张图像都必须反映用户的偏好(由学习到的奖励度量),并且在视觉上与其他图像不同以保持多样性。现有方法存在局限:它们要么分别处理奖励与多样性,要么将其合并为一个聚合分数,从而使高多样性能够抵消低奖励。在本文中,我们通过将生成形式化为满足化来应对这些局限:每张图像(候选)必须满足一个奖励下限,且整批图像必须满足一个多样性阈值。奖励下限控制最差候选奖励与批次多样性之间的平衡;我们表明,改变该下限会定义出一条帕累托前沿。为了穿越该前沿,我们提出 SatisDive,一种无需训练的推理时方法。SatisDive 使用批次相对的奖励截断值来区分较低奖励与较高奖励的候选,强调对低于截断值的候选进行奖励提升,并在高于截断值的候选之间保持多样性。在 Pick-a-Pic 上,在 DreamSim 匹配的情况下,当以 FLUX.1-dev 作为基础模型、HPSv3 作为奖励时,SatisDive 将最差候选奖励相较 FK steering 提升最多 0.43;当以 SANA-1.6B 作为基础模型、ImageReward 作为奖励时,提升最多 0.70。更广泛地说,在二者重叠的 DreamSim 范围内,SatisDive 的满足度-多样性曲线在每种设置下都帕累托支配 FK steering 的曲线。
cs.AI / 3 / 2610.02478
Tropical Reinforcement Learning
热带强化学习
Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac
cs.AI
large language model
大语言模型相关
Abstract
Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
Chinese Translation
大语言模型的强化学习通常最大化期望回报,即将所有成功轨迹的概率相加。然而,经典的求和表述只能报告模型策略成功的频率,而无法说明究竟哪个解法真正奏效;并且由于概率之和为一,强化一个解法可能会让模型遗忘另一个从未被证明错误的解法。这使得期望回报并不适合组合式推理,因为在这种推理中,一个解法必须由模型在彼此分离、且往往失败的尝试中产生、却很少同时产生的推理步骤组装而成。为解决这一问题,我们提出热带强化学习,它基于一个简单的代数改变:不再将各个备选解法的概率相加,而是取它们的最大值,由此得到热带半环。此时,一个状态的价值就变为其最可能且已获验证的解法的对数概率,并附带一条可重放、可复用的显式路径。这实现了真正的组合,因为即使最佳前缀与最佳后缀来自不同的 rollout,只要它们在某个共享状态处相接,就可以被拼接起来。为将其付诸实践,我们提出 TROPIC,一种面向具有可验证结果的确定性、可重置环境的训练算法。在四个智能体任务(Sokoban、Countdown、FrozenLake、WebShop)上,TROPIC 比最强的同策略基线高出最多 16 个百分点。因此,改变强化学习的代数,而不仅仅是其估计器,可以显著提升语言模型中的组合式推理能力
cs.AI / 4 / 2610.02491
What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
一个 token 的成本是什么?对充分每 token 计算量的混合智能体测量
Zhixu Du, Weijia Han, Hai Helen Li, Yiran Chen
cs.AI
large language model
大语言模型相关
Abstract
Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92--95\% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10\% account for 64--80\% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6\% fewer draft tokens and approximately 20\% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.
Chinese Translation
大型语言模型在生成每个 token 时都会花费相同量的计算,而不管每个 token 的生成难度如何。诸如推测解码和模型路由之类的方法建立在这样一个前提之上:这些计算中有很多是不必要的,然而,单个 token 实际需要的计算量尚未被测量。我们通过混合智能体(MoA)的视角来测量它:一个由来自三个家族的十五个容量递增的语言模型组成的小组,其中每个智能体都尝试在正确的前序 token 条件下,逐 token 地复现一个参考序列。我们将成功的最小智能体的推理成本定义为该 token 的充分计算量,这给出了该 token 所需计算量的上界。在三个核心基准上,一个 0.5B 智能体复现了 92--95\% 的参考 token。在 Qwen、OLMo 和 R1 蒸馏小组中,最昂贵的 10\% 占据了估计 FLOPs 的 64--80\%。在所有 500 个 MATH-500 问题上,相对于最佳置信度路由基线,MoA 导出的映射帮助模型路由将预计延迟从 7.59 秒降低到 5.12 秒,同时略微提高了准确率。在相似准确率下,MoA 映射帮助起草比固定窗口起草少使用 32.6\% 的草稿 token,并且预计延迟低约 20\%。这些比较揭示了仍存在的分配余量,从而激励控制器利用充分计算量结构。
cs.AI / 5 / 2610.02510
On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor
面向商科教育的本地部署多课程 RAG 辅导:校园 AI 导师中的硬件—软件权衡
Sidney Shapiro, Joshua Lindemann
cs.AI · cs.IR
large language model
大语言模型相关
Abstract
Campus AI tutors based on retrieval-augmented generation (RAG) must ground answers in assigned course materials while keeping textbooks and student dialogue on institutional infrastructure. We present CourseChat, an on-premises, multi-course RAG tutor for undergraduate business education, deployed behind a campus web gateway and intended for use embedded in Moodle. Six isolated course offerings, each keyed by its own course reference number (CRN), share twin-edge AI hosts running a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. We report two generation-model bake-off rounds, a separate fixed-evidence source-fidelity comparison, and conversation and quiz audits. Several larger models failed the classroom speed gate, but a 12B model and a 7B alternative passed. A separate mixture-of-experts candidate improved some corrections while introducing new factual and continuity errors. We therefore retain the 8B production model pending a demonstrated overall improvement, rather than claiming that 8B is universally optimal. Software changes improved follow-up topic resolution while preserving course scope; 435 prebuilt questions across 65 modules decouple practice from live generation. The results support treating model choice, evidence selection, serving compatibility, and product design as a joint engineering decision. They do not establish learning gains: faculty ratings, peak-load capacity, and complete public-gateway acceptance remain separate evaluation needs.
Chinese Translation
基于检索增强生成(RAG)的校园 AI 导师必须将答案锚定在指定课程材料之中,同时把教材与学生对话保留在院校自有基础设施之上。我们提出 CourseChat,一个面向本科商科教育的本地部署、多课程 RAG 导师,它部署在校园 Web 网关之后,并计划以嵌入 Moodle 的方式使用。六个相互隔离的课程开设实例,各自以其课程参考编号(CRN)为键,共享双边缘 AI 主机,这些主机运行一个 FastAPI 服务、一个本地向量数据库,以及一个由 Ollama 提供服务的本地大语言模型(LLM)。我们报告两轮生成模型比选(bake-off)、一项独立的固定证据来源保真度比较,以及对话与测验审计。若干更大的模型未能通过课堂速度门槛,但一个 12B 模型和一个 7B 备选模型通过了。另一个独立的混合专家候选模型改善了一些纠正,同时引入了新的事实性错误与连贯性错误。因此,在出现经证实的整体改进之前,我们保留 8B 生产模型,而不是声称 8B 是普遍最优的。软件改动改善了后续话题的解析,同时保持了课程范围;横跨 65 个模块的 435 道预置题目将练习与实时生成解耦。这些结果支持将模型选择、证据选择、服务兼容性与产品设计视为一项联合工程决策。它们并未确立学习收益:教师评分、峰值负载容量以及完整的公共网关验收仍是彼此独立的评估需求。
cs.AI / 6 / 2610.02523
Hypothesis-guided discovery of cognitive algorithms via program refinement
假设引导的通过程序精化发现认知算法
Huiwen Alex Yang, Mark K. Ho, Bill D. Thompson
cs.AI
large language model
大语言模型相关
Abstract
Developing cognitive models of algorithmic reasoning from behavioral data is a central problem in cognitive science that challenges current methods. Traditional approaches to cognitive modeling are interpretable and benefit from human expertise, but lack flexibility and scalability. Emerging techniques using large language models (LLMs) for de novo generation of cognitive models are scalable and flexible, but lack a role for human expertise and have mostly been applied to simpler tasks than algorithm recovery. We propose a hybrid system that treats discovery of cognitive algorithms as a program refinement problem. Human-created cognitive models are expressed as probabilistic programs and provided to a system of LLM agents with a mandate to: identify mismatches between model and behavior; propose code-level modifications within researcher-specified constraints; and verify structural fidelity. Revisions propagate to a probabilistic inference module that performs inference for latent variables and data likelihood computations. We evaluate the pipeline on human behavior in a problem-solving paradigm that exposes a variety of cognitive algorithms. Revised models consistently improve model fit relative to ancestral models and reveal a small set of recurring innovations that capture meaningful behavioral variability in this task.
Chinese Translation
从行为数据中构建算法推理的认知模型是认知科学中的一个核心问题,它对当前方法构成了挑战。传统的认知建模方法具有可解释性,并受益于人类的专业知识,但缺乏灵活性和可扩展性。使用大语言模型(LLM)从头生成认知模型的新兴技术具有可扩展性和灵活性,但缺乏人类专业知识的作用,且大多被应用于比算法还原更简单的任务。我们提出一个混合系统,将认知算法的发现视为一个程序精化问题。由人类创建的认知模型被表示为概率程序,并提供给一个由LLM智能体组成的系统,其任务是:识别模型与行为之间的不匹配;在研究者指定的约束内提出代码层面的修改;以及验证结构保真度。修改会传播到一个概率推断模块,该模块执行潜变量的推断和数据似然计算。我们在一个揭示多种认知算法的问题解决范式中,基于人类行为对该流程进行评估。相较于其前身模型,修订后的模型一致地改善了模型拟合,并揭示出一小组反复出现的创新,这些创新捕捉到了该任务中具有意义的行为变异性。
cs.AI / 7 / 2610.02622
CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges
CuBEs:文化情境化行为评估与文化盲大语言模型评判者的局限性
Hoda Ayad, Tanu Mitra, Abhishek Mukherji
cs.AI
large language model
大语言模型相关
Abstract
Evaluating the occurrence and triggers of large language model (LLM) behaviors - such as sycophancy, self-preference, or over-confidence - is critical for predicting real-world model deployment risks. However, existing situated behavioral evaluations typically ignore cultural context, limiting their generalizability across an increasingly global user base. To address this gap, we propose CuBEs - Culturally-situated Behavior Evaluations that probe for response patterns across diverse user cultures. We first extend an automated testing pipeline to inject cultural context into behavioral test scenarios and subsequent evaluation. We assess the cultural adaptability of this pipeline by building a human-labeled dataset that captures nuanced dimensions of behavior understanding across 12 distinct cultures. Our dataset reveals significant cross-cultural variations that one-size-fits all judgments fail to capture. Through evaluating 13 open- and closed-source LLMs, we find that introducing cultural situatedness in the evaluation scenario creates significant variation in the presence of a behavior. For example, while our baseline experiments testing for political bias capture localized Western political dimensions like the American conservative-progressive divide, non-Western culturally situated evaluations surface entirely different axes of bias such as religious and colonial political issues. Our findings demonstrate that standard, culturally-agnostic evaluations fail to capture these shifts, highlighting the necessity of culturally situated behavioral testing for global deployments.
Chinese Translation
评估大语言模型(LLM)行为——例如谄媚、自我偏好或过度自信——的发生和触发因素,对于预测真实世界模型部署风险至关重要。然而,现有的情境化行为评估通常忽略文化语境,限制了它们在日益全球化的用户群体中的泛化能力。为弥补这一空白,我们提出 CuBEs——文化情境化行为评估(Culturally-situated Behavior Evaluations),用于探查不同用户文化中的响应模式。我们首先扩展了一个自动化测试流水线,以将文化语境注入行为测试场景及后续评估中。我们通过构建一个人工标注数据集来评估该流水线的文化适应性,该数据集捕捉了跨 12 种不同文化的行为理解的细微维度。我们的数据集揭示了“一刀切”式判断无法捕捉的显著跨文化差异。通过评估 13 个开源和闭源 LLM,我们发现,在评估场景中引入文化情境性会使某一行为是否存在产生显著变化。例如,虽然我们测试政治偏见的基线实验捕捉到诸如美国保守派—进步派分歧等局部化的西方政治维度,但非西方的文化情境化评估却浮现出完全不同的偏见轴线,例如宗教和殖民政治议题。我们的发现表明,标准的、文化不可知的评估无法捕捉这些变化,凸显了针对全球部署开展文化情境化行为测试的必要性。
cs.AI / 8 / 2610.02684
Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves
随着患者证据演变,大型语言模型表现出不可靠的临床判断更新
Min Zeng, Rui Zhang
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.
Chinese Translation
大型语言模型(LLMs)在临床推理中日益受到探索,但它们是否会在患者证据演变时适当地修正判断仍不清楚。我们使用来自电子健康记录的匹配重症监护轨迹评估了纵向信念更新。在不同的大型语言模型中,当估计值发生变化时,以先前的判断为条件更常增加而非减少预测误差,这一结果在第二个终点上得到复现。受控干预揭示了两种失效模式。首先,在先前评估固定的情况下,模型对恶化呼吸证据的反应比对匹配的改善呼吸证据更强;这种不对称性在中等和强证据水平下经过余量归一化后仍然存在。其次,在当前证据固定的情况下,将先验风险从 10% 提高到 90% 使估计值偏移了 26.2 个百分点,表明先验模型信念具有因果影响。提示未能恢复可靠的更新。证据验证的纵向更新(EVLU)识别出更少、更可靠的修订,揭示了可靠性—覆盖度权衡。这些发现确立了纵向信念更新是大型语言模型可靠性的一个独立维度。
cs.AI / 9 / 2610.02687
Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning
将记忆与上下文解耦:面向令牌高效测试时持续学习的结构化记忆
Yehya Farhat, Michael Desmond, Anastasios Kyrillidis
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by improving model behavior through instructions, strategies, and evidence supplied at inference time. However, adapting context online typically requires a costly trial-and-error process, while queries are often processed independently, preventing useful experience from carrying forward. Memory systems address this limitation by retaining information across interactions, but approaches that continually append information to a shared context face increasing token costs, context-window limits, and performance degradation as the context expands. We introduce a unified formulation of context optimization and show that an agent memory system update can be interpreted as an optimization update procedure over the model's context. This perspective attempts to provide a principled framework for studying memory design and its efficiency. We then propose GraphMemory, a lightweight graph-based memory that accumulates, refines, organizes, and connects reusable strategies. For each query, GraphMemory retrieves only the relevant subgraph, enabling online context adaptation without exposing the model to the entire memory. Under bounded retrieval, the amount of retrieved memory remains constant as the number of processed examples grows. Experiments show that GraphMemory achieves competitive downstream performance while using approximately 81-85% fewer memory-construction tokens than our baselines.
Chinese Translation
大语言模型(LLMs)正越来越多地部署在企业、科学和医学应用中,在这些应用中,智能体必须纳入领域特定知识并从经验中适应。上下文工程通过推理时提供的指令、策略和证据来改进模型行为,为权重更新提供了一种实用的替代方案。然而,在线适配上下文通常需要代价高昂的试错过程,而查询往往被独立处理,使有用的经验无法延续。记忆系统通过在交互间保留信息来解决这一局限,但持续将信息追加到共享上下文的方法会随着上下文扩展而面临不断增长的令牌成本、上下文窗口限制和性能退化。我们引入上下文优化的统一形式化表述,并表明智能体记忆系统更新可以被解释为对模型上下文的一种优化更新过程。这一视角试图为研究记忆设计及其效率提供一个有原则的框架。我们随后提出 GraphMemory,一种轻量级的基于图的记忆,它累积、精炼、组织和连接可复用的策略。对于每个查询,GraphMemory 仅检索相关子图,从而能够进行在线上下文适配,而无需让模型暴露于整个记忆。在有界检索下,随着已处理样本数量的增长,检索到的记忆量保持恒定。实验表明,GraphMemory 在实现具有竞争力的下游性能的同时,相比我们的基线少使用约 81-85% 的记忆构建令牌。
cs.AI / 10 / 2610.02703
Learning to Revise Reasoning with Segment-wise On-Policy Distillation
通过分段式同策略蒸馏学习修正推理
Yuxiang Zhang, Ding Cao, Shuting Cui, Lei Wang, Weijieying Ren, Tianxiang Zhao
cs.AI
large language model
大语言模型相关
Abstract
On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student's step could be revised to improve subsequent reasoning. Furthermore, this paradigm can become less effective when the student produces a degenerate reasoning prefix, as subsequent teacher supervision remains conditioned on that prefix and may reinforce poor reasoning patterns. In this work, we focus on learning reasoning revision with segment-wise OPD to rework intermediate reasoning steps and better support subsequent reasoning. Through controlled reasoning interventions, we find that replacing student segments with teacher redrafts improves subsequent reasoning accuracy. Therefore, we address the problem of turning teacher redrafts into explicit supervision for learning to revise reasoning. We propose Segment-wise On-Policy Distillation (Seg-OPD), which selects student segments based on an uncertainty metric and obtains corresponding teacher redrafts. Seg-OPD trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise OPD supervision. Extensive experiments on mathematical reasoning and competitive programming tasks show that Seg-OPD-trained students achieve higher revision success rates than baselines. Seg-OPD consistently outperforms the compared state-of-the-art baselines in reasoning accuracy with an average relative improvement of 5.22% across diverse models and tasks. Code is available at https://anonymous.4open.science/r/Seg-OPD.
Chinese Translation
同策略蒸馏(OPD)通过在学生模型自身的采样轨迹上训练学生模型,并利用来自教师模型的密集 token 级监督,来改进大型语言模型的推理。然而,token 级 OPD 并没有明确提供一个连贯的替代推理步骤,来展示学生的步骤如何能够被修正以改进后续推理。此外,当学生产生退化的推理前缀时,这种范式可能变得不那么有效,因为后续的教师监督仍然以该前缀为条件,并可能强化不良推理模式。在这项工作中,我们关注利用分段式 OPD 学习推理修正,以重做中间推理步骤,并更好地支持后续推理。通过受控的推理干预,我们发现用教师重写稿替换学生片段能够提高后续推理准确率。因此,我们解决将教师重写稿转化为用于学习修正推理的显式监督这一问题。我们提出分段式同策略蒸馏(Seg-OPD),它基于不确定性度量选择学生片段,并获取相应的教师重写稿。Seg-OPD 训练学生模型相较于其配对的学生片段更偏好教师重写稿,同时保留密集的 token 级 OPD 监督。在数学推理和竞赛编程任务上的大量实验表明,经过 Seg-OPD 训练的学生模型比基线取得更高的修正成功率。Seg-OPD 在推理准确率上始终优于所比较的最先进基线,在不同模型和任务上平均相对提升 5.22%。代码可在 https://anonymous.4open.science/r/Seg-OPD 获取。
cs.AI / 11 / 2610.02792
Law And Order: Tax Law Autoformalization
法律与秩序:税法自动形式化
Sophia Simeng Han, Yoshiki Takashima, Anjiang Wei, Zhaoyu Li, Michael Genesereth
cs.AI
large language model
大语言模型相关
Abstract
Legal systems are increasingly implemented through software, yet scalable methods for translating legal texts into accurate symbolic representations remain underdeveloped. We study this problem through tax law, where forms and filing instructions define large computational structures involving arithmetic, branching, recursion, and tabular reasoning. We propose Law&Order, a neuro-symbolic framework for automatically formalizing tax forms and instructions into executable symbolic programs. Our approach establishes two forms of correspondence between law and logic: structural correspondence, which aligns legal and symbolic components such as cells and schedules, and denotational correspondence, which requires symbolic components to implement the computations specified by their legal counterparts. We combine large language model synthesis with cell-level verification and iterative localized error repair using human-written OpenTaxSolver tax returns. We then evaluate the resulting formalizations on independently authored, held-out TaxCalcBench returns, that are never exposed during generation or repair. Although the most advanced LLM achieves only 66% accuracy, Law&Order achieves 100% cell-level and form-level accuracy on 51 held-out returns, demonstrating the effectiveness of combining LLM-based synthesis with symbolic verification for scalable and verifiable large-scale legal autoformalization compared with using an LLM alone.
Chinese Translation
法律系统正越来越多地通过软件实现,然而将法律文本转化为准确符号表示的可扩展方法仍然不发达。我们通过税法研究这一问题,其中税务表单和申报说明定义了涉及算术、分支、递归和表格推理的大型计算结构。我们提出 Law&Order,一个神经符号框架,用于自动将税务表单和说明形式化为可执行的符号程序。我们的方法在法律与逻辑之间建立两种对应关系:结构对应,它将诸如单元格和附表之类的法律组件与符号组件对齐;以及指称对应,它要求符号组件实现其法律对应物所指定的计算。我们结合大语言模型合成与单元格级验证,以及使用人工编写的 OpenTaxSolver 纳税申报表进行迭代式局部错误修复。然后,我们在独立编写、留出的 TaxCalcBench 申报表上评估所得形式化结果,这些申报表在生成或修复期间从未被暴露。尽管最先进的 LLM 仅达到 66% 的准确率,Law&Order 在 51 份留出申报表上达到 100% 的单元格级和表单级准确率,证明与单独使用 LLM 相比,将基于 LLM 的合成与符号验证相结合,对于可扩展且可验证的大规模法律自动形式化是有效的。
cs.AI / 12 / 2610.02827
MLCommons Jailbreak Benchmark v1.0
MLCommons 越狱基准 v1.0
Carsten Maple, Cagatay Yucel, Isaac Holeman, Chris Knotz, Peter Mattson, James Goel, Jonathan Petit, Sean McGregor, James Ezick, Abhishek Kumar, Alicia Parrish, Murali Emani, Kashyap Iyer, Faiza Khan Khattak, Washington Mbonu, Daniel Machlab, Eileen Long, Shaona Ghosh, Jibin Varghese, Roman Lutz, Andrew Gruen, Bennett Hillenbrand, Prabal Gupta, Mohammed Serrhini, Dhivya Nagasubramanian, Aakash Gupta, Jun, Lu, Kurt Bollacker, Chang Liu, Jonathan Petit, Cong Chen, Jean-Philippe Monteuuis, Brent Miller, Apurv Verma, Roman Eng, Armstrong Foundjem, Mohammed Serrhini
cs.AI
large language model
大语言模型相关
Abstract
Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally refuse to provide. The MLCommons Jailbreak Benchmark v1.0 provides an end-to-end methodology for evaluating the robustness of large language models to single-turn, text-based jailbreak attacks. It combines criteria-driven system and attack selection, paired baseline and adversarial evaluation, human annotation, automated evaluator calibration, scoring, grading, and risk-calibrated disclosure within a single benchmarking pipeline. The benchmark evaluates eight open-weight systems using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy. Responses are assessed using the AILuminate Assessment Standard v1.4, and robustness is measured through the Resilience Gap: the change in safety performance between baseline and adversarial conditions. Across all evaluated systems and attacks, the unsafe-response rate increased from 11.08% under baseline conditions to 18.65% under jailbreak conditions, producing an average Resilience Gap of 7.57%. Accessible systems showed a larger mean gap, while attack effectiveness varied substantially across attack categories and hazards. The benchmark also examines evaluator reliability and sources of measurement error. Beyond reporting results, Jailbreak Benchmark v1.0 establishes a reproducible methodological foundation for comparative jailbreak evaluation and for future expansion across systems, attacks, hazards, and evaluation methods.
Chinese Translation
现代 AI 系统被设计为拒绝有害请求。越狱(jailbreak)是一种精心构造的提示,旨在绕过这些防护措施并诱导系统产出其通常会拒绝提供的输出。MLCommons 越狱基准 v1.0 提供了一套端到端方法,用于评估大语言模型对单轮、基于文本的越狱攻击的鲁棒性。它将准则驱动的系统与攻击选择、配对的基线评估与对抗评估、人工标注、自动化评估器校准、评分、分级以及风险校准的披露整合在单一的基准测试流程之中。该基准使用 264 个种子提示对八个开放权重系统进行了评估,这些提示涵盖十一个危害类别,并采用了取自 MLCommons 越狱分类体系(Jailbreak Taxonomy)的代表性攻击。回复使用 AILuminate 评估标准 v1.4 进行评估,鲁棒性则通过韧性差距(Resilience Gap)来衡量:即基线条件与对抗条件之间安全性能的变化。在所有被评估的系统与攻击中,不安全回复率从基线条件下的 11.08% 上升至越狱条件下的 18.65%,产生了 7.57% 的平均韧性差距。可访问系统表现出更大的平均差距,而攻击有效性在不同攻击类别与危害之间差异显著。该基准还考察了评估器的可靠性以及测量误差的来源。除报告结果之外,越狱基准 v1.0 还为比较性越狱评估以及未来在系统、攻击、危害和评估方法上的扩展建立了可复现的方法学基础。
cs.AI / 13 / 2610.02844
DNAlign: Dynamic Null-Space Safe Alignment for LLMs
DNAlign:面向 LLMs 的动态零空间安全对齐
Jisheng Dang, Yushuo Zhao, Dewei Liu, Junfeng Fang, Bimei Wang, Tiantian Rao, Hong Peng, Bin Hu, Tat-Seng Chua
cs.AI
large language model
大语言模型相关
Abstract
Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a persistent trade-off between safety and utility. We propose DNAlign, a lightweight alignment framework that integrates control-theoretic optimization with null-space projection. By treating the LLM as a dynamic system, the proposed framework introduces controllable perturbations to steer generation toward safe behavior. A key component is the projection module, which restricts these perturbations to the harmful-related subspace derived from neutral hidden states, thereby preserving general knowledge and response quality. A value function trained on human preference data adaptively optimizes the control signals to align with human safety preferences. Extensive evaluations across multiple LLM backbones demonstrate that our framework consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility. It achieves superior overall performance compared to prior alignment baselines without sacrificing generation diversity. These results indicate that the proposed framework provides an effective and practically deployable solution for safe LLM alignment. Code is available at https://anonymous.4open.science/r/DNAlign.
Chinese Translation
确保大语言模型(LLMs)安全可靠的部署仍然是一项根本性挑战。现有的安全对齐方法要么带来高昂的计算成本,要么无意中破坏模型的核心知识,导致其在良性任务上的流畅性和事实准确性下降。这揭示了安全性与效用性之间长期存在的权衡。我们提出 DNAlign,一个将控制论优化与零空间投影相结合的轻量级对齐框架。通过将 LLM 视为动态系统,所提出的框架引入可控扰动,以引导生成朝向安全行为。一个关键组件是投影模块,它将这些扰动限制在由中性隐藏状态导出的有害相关子空间中,从而保持通用知识和响应质量。一个基于人类偏好数据训练的价值函数自适应地优化控制信号,以与人类安全偏好对齐。在多个 LLM 主干上的广泛评估表明,我们的框架在保持流畅性、连贯性和事实效用的同时,持续减少有害输出。与先前的对齐基线相比,它在不牺牲生成多样性的情况下实现了更优的整体性能。这些结果表明,所提出的框架为安全的 LLM 对齐提供了一种有效且实际可部署的解决方案。代码可在 https://anonymous.4open.science/r/DNAlign 获取。
cs.AI / 14 / 2610.02867
TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control
TACD:通过终端放大控制蒸馏高效文本到动作模型
Wei-Jin Huang, Yuan-Ming Li, Kun-Yu Lin, Wang Luo, Yinlin Zhu, Yue Yu, Shenghao Ye, Junbin Yuan, Fa-Ting Hong, Qing Zhang, Wei-Shi Zheng
cs.AI
diffusion
扩散模型相关
Abstract
Recent text-to-motion models have improved motion quality and instruction following, yet many-step denoising and large model components make deployment slow and memory-intensive. We present Terminal-Amplification-Controlled Distillation (TACD), an on-policy approach for training efficient motion generators from text prompts and pretrained teachers, without real-motion training data. Building on segmented on-policy flow distillation, we supervise clean-motion predictions along student-generated trajectories. We identify a failure mode in which velocity matching on a fixed supervision grid repeatedly overweights errors near the denoising endpoint, degrading few-step generation. TACD ties the latest teacher query to the student's step size, bounding the effective loss weights in clean-motion space without changing inference. Experiments on HumanML3D and KIT-ML demonstrate improved few-step generation, including a 58% reduction in eight-step HY-Motion student FID relative to distillation without this bound. For diffusion teachers, the endpoint-matching form of TACD yields four-step students with lower FID and matched or improved text-motion retrieval relative to their 50-step teachers on HumanML3D. On HY-Motion and Kimodo, eight-step students with compact components achieve 7.7-11.9x end-to-end speedups and reduce peak GPU memory by 3.8-6.7x relative to their teachers. Project page: https://vkgo.github.io/TACD/
Chinese Translation
近期的文本到动作模型提升了动作质量和指令遵循能力,然而多步去噪和大型模型组件使部署变得缓慢且内存密集。我们提出终端放大控制蒸馏(Terminal-Amplification-Controlled Distillation, TACD),一种同策略方法,用于仅从文本提示和预训练教师模型训练高效动作生成器,而无需真实动作训练数据。基于分段同策略流蒸馏,我们沿学生生成的轨迹监督干净动作预测。我们识别出一种失败模式:在固定监督网格上进行速度匹配会反复过度加权去噪终点附近的误差,从而降低少步生成性能。TACD 将最新的教师查询与学生的步长绑定,在干净动作空间中有界地限制有效损失权重,而不改变推理。在 HumanML3D 和 KIT-ML 上的实验表明少步生成得到改进,包括相对于没有该界限的蒸馏,八步 HY-Motion 学生模型的 FID 降低了 58%。对于扩散教师模型,TACD 的终点匹配形式在 HumanML3D 上产生的四步学生模型具有更低的 FID,并且相对于其 50 步教师模型具有持平或提升的文本-动作检索性能。在 HY-Motion 和 Kimodo 上,具有紧凑组件的八步学生模型实现了 7.7-11.9 倍的端到端加速,并且相对于其教师模型将峰值 GPU 内存降低了 3.8-6.7 倍。项目主页:https://vkgo.github.io/TACD/
cs.AI / 15 / 2610.02885
PsyEvo: A Personalized Counseling Agent That Self-Evolves at Test Time
PsyEvo:一种在测试时自演化的个性化心理咨询智能体
Yuting Yan, Shihao Xu, Junhao Yu, Mingcong Zuo, Lu Chen, Nan Xiang, Haiyang Geng, Dongjie Tao, Minghao Wang
cs.AI
large language model
大语言模型相关
Abstract
Mental health disorders affect a substantial proportion of the global population, yet a persistent shortage of trained practitioners leaves the majority without adequate care. Large language model (LLM)-based counselors present a promising direction for delivering scalable conversational psychological support. Offline model training alone leaves limited room to adapt to individual clients or to learn from ongoing therapeutic interaction at test time. We introduce PsyEvo, an LLM-based counseling framework that enables both client-specific personalization and response-policy improvement at test time through three components: Hierarchical Bayesian Skill Policy (HBSP) personalizes what intervention to apply by maintaining a per-client skill posterior updated from session feedback; Inter-session Listwise Preference Optimization (LiPO) improves how the selected skill is expressed by updating a shared response adapter from cross-client preference evidence; and State-conditioned Ordinal Credit Assignment (SOCA) supplies candidate preferences and trajectory credit to the two components through consistency-checked comparisons and ordinal projection. In simulated-client evaluation with shared online cohort adaptation, PsyEvo obtains 7.684 Overall on PsychEval and exceeds every component variant in each of three matched runs. Removing individual components lowers mean overall score by 0.138--0.171 under the shared configuration, supporting conditional contributions within the complete scaffold. Our code is available at https://github.com/Lingxi-mental-health/PsyEvo
Chinese Translation
心理健康障碍影响着全球相当大比例的人口,然而训练有素的从业者的持续短缺使大多数人无法获得充分的照护。基于大语言模型(LLM)的心理咨询师为提供可扩展的对话式心理支持呈现了一个有前景的方向。仅靠离线模型训练,所能留出的空间有限,既难以适应个体来访者,也难以在测试时从持续进行的治疗互动中学习。我们提出 PsyEvo,一个基于 LLM 的心理咨询框架,它通过三个组件在测试时同时实现针对特定来访者的个性化与响应策略的改进:分层贝叶斯技能策略(HBSP)通过维护一个由会谈反馈更新的、针对每位来访者的技能后验,来个性化“应采用何种干预”;跨会话列表式偏好优化(LiPO)通过依据跨来访者的偏好证据更新一个共享的响应适配器,来改进“所选技能如何被表达”;而状态条件化序数信用分配(SOCA)则通过经一致性检验的比较与序数投影,为这两个组件提供候选偏好与轨迹信用。在带有共享在线队列适应的模拟来访者评估中,PsyEvo 在 PsychEval 上获得 7.684 的总体得分,并在三次匹配运行中的每一次都超过了每一个组件变体。在共享配置下,移除各个单独组件会使平均总体得分下降 0.138--0.171,这支持了各组件在完整框架内具有条件性的贡献。我们的代码可在 https://github.com/Lingxi-mental-health/PsyEvo 获取。
cs.AI / 16 / 2610.02968
Reasoning with Evidence, Not Merely Rationales: Verifiable Preference Proofs for LLM-Based Recommendation
以证据而非仅以理据进行推理:面向基于LLM推荐的可验证偏好证明
Yu Hou, Nathaniel Kang, Pengkai Wang, Hua Li
cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) can infer user preferences from interaction histories and reviews, yet the rationales they generate may not reflect the information actually used for recommendation. A preference claim may be weakly supported by its selected evidence, or may have little effect on the final ranking. We refer to these two failures as the grounding-influence gap. We introduce PROVE-REC, a general framework for verifiable preference reasoning in LLM-based recommendation. Pass A converts the complete pre-target history into a compact preference proof consisting of positive and avoidance claims linked to selected evidence entries. Pass B predicts the next item using only the proof and its selected evidence, preventing the recommender from bypassing the reasoning path. To verify evidence-to-proof grounding, we compare the effect of masking selected evidence with masking a comparable control entry. To verify proof-to-recommendation influence, we remove a preference claim and measure the resulting decrease in the target item's ranking margin. A ranking-preservation objective further retains useful information from the complete history. Comprehensive experiments on wide-ranging real-world datasets demonstrate that PROVE-REC consistently outperforms strong sequential, generative, and LLM-enhanced baselines, with improvements of up to 7.45%. Controlled ablations confirm the effectiveness of the two-pass architecture and verification objectives. Moreover, PROVE-REC produces claims that are more strongly grounded in historical evidence and more influential to recommendation while preserving ranking quality.
Chinese Translation
大语言模型(LLMs)能够从交互历史和评论中推断用户偏好,但它们生成的理据可能并不反映实际用于推荐的信息。一个偏好主张可能仅由其选定的证据提供较弱支持,或者可能对最终排序几乎没有影响。我们将这两类失败称为依据—影响鸿沟。我们提出PROVE-REC,一个用于基于LLM推荐中可验证偏好推理的通用框架。Pass A将完整的目标前历史转换为一个紧凑的偏好证明,该证明由与选定证据条目相关联的正面主张和规避主张组成。Pass B仅使用该证明及其选定证据来预测下一个物品,从而防止推荐器绕过推理路径。为了验证证据到证明的依据性,我们比较掩蔽选定证据与掩蔽一个可比较的对照条目的效果。为了验证证明到推荐的影响,我们移除一个偏好主张,并测量目标物品排序边际因此产生的下降。一个排序保持目标进一步保留来自完整历史的有用信息。在广泛的真实世界数据集上的全面实验表明,PROVE-REC始终优于强大的序列、生成式和LLM增强基线,提升最高达7.45%。受控消融实验证实了两阶段架构和验证目标的有效性。此外,PROVE-REC生成的主张更有力地以历史证据为依据,并对推荐更具影响力,同时保持排序质量。
cs.AI / 17 / 2610.02972
CreateScore: Domain-Theory-Informed Bayesian Routing for LLM-Based CV Screening
CreateScore:面向基于LLM的简历筛选的领域理论指导的贝叶斯路由
Rupsa Roy
cs.AI · stat.AP
large language model
大语言模型相关
Abstract
Large language models (LLMs) can support rubric-based screening of CVs, but applying a high-capability model to every candidate and criterion is costly. We present CreateScore, a domain-theory-informed Bayesian network for criterion-level LLM routing. A hand-specified directed acyclic graph with Dirichlet-multinomial conditional probability tables converts CV evidence into posterior uncertainty; low-uncertainty decisions are resolved by a local 8B model and uncertain ones are escalated to a 120B reference model. The graph is causally motivated, but the system performs standard Bayesian conditioning, not causal inference. The escalation threshold is calibrated on a training fold (target: 70% resolved locally) and then frozen. On 200 synthetic Data Science CVs (139 training and 61 test candidates, five criteria), 77.7% of criterion decisions were resolved locally (237 of 305). Relative to a reference condition in which the 120B model adjudicated every criterion, routed escalation reduced token use by 65.2% and raised exact score agreement from 32.8% (8B alone) to 42.6% (95% CI 31.0-55.1%); at n = 61 the gain was not statistically distinguishable. The uncertainty signal did not, however, identify the decisions on which the 8B model erred: disagreement with the reference was 16.2% among escalated and 19.4% among locally resolved decisions (AUROC 0.47, 95% CI 0.39-0.56), no better than random selection. We also document how an earlier evaluation was invalidated when truncated reasoning-model outputs were silently replaced by local labels, and we recommend safeguards for cascade evaluation. CreateScore is supported as an auditable cost-reduction mechanism, not yet as a targeted error detector, and is not an autonomous hiring system.
Chinese Translation
大语言模型(LLMs)能够支持基于评分标准的简历筛选,但将高能力模型应用于每一位候选人和每一项标准成本高昂。我们提出CreateScore,一个领域理论指导的贝叶斯网络,用于标准级别的LLM路由。一个手工指定的有向无环图,带有Dirichlet-多项条件概率表,将简历证据转换为后验不确定性;低不确定性决策由本地8B模型解决,不确定的决策则上报至120B参考模型。该图具有因果动机,但系统执行的是标准贝叶斯条件化,而不是因果推断。上报阈值在一个训练折上进行校准(目标:70%在本地解决),然后被冻结。在200份合成数据科学简历(139名训练候选人和61名测试候选人,五项标准)上,77.7%的标准决策在本地解决(305项中的237项)。相对于由120B模型裁决每一项标准的参考条件,路由式上报将token使用量减少了65.2%,并将精确分数一致率从32.8%(仅8B)提高到42.6%(95% CI 31.0-55.1%);在n = 61时,这一增益在统计上无法区分。然而,不确定性信号并未识别出8B模型判断错误的那些决策:与参考模型的不一致率在上报决策中为16.2%,在本地解决决策中为19.4%(AUROC 0.47,95% CI 0.39-0.56),并不优于随机选择。我们还记录了当被截断的推理模型输出被本地标签静默替换时,早先的一项评估如何失效,并建议为级联评估设置保障措施。CreateScore被支持作为一种可审计的成本削减机制,而尚不是一种有针对性的错误检测器,并且不是自主招聘系统。
cs.AI / 18 / 2610.02975
Reliable Self-Evolution with Imperfect Proxy Rewards
使用不完美代理奖励的可靠自进化
Kangjun Noh, Soyu Kim, Kyungwoo Song
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM)-based self-evolving search is a promising approach to scientific discovery. However, high-fidelity evaluation of every candidate is prohibitively expensive in some domains. Self-evolving systems in such settings therefore rely on low-cost but imperfect proxy rewards, which may assign high scores to infeasible candidates. These false positives may contaminate both the final output and the feedback used to guide subsequent generations. This motivates statistically calibrated reward intervals for more reliable self-evolving search. We propose Conformal Interval-Driven Self-Evolution (CISE), which constructs candidate-specific reward intervals using conditional conformal inference and iteration-wise online density-ratio estimation. CISE uses conservative interval-based rewards for evolutionary feedback and returns candidates only when all required property intervals lie entirely within their respective feasible regions. We derive fixed-iteration coverage results under explicit assumptions of independence and covariate shift. We evaluate CISE on three self-evolving search tasks in materials science. In our experiments, all candidates returned by CISE are true positives under high-fidelity evaluation, whereas the baselines return more candidates but include false positives. These results highlight the value of a smaller, more precise shortlist when downstream validation budgets are limited. Our repository is available at https://github.com/MLAI-Yonsei/CISE.git.
Chinese Translation
基于大语言模型(LLM)的自进化搜索是一种有前景的科学发现方法。然而,在某些领域中,对每个候选进行高保真评估的成本高得令人望而却步。因此,此类场景中的自进化系统依赖低成本但不完美的代理奖励,而这些奖励可能会给不可行的候选赋予高分。这些假阳性可能同时污染最终输出以及用于指导后续生成的反馈。这促使我们采用经过统计校准的奖励区间,以实现更可靠的自进化搜索。我们提出共形区间驱动的自进化(Conformal Interval-Driven Self-Evolution, CISE),它使用条件共形推断和逐迭代在线密度比估计来构建针对具体候选的奖励区间。CISE 使用保守的基于区间的奖励作为进化反馈,并且仅当所有所需性质区间都完全位于各自可行区域内时才返回候选。我们在显式的独立性和协变量偏移假设下推导了固定迭代的覆盖性结果。我们在材料科学中的三个自进化搜索任务上评估了 CISE。在我们的实验中,CISE 返回的所有候选在高保真评估下都是真阳性,而基线返回更多候选,但包含假阳性。这些结果凸显了当下游验证预算有限时,规模更小但更精确的候选短名单的价值。我们的代码仓库位于 https://github.com/MLAI-Yonsei/CISE.git。
cs.AI / 19 / 2610.02976
Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation
面向音视频幻觉缓解的相关证据解码
Hyunjae Ra, Aecheon Jung, Jungin Park, Sungeun Hong
cs.AI
large language model
大语言模型相关
Abstract
Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction. Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informative modality. For example, when asked which instrument is heard, a model may correctly predict violin from the audio alone. Once a video showing a guitar is added, its confidence in violin may drop. In this paper, we introduce Relevant Evidence Decoding (RED), a training-free method that identifies question-relevant evidence and selectively strengthens its contribution. RED uses pointwise mutual information to quantify the predictive support provided by audio and video beyond the question alone. It decomposes their joint contribution into audio, video, and residual interaction components. A question-only inference pass determines the required evidence type, after which the model augments the original audio-visual prediction with the corresponding PMI contribution. Across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves accuracy over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.
Chinese Translation
音视频大语言模型(AV-LLMs)仍然容易产生跨模态幻觉,即一种模态会错误地影响关于另一种模态的预测。尽管对比解码减少了视觉-语言模型中的幻觉,但将其直接扩展到 AV-LLMs 忽略了一个关键挑战:不同问题需要不同的感知证据,包括音频、视频或它们的交互。值得注意的是,我们观察到,即使模型能够从单一有信息量的模态中恢复出正确答案,联合音视频推理也可能削弱预测。例如,当被问及听到的是哪种乐器时,模型可能仅根据音频就正确预测出小提琴。一旦加入一段显示吉他的视频,其对小提琴的置信度可能会下降。在本文中,我们提出相关证据解码(Relevant Evidence Decoding,RED),这是一种无需训练的方法,它识别与问题相关的证据并有选择地增强其贡献。RED 使用逐点互信息来量化音频和视频在仅有问题之外所提供的预测支持。它将其联合贡献分解为音频、视频和残差交互分量。仅问题的推理过程确定所需的证据类型,之后模型用相应的 PMI 贡献增强原始音视频预测。在三个音视频幻觉基准和三个 AV-LLMs 上,RED 相比标准解码提高了准确率:在 CMM 上最高提升 7%,在 AVHBench 上提升 6.3%,在 SVHalluc 上提升 3.8%,平均相对首 token 时间为标准解码的 1.5 倍。
cs.AI / 20 / 2610.02982
PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation
PLCWorld:在闭环工厂仿真中对LLM生成的PLC程序进行基准测试
Yunji Kim, Yunseok Lee, Hyunwoo Seo, Jaerim Choi, Woojin Lee
cs.AI
large language model
大语言模型相关
Abstract
Programmable logic controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. Evaluating whether large language model (LLM)-generated PLC programs satisfy task requirements and safety constraints requires observing how their commands affect device and workpiece states. We introduce PLCWorld, a common closed-loop execution environment and benchmark that couples Structured Text (ST) execution with simulated plant responses and sensor feedback. Grounded in control relations identified in industrial PLC programs and engineering documentation, PLCWorld contains 100 synthetic tasks and 473 registered task-condition pairs across Motion Control and Material Handling, with difficulty defined by control-dependency scope. A common protocol reports Task Success and Safety Violation separately. Validation combines practitioner review, reference and alternative programs, targeted counterexamples, specification-evaluator alignment checks, and comparisons with independent ST runtimes. Reference and alternative programs satisfy their applicable cases, while all 542 targeted counterexamples activate their designated evaluator rules under at least one registered condition. Execution Gap relates submission-profile acceptance to subsequent task failure or observed Safety Violation. Across the constructed task groups, direct GPT-5.5 achieves 82.70% Task Success on Easy cases but 25.10% on Hard cases. Evaluations of six LLMs and four adapted generation-and-verification workflows further expose differences between completion, safety, and generation cost. Our code, simulation environment, benchmark tasks, and baseline implementations are publicly available at https://yunji0516.github.io/PLCWorld/.
Chinese Translation
可编程逻辑控制器(PLC)通过读取传感器输入并发出控制命令来协调工业设备。评估大语言模型(LLM)生成的PLC程序是否满足任务要求与安全约束,需要观察其命令如何影响设备与工件状态。我们提出PLCWorld,一个通用的闭环执行环境与基准,它将结构化文本(ST)执行与仿真的工厂响应及传感器反馈耦合起来。PLCWorld立足于从工业PLC程序和工程文档中识别出的控制关系,包含跨运动控制与物料搬运的100个合成任务和473个已注册的任务-条件对,其难度由控制依赖范围定义。一个通用协议分别报告任务成功与安全违规。验证结合了从业者评审、参考程序与替代程序、针对性反例、规范-评估器一致性检查,以及与独立ST运行时的比较。参考程序与替代程序均满足其适用的案例,而全部542个针对性反例都在至少一个已注册条件下激活了其指定的评估器规则。执行差距将提交画像的接受情况与随后的任务失败或观测到的安全违规相关联。在构建出的各任务组中,直接使用GPT-5.5在简单案例上达到82.70%的任务成功率,但在困难案例上仅为25.10%。对六个LLM和四种经适配的生成-验证工作流的评估,进一步揭示了在完成度、安全性与生成成本之间的差异。我们的代码、仿真环境、基准任务和基线实现已公开发布于 https://yunji0516.github.io/PLCWorld/。
cs.AI / 21 / 2610.03029
SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation
SoftGene:面向可解释基因集注释的蛋白质语言模型增强软提示
Drew Ross, Arya Hadizadeh Moghaddam, Dongjie Wang, Xiaoyu Zhang, Zijun Yao
cs.AI
large language model
大语言模型相关
Abstract
Gene set analysis is a cornerstone of functional genomics, yet it remains labor-intensive and heavily dependent on manual curation and expert biological interpretation. While Large Language Models (LLMs) have emerged as powerful tools for genomic reasoning and annotation, most existing approaches rely on symbolic gene names and fail to capture domain-specific biological structure, particularly protein sequence information that governs molecular activity, interactions, and downstream gene function. In this work, we propose SoftGene, a novel framework for LLM-based gene set annotation that leverages the hierarchical structure of gene sets. First, we use a hierarchical attention-based encoder built on ESM, a protein language model, to represent each gene set using protein-level amino acid sequence information. Second, we construct a hybrid prompting scheme that combines soft prompts derived from gene set embeddings with hard prompts containing auxiliary context generated by an LLM, and feed the resulting prompt into a local LLM for annotation. We evaluate our framework on two benchmark datasets: Gene Ontology (GO) and the Molecular Signatures Database (MSigDB). Our results show that integrating protein-sequence representations with textual context improves gene set annotation overall, while per-domain analyses reveal that the contribution of protein embeddings varies across biological domains.
Chinese Translation
基因集分析是功能基因组学的基石,但它仍然劳动密集,并且严重依赖人工整理和专家生物学解释。尽管大语言模型(LLMs)已成为基因组推理和注释的强大工具,但大多数现有方法依赖符号化的基因名称,未能捕捉领域特异性的生物结构,尤其是决定分子活性、相互作用和下游基因功能的蛋白质序列信息。在这项工作中,我们提出 SoftGene,一个用于基于 LLM 的基因集注释的新型框架,它利用了基因集的层次结构。首先,我们使用一个建立在 ESM(一种蛋白质语言模型)之上的基于层次化注意力的编码器,利用蛋白质层面的氨基酸序列信息来表示每个基因集。其次,我们构建了一种混合提示方案,它将由基因集嵌入得到的软提示与包含由 LLM 生成的辅助上下文的硬提示相结合,并将得到的提示输入本地 LLM 进行注释。我们在两个基准数据集上评估我们的框架:基因本体(GO)和分子特征数据库(MSigDB)。我们的结果表明,将蛋白质序列表示与文本上下文相结合总体上改善了基因集注释,而分域分析揭示,蛋白质嵌入的贡献在不同生物学领域之间有所差异。
cs.AI / 22 / 2610.03033
When Numbers Start Talking: Numerical Signalling and Strategic Behaviour Among LLMs
当数字开始说话:LLM 之间的数值信号传递与策略行为
Alessio Buscemi, Daniele Proverbio, Alessandro Di Stefano, The Anh Han, German Castignani, Pietro Liò
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM)-based agents increasingly operate in multi-agent systems (MAS) characterised by strategic interaction. However, little is known about whether, and to what extent, different types of messages affect the outcomes of strategic games. By investigating AI agents based on four popular LLMs, playing four games with different cooperation equilibria, we study whether messages of different kinds (natural language, numerical signals, or random sequences) significantly modify the levels of cooperation in each game, also depending on the agents' assigned personalities. We observe that structured messages alter the final payoffs for most games and LLMs, but without a predictable pattern; this challenges the assumption that AI agents can converge to stable equilibria regardless of additional capabilities. Moreover, we observe that agent-generated numerical messages depart from randomness, most strongly and consistently when agents are explicitly instructed to communicate; however, they introduce an additional interpretability challenge, as their symbol distributions are mostly associated with the payoff structure and typically become more concentrated with repetition, but are overall difficult for humans to interpret. Monitoring for coordination of AI agents through restricted channels should thus prioritise message-level fingerprints, which generalise across models, over behavioural decisions, which do not.
Chinese Translation
基于大语言模型(LLM)的智能体日益在多智能体系统(MAS)中运行,这些系统以策略互动为特征。然而,关于不同类型的消息是否以及在多大程度上影响策略博弈的结果,人们知之甚少。通过研究基于四种流行 LLM 的 AI 智能体,并让它们玩四种具有不同合作均衡的博弈,我们研究不同类型的消息(自然语言、数值信号或随机序列)是否显著改变每个博弈中的合作水平,同时这也取决于智能体被分配的人格。我们观察到,结构化消息会改变大多数博弈和 LLM 的最终收益,但不存在可预测的模式;这挑战了如下假设:AI 智能体无论是否具备额外能力都能收敛到稳定均衡。此外,我们观察到,智能体生成的数值消息偏离随机性,当智能体被明确指示进行交流时,这种偏离最为强烈且最为一致;然而,它们引入了额外的可解释性挑战,因为其符号分布大多与收益结构相关,并且通常会随着重复而变得更加集中,但总体上难以供人类解释。因此,通过受限渠道监测 AI 智能体的协调时,应优先考虑能够跨模型泛化的消息级指纹,而非不能跨模型泛化的行为决策。
cs.AI / 23 / 2610.03056
MOF-VERIFY: A Failure-Aware Agentic Harness for MOF Hypothesis Verification
MOF-VERIFY:一种用于MOF假设验证的失败感知智能体化框架
Donghyun Lee, Taehoon Lee, Geonhee Ahn, Jieun Kim, Jihyun Park, Suyeon Cho, Yoona Kim, Chaerim Shin, Hoi Ri Moon, Jonggeol Na, Sukho Hong, Jihwan Oh, Soo Kyung Kim
cs.AI · cs.CE
large language model
大语言模型相关
Abstract
Large language models are increasingly used as reasoning components in AI-driven materials Co-Scientists, yet the reliability of the resulting verification pipeline remains unclear. Metal-organic frameworks (MOFs) provide a particularly challenging setting because structures may appear under different identifiers, synthesis outcomes depend strongly on experimental conditions, evidence is distributed across heterogeneous sources, and some hypotheses require computation rather than literature alone. We introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification. T-MOF-1-3 are evaluated under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, while T-MOF-4 separately evaluates computational verification. Guided by these diagnosed failure modes, we develop MOF-Verify, a failure-aware agentic harness that targets structural, literature, evidence-sufficiency, and computational bottlenecks before producing a final verdict. Across multiple backbone LLMs, MOF-Verify substantially improves hypothesis-verification performance over direct inference and retrieval-based baselines. Benchmark datasets are released at https://github.com/IMMS-Ewha/MOF-Verify-Benchmark.
Chinese Translation
大语言模型正越来越多地被用作AI驱动的材料协同科学家中的推理组件,然而由此产生的验证流程的可靠性仍不清楚。金属有机框架(MOFs)提供了一个尤其具有挑战性的场景,因为结构可能以不同标识符出现,合成结果强烈依赖于实验条件,证据分布在异构来源中,而且一些假设需要计算而非仅靠文献。我们引入了一个诊断基准,包含四个任务族,涵盖结构锚定、合成条件验证、证据充分性验证以及基于MLIP的计算验证。T-MOF-1-3在闭卷、启用检索和oracle证据设置下进行评估,以定位知识访问、证据获取和推理中的失败,而T-MOF-4单独评估计算验证。在这些被诊断出的失败模式的指导下,我们开发了MOF-Verify,一种失败感知的智能体化框架,它在产生最终判定之前针对结构、文献、证据充分性和计算瓶颈进行处理。在多个骨干LLM上,MOF-Verify相较于直接推理和基于检索的基线,显著提升了假设验证性能。基准数据集发布在 https://github.com/IMMS-Ewha/MOF-Verify-Benchmark。
cs.AI / 24 / 2610.03316
Multi-Task Evolution for Zero-Shot Cross-Problem Generalization using LLMs
使用LLMs的零样本跨问题泛化的多任务进化
Zhouliang Xie, Changliang Zhou, Genghui Li, Zhenkun Wang
cs.AI
large language model
大语言模型相关
Abstract
Designing effective heuristics for diverse combinatorial optimization problems requires substantial expertise and repeated search. Large language models (LLMs) automate heuristic generation and refinement, but heuristic search typically depends on evaluation feedback from the problem being optimized. Generalizing to new problem definitions using only source-task feedback therefore remains a central challenge. We introduce MECo, an LLM-driven multi-task evolutionary framework for zero-shot cross-problem generalization. MECo maintains task-conditioned heuristic populations and uses a transfer gap based on cross-task population performance to guide their interactions. These interactions enable the transfer and recombination of heuristics. A complementary selection criterion then constructs a compact heuristic set by rewarding each member's additional coverage of source combinations. The selected set is applied to target problems without further search or adaptation. Experiments on 32 problem variants across vehicle routing (VRP) and flexible job-shop scheduling (FJSP) show that MECo achieves the lowest mean costs compared with eight automated heuristic design (AHD) baselines under the same budgets. On out-of-domain problems, it outperforms the strongest baseline in each family. Moreover, integrating the framework of MECo with different AHD methods improves their ID and OOD performance in both families, supporting its effectiveness across different methods.
Chinese Translation
为多样化的组合优化问题设计有效的启发式方法需要大量专业知识和反复搜索。大型语言模型(LLMs)自动化了启发式方法的生成与改进,但启发式搜索通常依赖于来自被优化问题的评估反馈。因此,仅使用源任务反馈泛化到新的问题定义仍然是一个核心挑战。我们提出 MECo,一个由 LLM 驱动的多任务进化框架,用于零样本跨问题泛化。MECo 维护以任务为条件的启发式种群,并使用基于跨任务种群性能的迁移差距来指导它们之间的交互。这些交互使得启发式方法能够迁移和重组。随后,一种互补的选择准则通过奖励每个成员对源组合的额外覆盖来构建一个紧凑的启发式集合。所选集合被应用于目标问题,无需进一步搜索或适应。在车辆路径(VRP)和柔性作业车间调度(FJSP)的32个问题变体上的实验表明,在相同预算下,与八个自动化启发式设计(AHD)基线相比,MECo 达到了最低的平均成本。在域外问题上,它在每个问题族中都优于最强基线。此外,将 MECo 的框架与不同的 AHD 方法集成,提高了它们在两个问题族中的 ID 和 OOD 性能,支持了其在不同方法上的有效性。
cs.AI / 25 / 2610.03320
Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS
细化换来可懂度,搜索换来身份:测试时计算在掩码扩散 TTS 中换来了什么
Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh
cs.AI
diffusion
扩散模型相关
Abstract
Diffusion language models for text-to-speech combine two forms of computation: model depth (parameters) and refinement steps (inference budget). We ask whether they scale equally across capabilities. We train 15 masked-diffusion codec TTS models varying depth (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweep refinement steps T in [1,16] at inference, measuring zero-shot synthesis via ASR word error rate (intelligibility) and speaker verification (identity) on 174 held-out speakers. Against measured floors, refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range - a 1.86x asymmetry robust across multiple error metrics. Retraining at 3x and 6x schedule attenuates but does not reverse this gap (1.84 to 1.36 to 1.23x), because intelligibility saturates with steps while identity continues improving. Best-of-K search recovers speaker identity where refinement fails, with 64.6-79.0% win rates across four independent encoders. Depth and steps are not interchangeable: separable B(d)B(T) fits significantly better (Delta AICc=+69.3) than substitution models. Analysis shows 62% of remaining identity deficit lies in the codec, not the generator. We conclude that refinement and depth target different bottlenecks and should be optimized separately.
Chinese Translation
用于文本到语音的扩散语言模型结合了两种形式的计算:模型深度(参数)与细化步数(推理预算)。我们追问它们是否在各种能力上以同等程度扩展。我们在 2,000 小时语音上训练了 15 个深度各异的掩码扩散编解码器 TTS 模型(19–133M 参数,3 个随机种子),并在推理时扫描细化步数 T 于 [1,16] 区间,通过在 174 个留出说话人上以 ASR 词错误率(可懂度)和说话人验证(身份)衡量零样本合成效果。对照实测下限,细化弥合了 86.2% 的可懂度区间,但仅弥合了 46.4% 的身份区间——一种在多种误差度量下均稳健的 1.86 倍不对称性。以 3 倍和 6 倍调度重新训练会减弱但不会逆转这一差距(1.84 到 1.36 再到 1.23 倍),因为可懂度随步数增加而饱和,而身份仍在持续改善。Best-of-K 搜索在细化失败之处恢复了说话人身份,在四个独立编码器上取得 64.6–79.0% 的胜率。深度与步数不可互换:可分离的 $B(d)B(T)$ 拟合显著优于替代模型(Delta AICc=+69.3)。分析表明,剩余身份缺口的 62% 位于编解码器,而非生成器。我们得出结论:细化与深度针对的是不同的瓶颈,应当分别优化。
cs.AI / 26 / 2610.03326
Preserving Mathematical Reasoning in Compressed Diffusion Language Models via Trajectory-Aware Low-Rank Approximation
在压缩扩散语言模型中通过轨迹感知低秩近似保持数学推理
Tian Liang, Zishan Shao, Yiran Chen
cs.AI
diffusion
扩散模型相关
Abstract
Diffusion language model (dLLM) compression faces a known challenge because calibration is typically performed on clean, fully visible activations, whereas inference traverses partially masked intermediate states. For low-rank compression, this raises two questions. First, can low-rank optimality still be characterized when approximation quality is measured over trajectory-distributed states, and second, does the choice of calibration states affect mathematical reasoning preservation under compression? We address these questions by formulating a trajectory-aware low-rank objective over corruption levels and masking realizations. To estimate this objective efficiently, we propose Traj-MC, which estimates the trajectory second moment through Monte Carlo sampling and yields exact sampled-state optimality and population consistency. Under matched compression budgets, trajectory-aware calibration improves reconstruction over the generation trajectory and preserves substantially more mathematical reasoning than clean calibration on mathematical reasoning benchmarks. Our results connect trajectory-aware low-rank optimality to the reasoning capability retained after dLLM compression. Our code is available at: https://github.com/Zishan-Shao/traj-mc.git.
Chinese Translation
扩散语言模型(dLLM)压缩面临一个已知挑战,因为校准通常在干净、完全可见的激活上进行,而推理则遍历部分掩码的中间状态。对于低秩压缩,这引出了两个问题。第一,当近似质量在轨迹分布状态上被衡量时,低秩最优性是否仍可被刻画?第二,校准状态的选择是否影响压缩下的数学推理保持?我们通过构造一个在损坏水平和掩码实现上的轨迹感知低秩目标来回答这些问题。为了高效估计该目标,我们提出 Traj-MC,它通过蒙特卡洛采样估计轨迹二阶矩,并给出精确的采样状态最优性和总体一致性。在匹配的压缩预算下,轨迹感知校准改善了沿生成轨迹的重建,并且在数学推理基准上比干净校准保持显著更多的数学推理。我们的结果将轨迹感知低秩最优性与 dLLM 压缩后保留的推理能力联系起来。我们的代码可在:https://github.com/Zishan-Shao/traj-mc.git 获取。
cs.AI / 27 / 2610.03356
ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models
ReFract:使用文本世界模型对语言模型智能体的视角意识进行基准测试
Hainiu Xu, Vítor N. Lourenço, Mohnish Dubey, Yunfei Bai, Yulan He, Caroline Catmur, Aline Paes, Marco Caserta, Akash Chandrayan, Luca D'Angelo
cs.AI
large language model
大语言模型相关
Abstract
Large Language Model (LLM) agents are increasingly deployed in high-stakes settings such as industrial maintenance and equipment fault troubleshooting, where workers occupy a variety of roles. A capable agent must therefore act in a way that is calibrated to user's role: taking actions and providing information that respect the role's knowledge and capability boundaries. Unlike coding, where mistakes are usually recoverable, agent responses in these settings are enacted on physical equipment, and can therefore cause irreversible equipment damage, production loss, or personnel harm. Existing benchmarks, however, largely overlook the need for agents to infer what a role intends and acting only through tools that role may legitimately use, a capability which we term Perspective Awareness. To this end, we introduce ReFract, a benchmark of 150 expert-validated entries in which an agent must act differently in response to the same query depending on user's role. Entries of ReFract are grounded in anonymized queries from domain support conversations, against which we construct Text World Models that simulate the agent's operating environments and assemble perspective-aware action trajectories. State-of-the-art LLMs solve at most 69% of the tasks with more than 50% of their trajectories contain attempts of taking perspective-violating actions. ReFract exposes perspective awareness as a distinct, largely unsolved axis of agent evaluation and motivates agents that calibrate not just how to act, but for whom.
Chinese Translation
大语言模型(LLM)智能体正越来越多地部署在工业维护和设备故障排查等高风险场景中,在这些场景中,工作人员承担着各种各样的角色。因此,一个有能力的智能体必须以与用户角色相校准的方式行动:采取行动并提供尊重该角色的知识和能力边界的信息。与错误通常可恢复的编程不同,在这些场景中,智能体的响应会作用于物理设备,因此可能造成不可逆的设备损坏、生产损失或人员伤害。然而,现有基准在很大程度上忽视了智能体需要推断角色的意图,并仅通过该角色可以合法使用的工具采取行动,我们将这一能力称为视角意识。为此,我们引入 ReFract,这是一个包含 150 个经专家验证的条目的基准,其中智能体必须根据用户角色对同一查询做出不同的行动。ReFract 的条目基于来自领域支持对话的匿名化查询,我们针对这些查询构建了模拟智能体操作环境的文本世界模型,并组装了具有视角意识的行动轨迹。最先进的 LLM 最多只能解决 69% 的任务,且其超过 50% 的轨迹包含采取违反视角行动的尝试。ReFract 将视角意识揭示为智能体评估中一个独特的、很大程度上尚未解决的维度,并推动智能体不仅校准如何行动,还要校准为谁行动。
cs.AI / 28 / 2610.03383
CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models
CVE2AP:通过大语言模型自动生成PDDL编码的攻击路径
Lin Cui, Vincenzo Scotti, Raffaela Mirandola
cs.AI
large language model
大语言模型相关
Abstract
Attack Path (AP) modeling is fundamental to cybersecurity analysis, where the Planning Domain Definition Language (PDDL) has been widely adopted to encode APs into formal and machine-verifiable representations for automated reasoning about vulnerability exploitation, attack progression, and their potential impacts. However, existing AP modeling approaches largely rely on expert-driven manual construction, limiting their scalability and ability to keep pace with rapidly evolving cyber threats. Large language models (LLMs) are promising candidates, as their extensive pre-trained knowledge and reasoning capabilities enable them to interpret and transform threat intelligence into formal representations. In this paper, we propose \textbf{CVE2AP}, an LLM-based approach for automatically generating PDDL-encoded attack paths from natural language CVE (Common Vulnerability Exposure) descriptions. CVE2AP leverages structured prompting and incorporates an error-feedback mechanism that iteratively refines the generated paths using planner-reported syntactic and solvability errors. We conduct a systematic empirical evaluation across multiple LLMs and generation configurations, assessing generation quality across syntactic, solvability and semantic dimensions, together with token consumption and generation time. The results demonstrate that CVE2AP effectively generates high-quality PDDL-encoded attack paths, achieving up to 86.9\% syntax correctness, 78.6\% solvability, and 93.1\% semantic correctness under LLM-as-expert evaluation, while \texttt{GPT-5.5} offers the best quality-cost trade-off and error feedback yields the most consistent quality improvement.
Chinese Translation
攻击路径(AP)建模是网络安全分析的基础,其中规划领域定义语言(PDDL)已被广泛采用,用于将AP编码为形式化且机器可验证的表示,以对漏洞利用、攻击进展及其潜在影响进行自动推理。然而,现有的AP建模方法在很大程度上依赖专家驱动的手工构建,限制了其可扩展性以及跟上快速演变的网络威胁的能力。大语言模型(LLM)是有前景的候选方案,因为其广泛的预训练知识和推理能力使其能够解释威胁情报并将其转化为形式化表示。在本文中,我们提出\textbf{CVE2AP},一种基于LLM的方法,用于从自然语言CVE(通用漏洞暴露)描述中自动生成PDDL编码的攻击路径。CVE2AP利用结构化提示,并引入一种错误反馈机制,该机制使用规划器报告的语法错误和可解性错误来迭代细化生成的路径。我们在多个LLM和生成配置上进行了系统的实证评估,从语法、可解性和语义维度评估生成质量,并同时评估token消耗和生成时间。结果表明,CVE2AP有效生成高质量的PDDL编码攻击路径,在LLM作为专家评估下达到高达86.9\%的语法正确性、78.6\%的可解性和93.1\%的语义正确性,而\texttt{GPT-5.5}提供了最佳的质量-成本权衡,并且错误反馈带来了最一致的质量提升。
cs.AI / 29 / 2610.03430
Jumping the Line: Exploiting Length Predictions in LLM Scheduling
插队:利用 LLM 调度中的长度预测
Yuyang Dai, Rana Shahout, Mahmood Sharif
cs.AI
large language model
大语言模型相关
Abstract
Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as Shortest Job First prioritize shorter requests, but output lengths are unknown before generation, so practical schedulers rely on predicted lengths. We introduce JIL, an attack on prediction-based LLM schedulers that manipulates the scheduling signal to obtain higher priority and reduce completion time. Using TRAIL as a case study, JIL optimizes an adversarial suffix that causes a lightweight output-length probe to underestimate a request's length. We evaluate JIL on two datasets and four LLMs across varied request profiles and deployment configurations. JIL reduces predicted output lengths by up to 83.4 percent, and adversarial requests complete up to 1.53 times faster on average in end-to-end serving experiments. The reduction in predicted length is substantially larger than the change in actual output length, revealing a mismatch between the scheduler's estimate and the request's realized size. Response utility varies across models and tasks, exposing a trade-off between scheduling advantage and response quality. We also evaluate scheduler-side defenses and find that grouping length predictions into coarse intervals reduces JIL's scheduling advantage and mitigates delays to benign requests.
Chinese Translation
高效请求调度对于减少大语言模型(LLM)服务中的完成时间日益重要。诸如最短作业优先(Shortest Job First)之类的基于大小的策略会优先处理较短的请求,但在生成之前输出长度是未知的,因此实际调度器依赖于预测长度。我们提出 JIL,这是一种针对基于预测的 LLM 调度器的攻击,它操纵调度信号以获得更高优先级并减少完成时间。以 TRAIL 为案例研究,JIL 优化一个对抗性后缀,使轻量级输出长度探针低估请求的长度。我们在两个数据集和四个 LLM 上,跨多种请求画像和部署配置评估 JIL。JIL 将预测输出长度最多降低 83.4%,并且在端到端服务实验中,对抗性请求平均完成速度最高快 1.53 倍。预测长度的减少显著大于实际输出长度的变化,揭示了调度器的估计与请求实际实现大小之间的不匹配。响应效用因模型和任务而异,暴露出调度优势与响应质量之间的权衡。我们还评估了调度器侧防御,并发现将长度预测分组为粗粒度区间会降低 JIL 的调度优势,并缓解对良性请求的延迟。
cs.AI / 30 / 2610.03509
Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
高效推理训练并不总是损害 CoT 忠实度与可监控性
Samuel Lewis-Lim, Xingwei Tan, Mario Sanger, Zhixue Zhao, Nikolaos Aletras
cs.AI
large language model
大语言模型相关
Abstract
Chain-of-thought (CoT) reasoning allows humans to inspect how large language models reach their answers, and oversee model behaviour. This reasoning comes at an increased inference cost, motivating efficient methods that train models to solve tasks using fewer tokens. However, a common concern is that such training may cause models to skip important reasoning steps, so the CoT no longer faithfully reflects the model's decision. It is unclear whether or when this occurs in practice, since different efficiency methods apply length pressure to models' CoT in distinct ways, and faithfully explaining a model's decision takes more tokens on some tasks than others. To understand these dynamics, we fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward. We evaluate how efficient reasoning affects CoT faithfulness (i.e., how well the CoT reflects model decisions on related inputs) and monitorability (i.e., whether the CoT reveals when input interventions alter the output). We find that it affects faithfulness and monitorability differently. Faithfulness falls in most settings, primarily because the trained models are less consistent. Monitorability is more robust, as models keep acknowledging the influence on their answer even when the CoT is much shorter.
Chinese Translation
思维链(CoT)推理使人类能够检查大语言模型如何得出其答案,并监督模型行为。这种推理带来了更高的推理成本,从而推动了训练模型使用更少 token 解决任务的高效方法。然而,一个常见的担忧是,此类训练可能导致模型跳过重要的推理步骤,从而使 CoT 不再忠实反映模型的决策。目前尚不清楚这在实践中是否发生或何时发生,因为不同的高效方法以不同方式对模型的 CoT 施加长度压力,而且在某些任务上忠实解释模型决策比其他任务需要更多 token。为了理解这些动态,我们使用三种以不同方式施加长度压力的方法微调了多种模型,即固定生成预算、逐示例长度目标和组相对长度奖励。我们评估高效推理如何影响 CoT 忠实度(即 CoT 在相关输入上反映模型决策的程度)和可监控性(即 CoT 是否揭示输入干预何时改变输出)。我们发现它对忠实度和可监控性的影响不同。在大多数设置中,忠实度下降,主要因为训练后的模型一致性更低。可监控性更为稳健,因为即使 CoT 短得多,模型仍会承认对其答案的影响。
cs.AI / 31 / 2610.03591
HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
HazardWeaver:面向灾害分析智能体的科学路线选择
Wangshu Zhu, Xueqi Cheng, Liang Wu, Yushun Dong
cs.AI
large language model
大语言模型相关
Abstract
Understanding and assessing natural hazards is essential for disaster preparedness and risk reduction. Recent advances in large language models have spurred growing interest in AI agents for hazard analysis, particularly their ability to integrate scientific data, models, and tools into automated workflows. However, effective automation requires agents to determine which scientific methods are appropriate for a given event and executable with the available data and tools. As new evidence and execution results become available, these conditions can change, requiring agents to reconsider their choices. We formulate this problem as state-dependent scientific route selection and introduce HazardWeaver. Specifically, HazardWeaver first leverages the Hazard Knowledge Compiler to extract evidence-linked conditions governing scientific applicability, then its Hazard Capability Graph represents executable scientific capabilities and checks compatibility between their inputs and outputs. Using these complementary representations, the Hazard Weaver Agent component selects applicable and executable routes, carries out their workflows, and revises its decisions as the analysis state changes. To evaluate both the scientific outputs and the decisions that produce them, we introduce the Hazard Weaver Benchmark, comprising 141 instances across seven single-hazard domains and four multi-hazard interaction classes. The benchmark accommodates multiple valid scientific routes and evaluates output correctness, route validity, and justified abstention. Extensive experiments on this benchmark show that HazardWeaver outperforms existing agent systems, with the largest gains on tasks with multiple eligible scientific routes. Our code is publicly available at https://github.com/LabRAI/HazardWeaver.
Chinese Translation
理解和评估自然灾害对于灾害防备和风险降低至关重要。大语言模型的最新进展激发了人们对用于灾害分析的 AI 智能体的日益增长的兴趣,尤其是其将科学数据、模型和工具集成到自动化工作流中的能力。然而,有效的自动化要求智能体确定哪些科学方法适用于给定事件,并且能够利用可用数据和工具执行。随着新证据和执行结果的出现,这些条件可能发生变化,要求智能体重新考虑其选择。我们将这一问题形式化为状态依赖的科学路线选择,并提出 HazardWeaver。具体而言,HazardWeaver 首先利用 Hazard Knowledge Compiler 提取与证据关联的、支配科学适用性的条件,然后其 Hazard Capability Graph 表示可执行的科学能力,并检查其输入和输出之间的兼容性。使用这些互补的表示,Hazard Weaver Agent 组件选择适用且可执行的路线,执行其工作流,并随着分析状态的变化修订其决策。为了评估科学输出以及产生这些输出的决策,我们引入 Hazard Weaver Benchmark,它包含跨七个单一灾害领域和四个多灾害交互类别的 141 个实例。该基准允许多个有效的科学路线,并评估输出正确性、路线有效性以及有正当理由的弃权。在该基准上的大量实验表明,HazardWeaver 优于现有智能体系统,在具有多个符合条件科学路线的任务上提升最大。我们的代码已公开提供,网址为 https://github.com/LabRAI/HazardWeaver。
cs.AI / 32 / 2610.03626
Depth as Time in One-Step Generative Models
一步生成模型中的深度即时间
Arnold Caleb Asiimwe, William Yang, Sanghyuk Chun, Esin Tureci, Olga Russakovsky
cs.AI
diffusion
扩散模型相关
Abstract
The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textit{depth as time}: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model's own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \texttt{SiT-L/2} model can be compressed by $16.6\times$ in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \texttt{SiT-L/2} model by $16.6\times$ in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.
Chinese Translation
最近兴起的一步生成模型,通过蒸馏或学习到的流映射来压缩扩散的多步轨迹,已经达到了一个拐点,即它们能够生成高质量图像。在此,我们提出一个由这些进展自然引出的问题:当生成被压缩为单次前向传播时,多步扩散的去噪轨迹会发生什么?我们提供一种我们称为“深度即时间”的经验观察:多步扩散在采样步骤上所执行的去噪计算,似乎沿单次前向传播的深度展开,并且可以通过使用模型自身的输出头解码中间层来恢复。最有趣的是,我们表明这种逐深度计算取决于流映射被训练来解决的传输任务。最令人惊讶的情况是 MeanFlow,其中探测更短的传输区间会在单次网络评估中同时揭示去噪和再加噪。相比之下,未使用时间索引传输任务训练的生成器,例如漂移模型,不会表现出相同的逐深度去噪。因此,我们表明表现出逐深度去噪现象的模型在逐层计算上更可压缩:一个 MeanFlow \texttt{SiT-L/2} 模型可以在参数上压缩 $16.6\times$,成为单个时间条件块。我们为这种先去噪后再加噪的行为提供解释,并表明当我们将逐层计算明确视为一个流时,可以训练单个时间条件块跨层去噪,从而将一个 MeanFlow \texttt{SiT-L/2} 模型在参数上压缩 $16.6\times$。总之,这些结果表明,扩散的时间计算并未被一步生成消除,而是跨网络深度被重新组织。
cs.AI / 33 / 2610.03639
Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
大型语言模型了解哥伦比亚法律吗?针对哥伦比亚法律体系的可靠性基准
Rubén Manrique, Michelle Castellanos, Jorge Morales, Juan David Gutiérrez, Antonio Barreto Rozo, Joaquín Vélez Navarro
cs.AI · cs.CY
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian legal system. The benchmark comprises 1,042 items spanning ten areas of law and three question formats (closed multiple-choice, semi-open, and open-ended IRAC), built through a human-in-the-loop pipeline with multi-stage expert review. We evaluate 15 contemporary proprietary and open-weight models with format-appropriate metrics. Accuracy on closed questions ranges widely, from 0.905 (Gemini 3.1 Pro) to 0.577, but on free-text legal answers factual correctness never exceeds 0.45 (on a 0-1 scale) for any model. We find a dissociation between answer relevancy and correctness (Spearman rho = -0.46): models reliably sound responsive while frequently being wrong, a pattern of particular concern for non-expert users. Closed-question accuracy and free-text correctness are strongly rank-correlated (rho = 0.94), so cheap multiple-choice screening predicts model ranking but overstates absolute reliability. An independent rubric-based LLM judge and blind human expert scoring both reproduce the free-text ranking (rho >= 0.88). The judge further reveals that only about half of the norms models cite are correct; the rest are wrong or non-existent. Reliability varies systematically by legal area and follows an inverted-U across question complexity. Our results indicate that current LLMs require expert supervision for Colombian legal tasks, and that grounding answers in authoritative sources is a promising path to higher reliability. We release the benchmark construction pipeline to support reproducible evaluation.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于支持法律实践、教育和研究,但它们在除美国以外的国家法律体系中的可靠性在很大程度上仍缺乏记录。我们提出了一个经专家验证的基准,用于评估LLM在哥伦比亚法律体系上的可靠性。该基准包含1,042个项目,涵盖十个法律领域和三种问题格式(封闭式多项选择、半开放式和开放式IRAC),并通过带有多阶段专家评审的人在回路流程构建。我们使用与格式相适应的指标评估了15个当代专有模型和开放权重模型。封闭式问题的准确率差异很大,从0.905(Gemini 3.1 Pro)到0.577不等,但在自由文本法律答案上,任何模型的事实正确性都从未超过0.45(在0-1量表上)。我们发现答案相关性与正确性之间存在分离(Spearman rho = -0.46):模型可靠地听起来有回应性,却经常出错,这一模式对非专家用户尤其令人担忧。封闭式问题的准确率与自由文本的正确性在排名上高度相关(rho = 0.94),因此廉价的多项选择筛选可以预测模型排名,但会高估绝对可靠性。一个独立的基于评分量规的LLM评判器以及盲法人类专家评分都复现了自由文本排名(rho >= 0.88)。该评判器进一步揭示,模型引用的规范中只有大约一半是正确的;其余的是错误的或不存在的。可靠性因法律领域而系统性变化,并随问题复杂度呈倒U形。我们的结果表明,当前的LLM在哥伦比亚法律任务中需要专家监督,并且将答案建立在权威来源之上是通往更高可靠性的一条有前景的路径。我们发布基准构建流程,以支持可复现的评估。
cs.CL / 34 / 2610.02444
Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs
经由逐定理符号验证器的反例生成:当模仿有害而强化修复
Omar Farouk Zouak, Houssam Eddine Boukhalfa, Soumaya Lakehal, Shiv Katiyar, Samia Nefti-Meziani
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. The collapse replicates across four seeds and on Gemma-3-4B. Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions finds 97.7% accuracy. Code, data, verifier modules and annotations: https://github.com/ce-rlvr/SymCE.
Chinese Translation
大语言模型常常能正向地证明一个定理,却无法反驳一个与之密切相关的错误命题:这是一种证伪鸿沟,监督微调无法弥合它,甚至可能主动使其恶化。我们将反例生成形式化为针对确定性的逐定理 Python 验证器进行受约束的见证输出,并发布 SymCE,一个包含 4,707 个错误的本科代数与实分析猜想的语料库,每个猜想都配有一个可执行的验证器。该验证器同时充当奖励函数,使 SymCE 成为一个训练环境。在该神谕下,用 SFT 继以 GRPO 训练 Qwen3-4B,揭示出一个模仿陷阱:仅使用反例的 SFT 将真定理识别率从 0.27 坍缩至 0.00,而采用仅含稀疏结果奖励的 RLVR 修复了这一点并超过基线,达到 0.66。这种坍缩在四个随机种子以及 Gemma-3-4B 上都可复现。稀疏奖励与稠密奖励在域内成功率上统计上无法区分,却在一个留出的校准探针上相差 33 个百分点,我们将这种分离追溯到部分得分项。我们的 4B 模型优于所评估的每一个 7B 开源权重数学专用模型,与六个前沿商业 API 相比仍具竞争力,并能在提示不变的情况下迁移到 GSM8K、MATH-500 和 MMLU-college-math。对 177 个验证器判定的人工审核发现准确率为 97.7%。代码、数据、验证器模块与标注:https://github.com/ce-rlvr/SymCE。
cs.CL / 35 / 2610.02549
Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks
评估大型语言模型在时间抽取任务中的多维泛化能力
Fahmid Shahriar Iqbal, Ritam Dutt, Soumitra Das, Arnav Verma, Sagnik Ray Choudhury
cs.CL
large language model
大语言模型相关
Abstract
Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predicts better generalization. However, this relationship weakens under substantial distribution shifts. Inductive prompting performs most consistently across domain shift, adversarial perturbations, compositionality, and length increase, while gains from scale, architecture, and deductive and abductive prompting strategies are uneven and dimension-specific. We conclude that LLM generalization in temporal extraction tasks cannot be predicted from any single dimension alone and cannot be reliably inferred from in-domain or single-dimension evaluations, highlighting the need for reasoning strategies that generalize across dimensions.
Chinese Translation
时间和事件表达抽取是基础性的时间推理任务,但由于标注歧义、领域敏感性和模型行为不稳定,该问题仍然困难。现有评估关注域内性能,对分布偏移下的可靠性提供的洞见有限。我们评估了跨模型家族、架构和推理策略的多种模型配置在四个泛化维度上的表现,考察了从基础性能的迁移、跨维度相关性,以及规模、架构和提示的影响。这为提示下的大型语言模型在时间和事件表达抽取任务中如何泛化提供了系统性研究。我们发现,强的基础任务性能通常预示着更好的泛化。然而,在显著分布偏移下,这种关系会减弱。归纳式提示在领域偏移、对抗扰动、组合性和长度增加方面表现最为一致,而来自规模、架构以及演绎和溯因提示策略的增益则不均衡且因维度而异。我们得出结论:大型语言模型在时间抽取任务中的泛化无法仅从任何单一维度预测,也无法从域内或单维度评估中可靠推断,这凸显了对能够跨维度泛化的推理策略的需求。
cs.CL / 36 / 2610.02665
Large Language Continuous Diffusion Models
大型语言连续扩散模型
Zhihan Yang, Wei Guo, Jean-Marie Lemercier, Simon Welker, Yonggan Fu, Mohammad Mahdi Kamani, Sajad Norouzi, Julius Berner, Tomas Geffner, Karsten Kreis, Yongxin Chen, Molei Tao, John Thickstun, Pavlo Molchanov, Ante Jukić, Arash Vahdat, Morteza Mardani
cs.CL · cs.AI · cs.LG
diffusion
扩散模型相关
Abstract
Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across comprehensive math reasoning and coding evaluations against state-of-the-art discrete counterparts (masked dLMs and AR baselines), Sigma achieves competitive performance with discrete models on standard benchmarks (e.g., GSM8K, Minerva, HumanEval, MBPP) after pre-training and on challenging reasoning tasks (e.g., MATH-500, AIME) after supervised fine-tuning. Beyond performance parity, we uncover key structural properties unique to continuous dLMs: (i) embedding-space steering effectively governs the quality-diversity trade-off, yielding strong pass@k performance and (ii) continuous trajectories enable graceful degradation for low NFEs and efficient distillation. These establish continuous dLMs as a promising paradigm for efficient language generation.
Chinese Translation
尽管离散扩散语言模型(dLMs)在快速并行解码方面取得了成功,但其不平滑的高维空间阻碍了用于推理与推理加速的轨迹引导。为克服这一限制,我们提出了 Sigma,这是首个基于可引导的低维 ODE/SDE 潜轨迹构建的大规模(3B/8B)连续 dLM。Sigma 通过似然优化以分块方式训练,在联合去噪被高斯噪声破坏的词元嵌入的同时,学习最优的嵌入几何。为加速训练,Sigma 利用自回归(AR)模型的预训练权重进行热启动。在推理阶段,我们发现无分类器引导与分数温度对于高保真推理和编码至关重要。在针对最先进离散对应模型(掩码 dLM 与 AR 基线)的全面数学推理和编码评测中,Sigma 在预训练后在标准基准(如 GSM8K、Minerva、HumanEval、MBPP)上,以及在监督微调后在具有挑战性的推理任务(如 MATH-500、AIME)上,均取得了与离散模型相当的性能。除了性能相当之外,我们还揭示了连续 dLM 独有的关键结构特性:(i) 嵌入空间引导能够有效调控质量—多样性权衡,从而获得强劲的 pass@k 表现;(ii) 连续轨迹使得在低 NFE 下能够平滑退化,并支持高效蒸馏。这些特性确立了连续 dLM 作为高效语言生成的一种有前景的范式。
cs.CL / 37 / 2610.02744
EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models
EpiWorld:将 LLM 政策智能体扎根于流行病学世界模型
Zeeshan Memon, Yiqi Su, Kai Shu, Naren Ramakrishnan, Liang Zhao
cs.CL
large language model
大语言模型相关
Abstract
Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quantitative surveillance signals required to assess severity, and the institutional constraints that define admissible actions. We present EpiWorld, a closed-loop framework that grounds an LLM policy actor in a learned action-conditioned epidemiological world model and a tiered skill library of public-health protocols, surveillance tools, and adaptive lessons accumulated through after-action analysis. Given a candidate intervention, the world model predicts regional epidemic evolution and enables fast counterfactual rollouts that provide feedback for policy selection and refinement. Outcomes of simulated futures are distilled into reusable lessons while protocol constraints remain fixed, allowing the decision process to improve without sacrificing interpretability or controllability. We evaluate both the world model and the end-to-end framework on retrospective COVID-19 and Influenza datasets: the world model achieves the best out-of-distribution Peak-MAE among all forecasting baselines, and the closed-loop framework reduces cumulative hospitalisation by up to 59% across datasets and by an average of ~16% across six LLM backbones, outperforming reinforcement-learning and optimal-control policy baselines.
Chinese Translation
流行病干预政策是文本产物,人类决策者通过自然语言对其进行解释、辩护和修订,这使得大语言模型成为流行病政策推理的自然候选者。然而,一个朴素的 LLM 缺乏预测干预后果所需的流行病动力学、评估严重程度所需的定量监测信号,以及界定可允许行动的机构约束。我们提出 EpiWorld,一个闭环框架,它将 LLM 政策行动者扎根于一个学习得到的、以行动为条件的流行病学世界模型,以及一个分层技能库,该技能库包含公共卫生协议、监测工具和通过事后分析积累的自适应经验教训。给定一个候选干预,世界模型预测区域流行病演化,并支持快速的反事实推演,为政策选择和精炼提供反馈。模拟未来的结果被提炼为可复用的经验教训,同时协议约束保持固定,从而使决策过程得以改进,而不牺牲可解释性或可控性。我们在回顾性 COVID-19 和流感数据集上评估世界模型和端到端框架:世界模型在所有预测基线中取得了最佳的分布外 Peak-MAE,并且闭环框架跨数据集将累计住院人数最多减少 59%,跨六个 LLM 主干平均减少约 16%,优于强化学习和最优控制政策基线。
cs.CL / 38 / 2610.02772
Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing
通过聚焦视图改进非结构化知识编辑中的原子事实召回
Ding Wu, Ye Zhang, Haoyu Wang, Tianci Liu
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly serve as general-purpose interfaces to factual knowledge, but their parameters do not automatically reflect information that changes after pretraining. Knowledge editing (KE) provides a targeted alternative to costly retraining by modifying selected knowledge and preserving unrelated knowledge and general capabilities. Conventional KE uses structured factual triples, whereas unstructured KE (UKE) uses free-form passages containing multiple facts. Nonetheless, existing UKE editors exhibit a failure mode known as context reliance: edited LLMs can often reproduce the editing passage but fail to reliably recall its individual facts without the original passage context. We identify context-induced difficulty underestimation under the standard passage-level editing objective: later facts receive increasingly rich ground-truth context and consequently incur lower initial losses, making them appear easier to learn. In response, we propose FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding (RoPE) positions assigned to the keys of its preceding context. The perturbation is applied during editing and removed afterward, leaving the model's native positional encoding unchanged at inference time. We instantiate FOVEATED for both direct-optimization and locate-then-edit editors. We theoretically analyze how FOVEATED counteracts context-induced difficulty underestimation and empirically demonstrate consistent improvements across five KE editors, two LLM backbones, and three benchmarks.
Chinese Translation
大语言模型(LLMs)日益充当事实知识的通用接口,但其参数并不会自动反映预训练之后发生变化的信息。知识编辑(KE)通过修改选定的知识并保留无关知识与通用能力,为代价高昂的重新训练提供了一种有针对性的替代方案。传统的 KE 使用结构化的事实三元组,而非结构化 KE(UKE)则使用包含多个事实的自由形式段落。然而,现有的 UKE 编辑器表现出一种被称为上下文依赖的失效模式:编辑后的 LLM 往往能够复现编辑段落,但在没有原始段落上下文时却无法可靠地召回其中的单个事实。我们识别出,在标准的段落级编辑目标下存在上下文诱发的难度低估:靠后的事实获得越来越丰富的真值上下文,因而产生更低的初始损失,使它们看起来更容易学习。为此,我们提出 FOVEATED,这是一个即插即用的框架,它通过随机偏移分配给某一句子前文上下文的键的旋转位置嵌入(RoPE)位置,来为该句子构造聚焦视图。该扰动在编辑期间施加,随后被移除,从而在推理时保持模型原生的位置编码不变。我们分别为直接优化型编辑器和先定位后编辑型编辑器实例化了 FOVEATED。我们从理论上分析了 FOVEATED 如何抵消上下文诱发的难度低估,并通过实验证明,它在五种 KE 编辑器、两种 LLM 主干模型和三个基准上均带来了一致的改进。
cs.CL / 39 / 2610.02775
Automatic Evaluation of Mental Health Stigma in Online Communication
在线交流中心理健康污名的自动评估
Naomi Baes, Jemima Kang, Nick Haslam, Chris Groot, Alsa Wu, Luc Raszewski, Yulia Otmakhova
cs.CL
large language model
大语言模型相关
Abstract
Mental health stigma has profoundly harmful impacts but its complexity makes it difficult to evaluate. Stigma may involve explicit derogation, but also subtler forms of blame, fear, paternalistic pity, social distancing, structural exclusion, and discrimination. We introduce a theory-grounded benchmark for automatic evaluation of mental health stigma in online communication, consisting of naturally occurring online news and social media text annotated with a fine-grained taxonomy of stigma across multiple mental health conditions. Our annotation framework comprises a binary stigma-detection task and a multi-level taxonomy covering (i) stigma mode, (ii) domain, and (iii) specific components of certain forms of stigma. We apply this framework to texts mentioning six mental health conditions and evaluate large language models alongside stigma-related classifiers for detecting sentiment, toxicity, and hate speech. Results show that mental health stigma is not well captured by models trained to detect these neighboring constructs, and that LLMs often overpredict stigma unless given explicit operational rules - mirroring the importance of decision rules in human annotation. We release the publicly available part of benchmark, annotations, prototypical exemplars of stigma and code at: https://github.com/jemimakang/mh_stigma.
Chinese Translation
心理健康污名具有深远的危害性影响,但其复杂性使其难以评估。污名可能涉及显性的贬损,但也可能涉及更微妙形式的责备、恐惧、家长式的怜悯、社会疏远、结构性排斥和歧视。我们引入一个以理论为基础的基准,用于在线交流中心理健康污名的自动评估,该基准由自然产生的在线新闻和社交媒体文本组成,这些文本按照跨多种心理健康状况的细粒度污名分类体系进行了标注。我们的标注框架包括一个二元污名检测任务和一个多层级分类体系,涵盖(i)污名模式,(ii)领域,以及(iii)某些污名形式的具体组成部分。我们将该框架应用于提及六种心理健康状况的文本,并评估大型语言模型(LLMs)以及用于检测情感、毒性和仇恨言论的污名相关分类器。结果表明,心理健康污名并未被那些训练用于检测这些邻近构念的模型很好地捕捉,并且除非给出明确的操作规则,LLMs 往往会过度预测污名——这反映了决策规则在人工标注中的重要性。我们发布基准中可公开获取的部分、标注、污名的原型示例和代码,地址为:https://github.com/jemimakang/mh_stigma。
cs.CL / 40 / 2610.02819
Text-Centric Post-Training for Omni-Modal Reasoning
以文本为中心的全模态推理后训练
Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, Yu Wang
cs.CL
large language model
大语言模型相关
Abstract
Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.
Chinese Translation
提升全模态大语言模型中的联合音视频推理通常会带来大量的数据构建与训练成本。我们的诊断表明,尽管模型对所有相应的单跳问题都能给出正确答案,仍存在多跳推理困难,并表明在感知与推理目标的局部优化中存在部分解耦。这促使我们开展对这些能力给予不同侧重的后训练。纯文本推理训练在不同数据来源、模型规模和模型家族上都能带来增益。在表现最佳的纯文本配置下,监督微调后再进行强化学习(RL)使 Qwen2.5-Omni-7B 在九项推理得分上的几何平均值相比基座模型提升了 25.83%,并在 GPU 小时数减少 56.6% 的情况下超过了完整的原生音视频路线。在完全由纯文本 LLM 合成的数据上进行训练,在构建或训练中均不使用音视频数据的情况下,将该几何平均值提升了 21.01%。然而,纯文本训练会损害感知能力。因此,我们提出一种以文本为中心的后训练范式:纯文本训练提供主要的推理优化,随后用减少数据的原生音视频 RL 来精炼感知。该精炼过程所使用的输入 token 数比全数据音视频 RL 少约 90%,将感知恢复到基座水平之上,并保留了表现最佳的纯文本流程推理增益的 93.5%。
cs.CL / 41 / 2610.02829
Clinical Concept Centers in LLMs
大型语言模型中的临床概念中心
Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of representation than the text: internal representations not only encode substantially more than the output verbalizes, but the stated reasoning also systematically omits features that causally drive the answer. An evaluation of model behavior in terms of mechanistic interpretability has not been explored in clinical decision support. In this work, we extend behavioral evaluation into the latent space and ask whether clinical concepts exist as locatable, causally used representations inside open-weight LLMs. We find dedicated clinical concept centers in the latent space of all eleven open models we test. These concept centers are interpretable, firing only on their aligned clinical narratives, and meaningfully and causally drive model behavior in both constrained and open-ended settings. They are not just analytical representations, but circuits that can be utilized in clinical practice, and we explore their use from the perspective of both evaluation and performance. From the evaluation standpoint, models stay internally coherent and keep using the relevant concept centers even under adversarial role-based priming, while aligned priming improves downstream clinical performance. From a performance perspective, we simulate realistic deployment settings and find that steering models along these centers leads to meaningful downstream improvements. Finally, we conduct a blinded clinician validation and find the activation and usage of these concept centers predicts clinicians preferences.
Chinese Translation
大型语言模型正越来越多地被用于临床环境。然而,对这些模型可靠性和性能的研究几乎完全集中在语言基质上,对模型所说的话进行评分。机制可解释性发现,潜空间承载着比文本更高保真度的表征:内部表征不仅编码了远多于输出所表述的内容,而且所述推理也系统性地遗漏了因果驱动答案的特征。从机制可解释性角度对模型行为进行评估尚未在临床决策支持中得到探索。在这项工作中,我们将行为评估扩展到潜空间,并探究临床概念是否作为可定位、被因果性使用的表征存在于开放权重 LLM 内部。我们在测试的全部十一个开放模型的潜空间中都发现了专门的临床概念中心。这些概念中心是可解释的,仅在其对齐的临床叙述上激活,并在受限和开放式设置中都有意义且因果性地驱动模型行为。它们不仅是分析性表征,而且是可在临床实践中利用的回路,我们从评估和性能两个视角探索了它们的用途。从评估的角度看,即使在对抗性的基于角色的启动下,模型仍保持内部连贯,并继续使用相关的概念中心,而对齐的启动则改善下游临床性能。从性能的角度看,我们模拟真实部署设置,并发现沿这些中心引导模型会带来有意义的下游改进。最后,我们进行了盲法临床医生验证,并发现这些概念中心的激活和使用能预测临床医生的偏好。
cs.CL / 42 / 2610.02856
Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models
用于大语言模型平衡多任务后训练的自适应互蒸馏
Baohang Li, Xiaocheng Feng, Yichong Huang, Chengpeng Fu, Wenshuai Huo, Zekun Zhou, Zekun Yuan, Tingjia Zhang, Bing Qin
cs.CL
large language model
大语言模型相关
Abstract
Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across tasks, then uses task-wise validation scores to select an adjustment for each task and transfer direction. Across six benchmarks and three LLM backbones, both AMD models achieve higher average benchmark scores than supervised fine-tuning (SFT) baselines trained with the same sampling strategies. They also outperform the task-balancing methods evaluated in our experiments. Merging the two trained models can further improve their average benchmark score while yielding a single model for inference. The merged models outperform multi-task SFT by an average of 2.91 points across the three backbones.
Chinese Translation
大语言模型(LLMs)的多任务后训练旨在提升在训练数据量不等的各个任务上的性能。现有方法主要关注在单模型训练期间平衡任务贡献。不同的任务平衡策略能够产生具有互补优势的模型,为相互蒸馏创造机会。然而,跨模型监督的有用性可能因任务、迁移方向和训练阶段而异。我们提出自适应相互蒸馏(AMD),一种协同后训练框架,它联合训练两个采用不同任务平衡策略的模型。AMD 通过跨任务共享的短时训练探测来评估对蒸馏权重的候选调整,然后使用逐任务验证分数为每个任务和迁移方向选择一个调整。在六个基准和三个 LLM 主干模型上,两个 AMD 模型均取得了比使用相同采样策略训练的监督微调(SFT)基线更高的平均基准分数。它们也优于我们实验中评估的任务平衡方法。合并这两个训练后的模型可以进一步提升它们的平均基准分数,同时得到一个用于推理的单一模型。在三个主干模型上,合并后的模型平均超过多任务 SFT 2.91 分。
cs.CL / 43 / 2610.02877
Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty
超越分数对齐评估LLM作为评判者:残差评判难度的心理测量学分析
Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.
Chinese Translation
大语言模型(LLM)被广泛用作自动评判者,其效度通常通过与人类评分的对齐程度来评估。然而,总体一致性无法揭示人类与LLM是否认为相同的评估案例具有难度。在本文中,我们从心理测量学的视角研究摘要评估中的这一问题。我们分别对人类评分和LLM评分拟合多面Rasch模型(Many-Facet Rasch Models),从而将分数分解为潜在摘要质量、评分者严厉度、维度严厉度和评分量表阈值。在此分解的基础上,我们将残差难度定义为一种经模型调整的评判难度度量,并比较人类评判者与LLM评判者是否共享相同的难度结构。在SummEval上针对17个开放权重LLM评判者的实验中,我们发现潜在摘要质量上的中等程度对齐并不意味着残差难度上的对齐。人类评判者与LLM评判者在哪些摘要–维度单元仍然困难这一点上存在差异,且这种不匹配强烈依赖于维度。一致性(consistency)表现出明显的“LLM更难”偏移,而连贯性(coherence)则表现出“人类更难”偏移。我们进一步表明,人类认为容易但LLM认为困难的案例可以部分地从可观测的源文本–摘要属性中预测出来。这些发现表明,总体的人类对齐只反映了LLM作为评判者可靠性的一部分,而心理测量学的残差诊断有助于实现信息更丰富的评判者评估以及更有针对性的人类–LLM协作。
cs.CL / 44 / 2610.02949
Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method
通过多编程语言指令微调与集成方法增强生物医学命名实体识别
Songtao Li, Yijia Zhang, Jianyuan Yuan, Shidi Zhang, Fengyu Zhang, Hongfei Lin
cs.CL
large language model
大语言模型相关
Abstract
Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textual outputs, providing limited structural constraints for typed entity extraction. Second, high-quality biomedical annotations are limited, and learning from a single serialized output form may restrict structural diversity and reduce model robustness. Although external biomedical knowledge can be introduced to alleviate data scarcity, it often requires costly resource construction. To address these challenges, we propose MITE, a Multiple Programming Languages Instruction Tuning and Ensemble method for BioNER. MITE reformulates BioNER as a structure-to-structure generation task by representing both instructions and entity outputs in code-formatted representations. Specifically, each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, while preserving the same underlying entity semantics. These language-specific representations provide structurally diverse supervision without requiring external biomedical knowledge or additional annotations. During inference, MITE aggregates predictions from different code formats through an entity-level voting strategy, reducing language-specific prediction variance and improving robustness. Experiments on six widely used BioNER datasets demonstrate that MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization. Ablation and parameter analyses further verify the effectiveness and robustness of the proposed components.
Chinese Translation
指令微调已成为将大语言模型(LLMs)应用于生物医学命名实体识别(BioNER)的一种常见范式。然而,现有的指令微调方法仍然面临两个关键挑战。首先,传统的自然语言指令通常将 BioNER 标注序列化为扁平的文本输出,为带类型的实体抽取提供的结构约束有限。其次,高质量的生物医学标注有限,而仅从单一序列化输出形式中学习可能会限制结构多样性并降低模型鲁棒性。尽管可以引入外部生物医学知识来缓解数据稀缺问题,但这通常需要昂贵的资源构建。为解决这些挑战,我们提出 MITE,一种用于 BioNER 的多编程语言指令微调与集成方法。MITE 通过将指令和实体输出都表示为代码格式的表示,将 BioNER 重新表述为结构到结构的生成任务。具体而言,每个训练实例都被转换为多种编程语言格式,包括 Python、C++ 和 Java,同时保留相同的底层实体语义。这些特定于语言的表示提供了结构多样的监督,而无需外部生物医学知识或额外标注。在推理过程中,MITE 通过实体级投票策略聚合来自不同代码格式的预测,从而减少特定于语言的预测方差并提高鲁棒性。在六个广泛使用的 BioNER 数据集上的实验表明,MITE 持续优于具有代表性的基于 BERT 和基于 LLM 的基线,并展现出强大的跨数据集泛化能力。消融实验和参数分析进一步验证了所提出组件的有效性和鲁棒性。
cs.CL / 45 / 2610.02970
A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition
一种面向 Schema-as-Code 生物医学命名实体识别的指南增强多智能体框架
Songtao Li, Yijia Zhang, Shidi Zhang, Jianyuan Yuan, Fengyu Zhang, Hongfei Lin
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown promising potential for biomedical named entity recognition (BioNER) through instruction following and in-context learning. However, existing LLM-based BioNER methods still face two key limitations. First, retrieved demonstrations and external biomedical knowledge provide limited support for dataset-specific annotation semantics, leaving entity boundaries, type scopes, and annotation conventions ambiguous. Second, free-form generation lacks sufficient structural control, often leading to invalid formats, hallucinated mentions, duplicated entities, and boundary errors. To address these limitations, we propose GAMA, a guideline-augmented multi-agent framework for schema-as-code BioNER. GAMA first induces candidate annotation rules from labeled training instances and verifies them against annotated data to construct reliable dataset-specific guideline memory. Guided by these verified rules, a planning component generates ranked span-type hypotheses with rationales, and a coding component converts them into schema-constrained entity objects. A verification module then checks span grounding, type validity, and structural compliance, and performs dual-loop refinement to correct invalid or low-confidence predictions. Experiments on five widely used BioNER datasets with multiple LLM backbones show that GAMA consistently outperforms strong LLM-based baselines. Ablation and parameter analyses further verify the effectiveness of the proposed components.
Chinese Translation
大语言模型(LLMs)通过指令遵循和上下文学习,已在生物医学命名实体识别(BioNER)中展现出可观潜力。然而,现有的基于 LLM 的 BioNER 方法仍面临两个关键局限。第一,检索到的示例与外部生物医学知识对数据集特定的标注语义提供的支持有限,使实体边界、类型范围与标注惯例仍然模糊不清。第二,自由形式生成缺乏足够的结构控制,常导致无效格式、幻觉提及、重复实体与边界错误。为解决这些局限,我们提出 GAMA,一个面向 schema-as-code BioNER 的指南增强多智能体框架。GAMA 首先从带标注的训练实例中归纳候选标注规则,并依据标注数据对其加以验证,以构建可靠的数据集特定指南记忆。在这些经核验的规则指导下,一个规划组件生成带有理由的排序后的跨度-类型假设,一个编码组件将其转换为受 schema 约束的实体对象。随后,一个验证模块检查跨度定位、类型有效性与结构合规性,并执行双环精炼以纠正无效或低置信度的预测。在五个广泛使用的 BioNER 数据集上、使用多个 LLM 骨干进行的实验表明,GAMA 持续优于强大的基于 LLM 的基线。消融与参数分析进一步验证了所提出各组成部分的有效性。
cs.CL / 46 / 2610.02986
OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models
OLMo-Detect:一个面向大语言模型成员推断的多阶段、混杂因素受控基准
Tao Shi, Chaoyi Xiang, Qiongkai Xu, Jey Han Lau
cs.CL
large language model
大语言模型相关
Abstract
Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer from three limitations: limited coverage of training stages, insufficient distributional alignment between members and non-members, and lack of rigorous filtering of non-members against the training corpus. To address these limitations, we propose OLMo-Detect, a multi-stage, confounder-controlled benchmark built upon the fully open OLMo 2 pipeline. OLMo-Detect spans pre-training, mid-training, and post-training, explicitly aligns members and non-members on three key axes, and rigorously filters non-members via infini-gram. To assess robustness to distribution shifts, we further introduce OLMo-Detect (Shifted), a variant where members are misaligned with non-members. We evaluate 15 unsupervised and 3 supervised membership inference attacks (MIAs) across the OLMo 2 family, finding that: (i) overall performance is limited: the best unsupervised and supervised MIAs both reach an AUC of only 0.68, and supervised MIAs degrade under cross-domain evaluation; (ii) MIA performance peaks at mid-training and is lower at pre-training and post-training, a pattern driven by data type rather than a stage effect: curated math data is far more detectable than other types; (iii) overall scores improve from 1B to 13B but plateau at 32B; and (iv) no unsupervised MIA is robust to distribution shifts, with AUCs shifting by up to 0.42. Finally, we find that our findings on OLMo 2 generalize to OLMo 3 and non-OLMo models.
Chinese Translation
大语言模型(LLM)上的成员推断旨在确定给定文本样本是否被包含在某个 LLM 的训练数据中,而无需访问其训练语料库。尽管最近取得了进展,现有基准仍存在三个局限:对训练阶段的覆盖有限,成员与非成员之间的分布对齐不足,以及缺乏针对训练语料库对非成员进行严格过滤。为解决这些局限,我们提出 OLMo-Detect,一个建立在完全开放的 OLMo 2 流程之上的多阶段、混杂因素受控基准。OLMo-Detect 涵盖预训练、中期训练和后训练,在三个关键轴上显式对齐成员与非成员,并通过 infini-gram 严格过滤非成员。为评估对分布偏移的鲁棒性,我们进一步引入 OLMo-Detect (Shifted),这是一个成员与非成员未对齐的变体。我们在 OLMo 2 系列中评估了 15 种无监督和 3 种有监督成员推断攻击(MIA),发现:(i) 总体性能有限:最佳的无监督和有监督 MIA 都仅达到 0.68 的 AUC,并且有监督 MIA 在跨域评估下性能下降;(ii) MIA 性能在中期训练时达到峰值,而在预训练和后训练时较低,这一模式由数据类型而非阶段效应驱动:精选数学数据远比其它类型更容易被检测;(iii) 总体得分从 1B 到 13B 有所提升,但在 32B 时趋于平稳;以及 (iv) 没有任何无监督 MIA 对分布偏移具有鲁棒性,AUC 变化幅度高达 0.42。最后,我们发现我们在 OLMo 2 上的发现可推广到 OLMo 3 和非 OLMo 模型。
cs.CL / 47 / 2610.02999
OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
OmniConfess:诱导 token 坦白以缓解全模态幻觉
Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu
cs.CL
large language model
大语言模型相关
Abstract
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.
Chinese Translation
全模态大语言模型(OmniLLMs)统一了文本、图像、音频和视频,然而当生成依赖错误证据时便会产生幻觉。现有的推理时方法能够减少幻觉,但很少揭示究竟是哪些证据支撑着某一生成出的承诺。我们提出 OmniConfess,一种无需训练的全模态幻觉缓解方法。它固定一个候选回答,并在受控的逐通道证据干预下以 token 粒度对其重新打分,从而产生一个结构化的逐 token、逐通道的坦白,揭示该回答的证据依赖性。OmniConfess 利用这一坦白来保留有依据的内容,并纠正由无关或矛盾证据驱动的承诺。为评估 OmniConfess,我们构建了 OmniHalluBench,这是一个包含 3,540 个样例的基准,由六个数据集构建而成,涵盖文本、图像、音频和视频场景,以及判断式与自由形式生成两类任务。实验表明,OmniConfess 在异构的模态与任务设置下都能缓解幻觉。我们的代码与基准已公开于 https://github.com/RongHuiQiang/OmniConfess。
cs.CL / 48 / 2610.03039
HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning
HyperThink:用于高效推理的文本到参数超网络
Donggyun Kim, Jack Lu, Chanwoo Kim, Mengye Ren, Seunghoon Hong
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a single query-conditioned parameter update: a lightweight hypernetwork reads the question and predicts updates to a small subset of the base LLM's parameters, while a vector-quantized decoder constrains them to a finite set of reusable patterns to improve robustness and transfer. Trained end-to-end on outputs from the base model itself, HyperThink eliminates long thinking traces at test time: after one hypernetwork forward pass, the adapted model generates a concise step-by-step solution and final answer without an intermediate trace, using far fewer tokens while retaining strong reasoning performance. Empirically, HyperThink improves the low-latency region of the accuracy-latency trade-off on mathematical and general reasoning tasks, with its strongest gains in the near-non-thinking regime.
Chinese Translation
长篇思维轨迹能够显著提升大语言模型(LLMs)的多步推理性能,但它们会带来高昂的推理时开销,其延迟主要由顺序解码主导。我们提出 HyperThink,这是一种文本到参数的方法,它将这一推理计算摊销为单次由查询条件化的参数更新:一个轻量级超网络读取问题并预测对基础 LLM 参数中一小部分参数的更新,而一个矢量量化解码器将这些更新约束到一组有限的可复用模式上,以提升鲁棒性和迁移能力。HyperThink 在基础模型自身产生的输出上进行端到端训练,从而在测试时消除了长篇思维轨迹:在一次超网络前向传播之后,经适配的模型无需中间轨迹即可生成简洁的分步解答与最终答案,在保持强劲推理性能的同时使用的 token 数量远少得多。在实证上,HyperThink 改善了数学与通用推理任务中准确率—延迟权衡的低延迟区域,其最强增益出现在接近非思考的区间。
cs.CL / 49 / 2610.03052
The Geometry of Knowledge Accessibility in Large Language Models
大型语言模型中知识可及性的几何结构
Lihu Chen
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) contain broad knowledge, but they cannot access all of it reliably. We study this problem through knowledge accessibility, which describes whether the knowledge needed for a query can be recalled from the model. We find that knowledge accessibility has a simple geometric structure in the model's representation of the query alone, before any generation. More accessible queries are closer to a center in the representation space, while less accessible queries are farther away. This geometry reveals a knowledge boundary that separates more accessible queries from less accessible ones. Accessibility consistently decreases with distance from the center, and this distance-based ordering transfers across datasets even when the centers differ. Controlled experiments further show that the centered geometry is more closely related to knowledge accessibility than to reasoning difficulty. The geometry also reveals when different interventions are useful. Query rewriting helps more for accessible queries, chain-of-thought reasoning helps more near the boundary, and retrieval gives larger gains beyond the boundary. These findings not only provide a new geometric view of how knowledge is organized in language models, but also suggest a useful pre-generation signal for adaptive inference.
Chinese Translation
大型语言模型(LLMs)包含广泛的知识,但它们无法可靠地获取其中全部知识。我们通过知识可及性来研究这一问题,它描述的是查询所需的知识能否从模型中被召回。我们发现,在模型仅对查询的表示中、在任何生成发生之前,知识可及性具有一种简单的几何结构。可及性更高的查询更接近表示空间中的一个中心,而可及性更低的查询则离该中心更远。这一几何结构揭示了一条知识边界,它将可及性更高的查询与可及性更低的查询分隔开来。可及性随着离中心距离的增加而一致地下降,并且这种基于距离的排序在不同数据集之间可以迁移,即使这些中心各不相同。受控实验进一步表明,这种以中心为参照的几何结构与知识可及性的关联比其与推理难度的关联更为紧密。该几何结构还揭示了不同干预手段在何时有用。查询改写对可及性较高的查询帮助更大,思维链推理在边界附近帮助更大,而检索在边界之外带来的增益更大。这些发现不仅为语言模型中知识如何组织提供了一种新的几何视角,还为自适应推理提示了一个有用的生成前信号。
cs.CL / 50 / 2610.03063
HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation
HARPO:面向忠实且创造性语言生成的幻觉感知强化学习
Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.
Chinese Translation
大型语言模型(LLMs)容易生成幻觉内容,这损害了它们在知识密集型任务中的可靠性。为在不牺牲创造性的情况下应对这一挑战,我们提出 HARPO,一个旨在联合优化忠实性与创造性的强化学习框架。HARPO 包含一个幻觉感知生成式奖励模型(HA-GRM),该模型通过可验证反馈训练,以评估忠实性和写作质量。选择性激活机制(SAM)仅对由 HA-GRM 判定为无幻觉的输出激活写作奖励,而数据课程逐步将训练从创意写作转向以幻觉为中心的任务。在 RAGTruth 上,我们基于 Qwen3-4B 的 HA-GRM 实现了 78.08% 的响应级 F1 分数,相比之下,监督微调基线为 66.37%。在参数从 1.7B 到 8B 的 Qwen2.5 和 Qwen3 模型上的实验显示,忠实生成和写作质量均有提升。在 Qwen3-4B 上,HARPO 将 MultiHopRAG 上由 HA-GRM 判定的幻觉率从 3.29% 降低到 1.02%,同时将 Arena-Hard-v2.0 创意写作分数从 16.95% 提高到 27.54%。
cs.CL / 51 / 2610.03136
Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation
探究推理语言对齐在单语检索增强生成中的作用
Oliver Hauck, Mario Sanz-Guerrero, Katharina von der Wense
cs.CL
large language model
大语言模型相关
Abstract
Reasoning traces improve large language models (LLMs), but current models are trained to reason mostly in English. It has been shown that forcing a model to reason in another language degrades accuracy, even when the reasoning language matches the language of the prompt -- but only for a setting where the model reasons over a short prompt. Here, we ask whether the same holds for retrieval-augmented generation (RAG), where the model must read and integrate a large amount of retrieved evidence in the target language. To study this, we build a fully monolingual German RAG question-answering testbed over the fictional world of the tabletop role-playing game The Dark Eye, a domain that is richly documented in German but too niche for the model to answer from memory, so that it has to rely on retrieval. Varying the forced reasoning language of an agentic RAG system on this testbed, we find that aligning the reasoning language with the language of the query and the retrieved documents helps. Forced German reasoning outperforms forced French, although the model benchmarks higher in French, so the benefit comes from alignment and not from language proficiency. The advantage grows when the retrieved context is richer and structure-aware. However, forced German only reaches the level of the model's native, unconstrained English reasoning without surpassing it, showing that native multilingual reasoning is needed. We publicly release the testbed and QA benchmark.
Chinese Translation
推理轨迹能提升大语言模型(LLM)的表现,但当前模型主要被训练为用英语进行推理。已有研究表明,强制模型用另一种语言进行推理会降低准确率,即使推理语言与提示语言相匹配也是如此——但这仅针对模型在短提示上进行推理的设定。在此,我们追问这一结论是否同样适用于检索增强生成(RAG),在该设定中,模型必须阅读并整合大量以目标语言呈现的检索证据。为研究这一问题,我们围绕桌面角色扮演游戏《黑暗之眼》(The Dark Eye)的虚构世界构建了一个完全单语的德语 RAG 问答测试平台;该领域有丰富的德语文献记录,但对模型而言过于小众,无法凭记忆作答,因此它必须依赖检索。在该测试平台上改变智能体式 RAG 系统被强制的推理语言,我们发现,使推理语言与查询语言以及检索文档的语言保持一致会有所帮助。强制使用德语推理优于强制使用法语推理,尽管该模型在法语上的基准表现更高,因此这一收益来自语言对齐,而非语言熟练程度。当检索到的上下文更丰富且具备结构感知时,这一优势会进一步扩大。然而,强制使用德语仅能达到模型原生的、无约束的英语推理水平,而无法超越它,这表明模型需要具备原生的多语言推理能力。我们公开发布该测试平台与问答基准。
cs.CL / 52 / 2610.03163
Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer
预测用于少样本作者风格迁移的引导向量与适配器权重
Leonard Popp, Danni Liu, Supriti Sinhamahapatra, Jan Niehues
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Adapting large language models to an individual author's style from a few examples is challenging, and scientific writing sharpens the difficulty: formal conventions leave little surface variation, and authors write about their own topics, so extracted ``style'' easily entangles with content. We study style-conditioned abstract generation from a few example abstracts per author and propose three methods: (1) contrastive activation steering, (2) a network that predicts steering vectors, and (3) a hypernetwork that predicts LoRA adapters. We find a consistent trade-off between style imitation and output quality: fine-tuning buys most of the available style signal but forfeits fluency, while the hypernetwork achieves the best trade-off on both seen and unseen authors. Our steering operates at author level, contrasting an author's abstracts against style-neutral generations for the same content. This holds topic fixed, removes the need for a predefined style inventory, and outperforms inventory-based steering. % [EDIT 1a] softened "no single optimal axis" claim Moreover, our analyses demonstrate that manually extracted and predicted steering vectors are near-orthogonal yet score comparably, indicating that style conditioning here can admit at least two unrelated directions rather than requiring one particular axis.
Chinese Translation
仅凭少量示例将大语言模型适配到单个作者的风格颇具挑战,而科学写作更使这一难度加剧:正式写作规范几乎不留表层变化的空间,且作者书写的是自己的主题,因此所提取的``风格''很容易与内容纠缠在一起。我们研究在每位作者仅有少量示例摘要的条件下进行风格条件化的摘要生成,并提出三种方法:(1)对比式激活引导,(2)一个预测引导向量的网络,以及(3)一个预测 LoRA 适配器的超网络。我们发现风格模仿与输出质量之间存在一种一致的权衡:微调获取了大部分可得的风格信号,却牺牲了流畅性,而超网络在已见作者与未见作者上均取得了最佳的权衡。我们的引导在作者层面运作,将某位作者的摘要与针对相同内容生成的风格中性文本进行对比。这保持了主题固定不变,免除了对预定义风格清单的需求,并且优于基于清单的引导。% [编辑 1a] 弱化了“不存在单一最优轴”的论断。此外,我们的分析表明,人工提取的引导向量与预测得到的引导向量近乎正交,却得分相当,这表明此处的风格条件化至少可以容纳两个互不相关的方向,而不必要求某一个特定的轴。
cs.CL / 53 / 2610.03215
StanceEval 2026: The Second Stance Detection Shared Task
StanceEval 2026:第二届立场检测共享任务
Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Asma Yamani, Maram Kurdi, Saad Ezzini, Ahmed Ashraf, Maged Al-Shaibani, Nora Alturayeif
cs.CL
large language model
大语言模型相关
Abstract
StanceEval 2026 is the second edition of the StanceEval shared task series on stance detection in Arabic social media text. Stance detection aims to identify a writer's stance toward a given topic. Given a tweet and a target, participating systems must determine whether the writer's stance is Favor, Against, or None. This edition focuses on cross-target generalization across two distinct evaluation tracks: Track 1 evaluates thematically related cross-target transfer (testing on Women Driving, related to Women Empowerment from training data), while Track 2 evaluates cross-domain transfer to completely unseen targets (E-Cars and Trimester System). The shared task attracted 80 registered teams from 12 countries. During the evaluation phase, 30 unique teams submitted entries, with 21 teams officially ranked in Track 1 and 13 in Track 2 following validation filtering, and 20 teams submitting system-description papers. Participating teams employed diverse methodologies, including fine-tuned pretrained language models, prompt-based and retrieval-augmented large language models (LLMs), fine-tuned LLMs, and hybrid cascades. Top systems achieved impressive $F_{avg2}$ scores of 0.8994 on Track 1 and 0.9400 on Track 2, substantially outperforming the strongest baselines (0.7366 and 0.7475, respectively), where $F_{avg2}$ denotes the macro-averaged F1 score over the Favor and Against classes. Counterintuitively, performance on the unseen targets was higher than on the related target, a disparity could be driven by extreme target polarization, class imbalance, and dialectal or sarcastic nuance across topics.
Chinese Translation
StanceEval 2026 是关于阿拉伯语社交媒体文本立场检测的 StanceEval 共享任务系列的第二届。立场检测旨在识别作者对给定主题的立场。给定一条推文和一个目标,参赛系统必须判定作者的立场是支持(Favor)、反对(Against)还是无(None)。本届任务聚焦于跨目标泛化,设有两个不同的评测赛道:赛道 1 评估主题相关的跨目标迁移(在 Women Driving 上测试,其与训练数据中的 Women Empowerment 相关),而赛道 2 评估向完全未见目标的跨领域迁移(E-Cars 与 Trimester System)。该共享任务吸引了来自 12 个国家的 80 支注册队伍。在评测阶段,共有 30 支不同的队伍提交了结果,经过有效性过滤后,赛道 1 有 21 支队伍、赛道 2 有 13 支队伍获得正式排名,另有 20 支队伍提交了系统描述论文。参赛队伍采用了多样化的方法,包括微调的预训练语言模型、基于提示与检索增强的大语言模型(LLMs)、微调的 LLMs,以及混合级联方法。顶尖系统取得了令人瞩目的 $F_{avg2}$ 分数,赛道 1 为 0.8994,赛道 2 为 0.9400,大幅超越了最强基线(分别为 0.7366 和 0.7475),其中 $F_{avg2}$ 表示在 Favor 和 Against 类别上的宏平均 F1 分数。与直觉相反,在未见目标上的表现高于在相关目标上的表现,这种差异可能由极端的目标极化、类别不平衡以及跨主题的方言或讽刺细微差别所驱动。
cs.CL / 54 / 2610.03240
Collective Bias Mitigation via Model Routing and Collaboration
通过模型路由与协作实现的集体偏见缓解
Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang, Yue Liu, Dong Huang, Huijun Liu, Bin Ji, Jie M. Zhang, See-Kiong Ng
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model's intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained model behavior and fostering knowledge sharing among diverse LLMs. This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses. Experiments show CBM substantially outperforms standalone baselines (e.g., in the top-7 setting, Committee lowers the age bias score from 0.25 to 0.10). Our Debating and Committee topologies achieve substantial bias reduction, with the latter balancing mitigation effectiveness and inference cost, highlighting the potential of CBM for fairer LLMs.
Chinese Translation
大语言模型(LLMs)日益被部署于公共卫生、金融和治理领域,这既要求准确性,也要求与社会价值保持一致。尽管近来取得了诸多进展,LLMs 仍常常延续或放大其训练数据中固有的偏见,给公平性带来挑战。虽然自去偏(self-debiasing)鼓励 LLM 识别并纠正自身的偏见,但仅依赖单一模型的内在知识可能不足以解决根深蒂固的刻板印象。为解决这一局限,我们提出了集体偏见缓解(Collective Bias Mitigation,CBM)框架,该框架通过学习细粒度的模型行为并促进多样 LLM 之间的知识共享来缓解偏见。本工作是首个系统性地探索如何有效选择与组织不同的 LLM 以培养更公平 LLM 响应的研究。实验表明,CBM 大幅优于单模型基线(例如,在 top-7 设置下,Committee 将年龄偏见分数从 0.25 降至 0.10)。我们的 Debating 与 Committee 拓扑结构实现了显著的偏见降低,其中后者在缓解效果与推理成本之间取得平衡,凸显了 CBM 在实现更公平 LLM 方面的潜力。
cs.CL / 55 / 2610.03268
Shrome at Touché: Soft-Vote Ensembling and Counter-Causal Augmentation for Causality Extraction
Shrome 在 Touché:用于因果抽取的软投票集成与反因果增强
Roham Zendehdel Nobari, Shayan Sooratgar
cs.CL
large language model
大语言模型相关
Abstract
Touché 2026 extends causality extraction to counter-causal claims: news sentences whose surface form appears causal but whose meaning denies the causation, as in "It is falsely believed that X caused Y." A system that relies on surface cues such as "caused" or "led to" will accept such a sentence as causal and give it the wrong polarity. On the Countercausal News Corpus (CCNC), the task has three subtasks: deciding whether a sentence is causal (detection), locating its cause and effect spans (extraction), and labeling its polarity as procausal, counter-causal, or uncausal. We build one model per subtask. Detection is a fine-tuned classifier with a single cross-task rule that uses the extracted spans to remove false positives. For extraction, we ensemble three RoBERTa-large BILOU+CRF taggers by averaging their token-level scores before decoding, rather than voting on the spans each tagger produces. For polarity, where labeled counter-causal examples are scarcest, we add training sentences generated by a large language model prompted with nine patterns of counter-causal expression adapted from Hagen et al., keeping only those that pass automatic structural checks. On the held-out CCNC test set, the system reaches F1 0.869 on detection and macro-F1 0.817 on polarity, and in the organizers' final causal-only evaluation of extraction it scores granularity-adjusted F1 0.728, the highest extraction score among all submissions including the organizers' baseline. The development split is used only for component selection and the ablations reported in the paper.
Chinese Translation
Touché 2026 将因果抽取扩展到反因果主张:即表层形式看似因果、但含义否认该因果关系的新闻句子,如“人们错误地认为 X 导致了 Y”。一个依赖诸如 "caused"(导致)或 "led to"(致使)这类表层线索的系统会把这样的句子判定为因果句,并给出错误的极性。在反因果新闻语料库(Countercausal News Corpus, CCNC)上,该任务包含三个子任务:判定一个句子是否为因果句(检测)、定位其原因与结果片段(抽取),以及将其极性标注为正因果、反因果或无因果。我们为每个子任务各构建一个模型。检测是一个微调后的分类器,并附带一条单一的跨任务规则,该规则利用抽取出的片段来去除假阳性。对于抽取,我们将三个 RoBERTa-large BILOU+CRF 标注器进行集成,在解码之前对它们的词元级分数取平均,而不是对每个标注器产生的片段进行投票。对于极性判定——其中带标注的反因果样本最为稀缺——我们加入了由大语言模型生成的训练句子,其提示词采用了改编自 Hagen 等人的九种反因果表达模式,并只保留那些通过自动结构检查的句子。在留出的 CCNC 测试集上,该系统在检测任务上达到 F1 0.869,在极性判定上达到 macro-F1 0.817;在组织方最终的仅因果抽取评估中,它取得粒度调整后的 F1 0.728,这是包括组织方基线在内的所有提交中抽取得分最高的。开发集划分仅用于组件选择以及论文中报告的消融实验。
cs.CL / 56 / 2610.03324
To Jev or Not? Evaluating the Accuracy and Efficiency of Structured Decision Models for Hate-Speech Moderation
Jev 还是不 Jev?评估结构化决策模型用于仇恨言论审核的准确性与效率
Demetris Paschalides, George Pallis, Marios D. Dikaiakos
cs.CL
large language model
大语言模型相关
Abstract
The scale of online content makes hate-speech moderation challenging, while Large Language Models (LLMs) enable harmful material to be produced and adapted more easily. Moderation therefore requires efficient classifiers that can accommodate different definitions of hate speech. Recent structured decision models accept natural-language criteria and select among specified answers, raising the question of whether they can meet these requirements without task-specific training. We present HATEDECIDE, an evaluation of six decision-model configurations on four hate-speech datasets against specialized moderation, zero-shot, commercial, and supervised baselines. We examine whether supplying a dataset's definition, or decomposing it into multiple questions, improves classification, and we measure their latency and cost. We find that commercial LLMs significantly outperform all decision models on only one dataset. Supplying definitions changes up to 28\% of predictions without consistently improving classification, and decomposition significantly improves performance in only 20\% of the comparisons. On a diagnostic set of test cases, the best hosted decision model comes within 1.6 macro-F1 points of the best commercial LLM at approximately 97\% lower inference cost. These results identify opportunities for inexpensive moderation, while showing that explicit criteria and additional questions do not reliably improve classification.
Chinese Translation
在线内容的规模使仇恨言论审核具有挑战性,而大型语言模型(LLMs)使有害材料能够更容易地被生成和改编。因此,审核需要高效分类器,能够适应仇恨言论的不同定义。近期的结构化决策模型接受自然语言标准,并在指定答案中进行选择,这提出了一个问题:它们能否在没有任务特定训练的情况下满足这些要求。我们提出 HATEDECIDE,这是一个在四个仇恨言论数据集上对六种决策模型配置进行的评估,并将其与专用审核、零样本、商业和有监督基线进行比较。我们考察提供数据集的定义,或将其分解为多个问题,是否会改善分类,并测量它们的延迟和成本。我们发现,商业 LLMs 仅在一个数据集上显著优于所有决策模型。提供定义会改变多达 28\% 的预测,但并未持续改善分类;而分解仅在 20\% 的比较中显著提升性能。在一组诊断性测试用例上,最佳托管决策模型与最佳商业 LLM 的差距在 1.6 个 macro-F1 点以内,而推理成本约低 97\%。这些结果指出了实现低成本审核的机会,同时表明明确的标准和额外问题并不能可靠地改善分类。
cs.CL / 57 / 2610.03329
SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models
SyntaxBench:面向大语言模型字符级推理的统计诊断框架
Mohsen Larni, Sobhan Ebrahimi Azar, Pouyan Nahed, Kazem Taghva
cs.CL · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.
Chinese Translation
大语言模型正越来越多地被用于小句法错误会产生影响的场景,然而字符级推理仍主要通过孤立的探测和总体准确率来评估。我们提出 SyntaxBench,一个用于字符级推理的诊断基准与统计评估框架。它包含五个核心任务:字符计数、字母包含、回文检测、编辑距离和最长字符串选择,外加 index_to_span,一个更困难的子串抽取压力测试。这五个核心任务使用配对的英语输入和字符长度匹配的随机字符串输入。index_to_span 文档共享 200-500 词的区间,并且不进行字符长度匹配。所有六个任务都使用零样本、单样本和四样本提示。我们在 11 种推理模式配置下评估了八个参数规模从 2B 到 32B 的开放权重模型。该框架报告精确匹配准确率和宽松准确率、Cohen's kappa、带优势比的配对 McNemar 检验、自助法置信区间、Kendall's tau、类别条件指标、分词分析以及多重比较校正检验。有三项发现尤为突出。第一,分词塑造准确率:随机字符串比英语字符串更具字符可见性(每 token 1.892 个字符 vs. 3.169 个字符),并且随着英语单词占据更多 token,字符计数准确率下降。第二,推理模式并非一致有帮助:在接近饱和的任务上,Gemma4-31B 在不同模式间几乎不变,而 Qwen3.6-27B 在回文检测上使用思考模式时表现更差(四样本时非思考 0.952 vs. 思考 0.886)。第三,index_to_span 在很大程度上仍未解决;最佳四样本精确匹配准确率为 6.75%。字符级评估需要受控输入、配对检验,以及对分词和推理模式的分析,而不仅仅是总体准确率。
cs.CL / 58 / 2610.03421
CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation
CLIMB:面向多模态检索增强生成的置信度引导互补证据
Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas
cs.CL
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model's parametric knowledge. Existing multimodal RAG systems commonly rely on Top-$K$ retrieval or reranking, which may return redundant passages and provide limited control over whether an answer update is sufficiently supported by the retrieved evidence. We propose \textit{CLIMB}, a training-free inference-time framework for multimodal RAG. CLIMB first constructs a compact complementary evidence pool using an MMR-style objective that balances query relevance and passage-level redundancy. It then performs confidence-controlled refinement within this fixed pool: an R/E/C critic scores passages by relevance, evidence specificity, and cross-modal alignment, while an evidence-grounded confidence estimator accepts an updated answer only when the estimated confidence increases. This design provides a simple stopping criterion and reduces unnecessary refinement without modifying the underlying retriever or MLLM. Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines. Ablations further indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.
Chinese Translation
多模态大语言模型(MLLMs)已展现出强大的视觉推理能力,但知识密集型的视觉问答往往需要图像和模型参数化知识之外的外部文本证据。现有多模态 RAG 系统通常依赖 Top-$K$ 检索或重排序,这可能返回冗余段落,并且对答案更新是否得到检索证据的充分支持提供有限的控制。我们提出了 \textit{CLIMB},一个面向多模态 RAG 的免训练、推理时框架。CLIMB 首先使用一个平衡查询相关性与段落级冗余的 MMR 式目标构建一个紧凑的互补证据池。然后,它在这个固定池内执行置信度控制的精炼:一个 R/E/C 评论器依据相关性、证据特异性和跨模态对齐对段落进行评分,而一个基于证据的置信度估计器仅在估计置信度提高时才接受更新后的答案。该设计提供了一个简单的停止准则,并在不修改底层检索器或 MLLM 的情况下减少了不必要的精炼。在 Encyclopedic-VQA 和 InfoSeek 上的实验表明,CLIMB 相较检索增强的多模态基线持续取得提升。消融实验进一步表明,互补池化、基于评论器的评分以及迭代式置信度控制的精炼各自都对最终性能有所贡献。
cs.CL / 59 / 2610.03525
Structured Composition of Verifiable Atomic Insights for Table-to-Report Generation
面向表格到报告生成的可验证原子洞见的结构化组合
Teng Lin, Xinyu Liu, Nan Tang
cs.CL
large language model
大语言模型相关
Abstract
Table-to-report generation refers to the task of automatically generating article-level analyt- ical reports from relational tables and is an essential capability for automated data science and decision support. Its central challenge lies in systematically discovering verifiable com- posite insights across tables, attributes, and analytical perspectives, and organizing them into coherent, complete, and traceable evidence chains. Existing methods primarily rely on sequential, reactive data agents or direct Large Language Model(LLM) generation. They suffer from exploration bias: early local observations constrain subsequent actions, causing models to focus prematurely on local analyzes and miss cross-table or cross-dimensional evidence. We propose ComInsight, which reformulates insight discovery as the composition of atomic evidences. We first define an atomic insight as the smallest executable analytical unit conforming to a predefined analysis pattern and enumerate all valid atomic insights from database schema and content. These atoms are then organized into a multi-relational insight graph, where nodes represent verified data facts and edges encode logical, temporal, or hierarchical relations. Finally, a set of composition operators systematically fuses atomic nodes into higher-order composite conclusions. Every composite output is accompanied by executable SQL and fine-grained provenance, ensuring full verifiability. Across three benchmarks InsightBench, DDR-Bench, and T2R-Bench, ComInsight consistently outperforms strong baselines in factual correctness, novelty, and structural completeness. We believe ComInsight offers a reliable, efficient, and explainable path toward table-to-report generation.
Chinese Translation
表格到报告生成指的是从关系表中自动生成文章级分析报告的任务,并且是自动化数据科学与决策支持的一项基本能力。其核心挑战在于系统性地发现跨表、属性和分析视角的可验证复合洞见,并将它们组织成连贯、完整且可追溯的证据链。现有方法主要依赖顺序式、反应式的数据智能体,或直接依赖大语言模型(LLM)生成。它们受到探索偏差的影响:早期的局部观察会限制后续行动,导致模型过早聚焦于局部分析,并遗漏跨表或跨维度的证据。我们提出 ComInsight,它将洞见发现重新表述为原子证据的组合。我们首先将原子洞见定义为符合预定义分析模式的最小可执行分析单元,并从数据库模式和内容中枚举所有有效的原子洞见。然后,这些原子被组织成一个多关系洞见图,其中节点表示已验证的数据事实,边编码逻辑、时间或层次关系。最后,一组组合算子系统性地将原子节点融合为高阶复合结论。每个复合输出都伴随可执行 SQL 和细粒度溯源信息,从而确保完全可验证性。在 InsightBench、DDR-Bench 和 T2R-Bench 这三个基准上,ComInsight 在事实正确性、新颖性和结构完整性方面始终优于强基线。我们相信,ComInsight 为表格到报告生成提供了一条可靠、高效且可解释的路径。
cs.CR / 60 / 2610.02349
MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication
MIRROR:面向 LLM 多智能体通信的多路径仲裁完整性
Ryuichi Yamafuji Lun, Jingzhen Wang, Shreyas Kolte, Ruiteng Li
cs.CR · cs.AI · cs.MA
large language model
大语言模型相关
Abstract
Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100% on structured tasks. Existing defenses rely on semantic validation, which requires additional inference and can block benign outputs, or on transport-layer encryption, which does not help when an intermediary legitimately terminates TLS. We present MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest. MIRROR uses unkeyed hashing and so authenticates nothing on its own, since an active on-path adversary can always recompute a digest over a payload it has modified. All integrity derives from the assumption that honest routes form a majority. The digest serves only to make witness routes constant-size and to bind the recovered payload to the quorum-agreed value under second-preimage resistance. We give the guarantee under a route-compromise bound alpha < 0.5, and extend it to correlated routes, where the quantity that matters is the size of the largest shared-failure group and not the route count. We further show that availability and integrity degrade at the same threshold: below alpha = 0.5, quorum-denial and message-dropping adversaries cannot block honest traffic. Across MMLU, HumanEval, and MBPP on two frameworks and four communication topologies, and in a MetaGPT deployment against a production API, MIRROR reduces ASR to 0% below the threshold at 1x LLM token cost. LLM-as-a-Judge costs 35x in the same deployment, and blocks up to 44.2% of benign outputs in the topology sweep.
Chinese Translation
智能体间通信是大型语言模型多智能体系统(LLM-MAS)的核心,但它引入了一个尚未被充分探索的漏洞:中间智能体(Agent-in-the-Middle,AiTM)攻击,这类攻击在传输途中操纵消息,而无需攻陷智能体本身。先前工作报告,在结构化任务上攻击成功率(ASR)接近 100%。现有防御依赖于语义验证,这需要额外的推理,并且可能阻断良性输出;或者依赖于传输层加密,而当中介合法地终止 TLS 时,传输层加密无济于事。我们提出 MIRROR,一种通信层完整性原语,它将单个规范化后的载荷复制到 k 条逻辑路由上,并且仅当严格多数路由报告相同摘要时才接受消息。MIRROR 使用无密钥哈希,因此其自身并不能认证任何东西,因为主动的路径上对手总能针对其已修改的载荷重新计算摘要。所有完整性均源自诚实路由构成多数的假设。摘要仅用于使见证路由保持恒定大小,并在第二原像抗性下将恢复出的载荷绑定到仲裁一致值。我们在路由被攻陷界限 alpha < 0.5 下给出该保证,并将其扩展到相关路由,其中重要的量是最大共享故障组的大小,而不是路由数量。我们进一步表明,可用性和完整性在同一阈值处退化:在 alpha = 0.5 以下,拒绝仲裁的对手和丢弃消息的对手无法阻断诚实流量。在 MMLU、HumanEval 和 MBPP 上,跨两个框架和四种通信拓扑,以及在针对生产 API 的 MetaGPT 部署中,MIRROR 在阈值以下将 ASR 降至 0%,且 LLM token 成本为 1x。在同一部署中,LLM-as-a-Judge 的成本为 35x,并在拓扑扫描中阻断多达 44.2% 的良性输出。
cs.CR / 61 / 2610.02418
Mitigating Private Data Leakage in LLMs with Whiteout
使用 Whiteout 缓解 LLM 中的私有数据泄露
Anna Yoo Jeong Ha, Ronik Bhaskar, Haitao Zheng, Ben Y. Zhao
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This leads to significant privacy risks, particularly for high-profile individuals such as executives, politicians, and judges. Existing mitigations largely rely on machine unlearning. However, these methods often remove more information than needed, degrade model utility and safety, and are highly vulnerable to attacks. This paper presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine PSIs, by overwriting them using precise and carefully designed obfuscation samples. We evaluate Whiteout on modern LLMs of varying sizes and makers, including a widely-used OpenAI model. Results show that Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, and outperforms existing alternatives. We also test Whiteout against a wide range of countermeasures, from black-box attacks like jailbreaking to white-box adaptive attacks like relearning and quantization. Finally, we conclude with a discussion on the security and ethical implications of Whiteout.
Chinese Translation
现代大型语言模型(LLM)是在海量且基本未经筛选的数据集上训练的,这些数据集包含从几乎所有可访问网站抓取的内容以及用户输入。因此,LLM 往往会记忆并重现个人敏感信息(PSI),例如出生日期、电话号码和家庭住址。这带来了重大的隐私风险,尤其是对于高管、政治人物和法官等公众人物而言。现有的缓解方法在很大程度上依赖于机器遗忘。然而,这些方法往往会删除超出所需范围的信息,降低模型的效用与安全性,并且极易受到攻击。本文提出了 Whiteout,这是一种实用工具,它可应个人的请求,通过使用精确且精心设计的混淆样本来覆写其真实 PSI,从而阻止 LLM 复述这些信息。我们在不同规模和不同厂商的现代 LLM 上评估了 Whiteout,其中包括一个被广泛使用的 OpenAI 模型。结果表明,Whiteout 能有效防止目标 PSI 的泄露,对模型效用和安全性的影响可忽略不计,并且优于现有的替代方案。我们还针对各类对抗手段对 Whiteout 进行了测试,从越狱等黑盒攻击到重新学习和量化等白盒自适应攻击。最后,我们以关于 Whiteout 的安全与伦理影响的讨论作为总结。
cs.CR / 62 / 2610.02432
Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
评估与改进大型语言模型对输入序列变化的鲁棒性
Narek Maloyan
cs.CR · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.
Chinese Translation
生产系统中的大型语言模型(LLM)面临提示注入、木马(后门)以及自动质量指标操纵。本论文开发了用于评估和改进 LLM 对对抗性输入序列变化鲁棒性的模型、方法和算法。我们提出 R_stab(f),一种基于小输入扰动下每步输出分布之间的 Jensen-Shannon 散度的生成式鲁棒性度量。对于局部化攻击,我们证明 V(h) <= 1 - R_class(h),其中 R_class(h) 是决策算子 h 在小扰动下保持其决策的概率。对于非局部化攻击,我们提出一个校准的经验模型。对于 LLM-as-a-Judge 系统,我们开发了 ASA,一种自适应进化黑盒攻击,其攻击成功率(ASR)最高可达 73.8%,在开放模型之间的迁移率最高可达 62.6%。在 Trojan Detection Challenge 2023 数据(Pythia-1.4B)上,替代触发器达到 REASR ~0.99,而真实触发器的召回率约为 ~0.17,相对于约 ~0.14 的基线。在 SaTML CTF 2024 上,我们系统化了多层防御的四类绕过,这些防御将 ASR 从 90% 降低到 15-25%。由 5-7 个异构模型组成的委员会将 Gemma-3-4B 的 ASR 降低了 47-55 个百分点,使用 7 个模型时降至 19.3%。对于基于模型上下文协议(MCP)的智能体系统,我们提出 AttestMCP,它使用受 HMAC 保护的数据包对工具调用进行证明,每次调用耗时低于 0.1 毫秒,以及 Commit Boundary 隔离模式。在包含 847 个场景的 MCPBench 基准上,它们将平均 ASR 从 53.7% 降低到 12.4%。这些方法已在 JudgeGuard 和 TrojanArmor 软件套件以及 MCPSec 模块中实现。
cs.CR / 63 / 2610.02544
CITADEL: CWE-Guided Insertion of Hardware Trojans via Analysis of DFG-Enabled LLMs
CITADEL:通过分析 DFG 赋能的 LLM 实现 CWE 引导的硬件木马插入
Jayanth Thangellamudi, Raghul Saravanan, Sudipta Paria, Swarup Bhunia, Sai Manoj P D
cs.CR
large language model
大语言模型相关
Abstract
The increasing sophistication of Hardware Trojans (HTs) and system-level vulnerabilities poses significant risks to modern integrated circuits. However, constructing realistic HT scenarios, remains a substantial burden: researchers must manually analyze complex RTL structures, identify plausible weaknesses, and craft stealthy, synthesizable insertions that preserve functional correctness. This paper introduces CITADEL CWE-Guided Insertion of Trojans via Analysis of DFG-Enabled LLMs, a framework that leverages Large Language Models (LLMs) and Data Flow Graphs (DFGs) to automate CWE-grounded HT synthesis. CITADEL uses structured CWE semantics together with DFG-derived structural context to assist the user in identifying relevant vulnerabilities, localize the module surrounding the chosen insertion point, and perform intent-conditioned RTL modification. The framework produces minimal, synthesizable, and interface-preserving HTs with ultra-rare triggers. Experimental evaluation across diverse RTL designs demonstrates that all generated HTs are 100% syntactically correct, remain undetectable under large-scale random simulation, and are functionally triggerable under their intended activation conditions. These results highlight CITADEL as a scalable and principled method for generating realistic HT benchmarks.
Chinese Translation
硬件木马(HTs)和系统级漏洞日益复杂,给现代集成电路带来了重大风险。然而,构建真实的 HT 场景仍然是一项巨大负担:研究人员必须手动分析复杂的 RTL 结构,识别可能的弱点,并精心设计隐蔽、可综合且保持功能正确性的插入。本文介绍了 CITADEL:通过分析 DFG 赋能的 LLM 进行 CWE 引导的木马插入,这是一个利用大型语言模型(LLM)和数据流图(DFG)来自动化基于 CWE 的 HT 合成的框架。CITADEL 使用结构化的 CWE 语义以及由 DFG 导出的结构上下文,来协助用户识别相关漏洞、定位所选插入点周围的模块,并执行以意图为条件的 RTL 修改。该框架生成具有超罕见触发器的最小化、可综合且保持接口的 HT。跨多种 RTL 设计的实验评估表明,所有生成的 HT 都 100% 语法正确,在大规模随机仿真下仍不可检测,并且在其预期激活条件下可在功能上被触发。这些结果凸显了 CITADEL 作为一种可扩展且有原则的方法,可用于生成真实的 HT 基准测试。
cs.CR / 64 / 2610.02817
RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models
RMCW:一种基于 Reed--Muller 码的面向语言模型的抗删除水印
Yi Wang, Baicheng Chen, Yu Wang, Jian Zhao, Yilei Chen, Tianxing He
cs.CR · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed--Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed--Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords. During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at https://github.com/BaichengDanny/RMCW.
Chinese Translation
大语言模型(LLM)水印提供了一种轻量级机制,用于识别由特定模型生成的文本,但其鲁棒性在后处理攻击下仍然脆弱。删除攻击尤其具有挑战性,因为它们会移动 token 位置,并破坏观测到的 token 与其原始水印位置之间的对齐。我们提出 Reed--Muller 码水印(Reed--Muller Code Watermarking,RMCW),一种基于 Reed--Muller 码的 LLM 水印方法。与全局码字恢复不同,RMCW 寻找幸存的局部代数结构,利用由 Reed--Muller 码字的仿射直线限制所诱导的 Reed--Solomon 一致性。在生成过程中,RMCW 通过秘密密钥化的词表划分将 Reed--Muller 结构注入序列中。在检测过程中,它将给定文本映射到密钥化词表桶,并使用 Berlekamp--Welch 检验来测试局部子序列是否具有低次 Reed--Solomon 一致性。在 C4 和 ELI5 数据集上使用 OPT-1.3B 和 Llama-3.1-8B-Instruct 进行的实验表明,RMCW 保持了较强的干净文本可检测性,并在多种删除和改写攻击下优于或媲美基线方法。我们的代码可在 https://github.com/BaichengDanny/RMCW 获取。
cs.CR / 65 / 2610.02861
Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes
遏制自主运维体:面向 Kubernetes 上 AI 智能体安全的纵深防御框架与参考架构
Simhadri Podala Narasimha
cs.CR · cs.NI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents are moving from chat interfaces into infrastructure operations, where they read telemetry, call tools, generate and execute code, and change the state of production Kubernetes clusters. This collapses a boundary that conventional cloud-native security assumes: the boundary between data and control. Content that an agent merely reads (a log line, a ticket, a tool description) can redirect what it does. This paper argues that the model must not be treated as a security boundary and that agent safety on Kubernetes is therefore an infrastructure problem: every guarantee must continue to hold under the assumption that the agent is fully compromised by prompt injection. We contribute (i) a threat model and ten-class threat taxonomy for agents operating on and within Kubernetes, aligned with emerging OWASP guidance for agentic applications; (ii) nine design principles, centered on complete mediation at the tool boundary and on breaking the combination of untrusted input, sensitive access, and external egress; (iii) a seven-layer defense-in-depth framework that maps each principle to native or widely adopted Kubernetes mechanisms: workload identity, RBAC and ValidatingAdmissionPolicy, gVisor/Kata sandboxing via the SIG Apps Agent Sandbox project, FQDN-aware egress policy, an agent/MCP gateway with policy-as-code over tool arguments, and eBPF runtime enforcement; (iv) a reference architecture with concrete policy artifacts and per-layer bindings for Amazon EKS, Azure Kubernetes Service, and Google Kubernetes Engine; and (v) a qualitative evaluation comprising a threat-control coverage matrix and four attack walkthroughs, with a proposed empirical methodology. We report no measured attack-success or overhead figures; instead we identify residual risks and the measurements needed to validate the framework.
Chinese Translation
大型语言模型(LLM)智能体正在从聊天界面进入基础设施运维领域,在那里它们读取遥测数据、调用工具、生成并执行代码,并改变生产 Kubernetes 集群的状态。这打破了传统云原生安全所假定的一个边界:数据与控制之间的边界。智能体仅仅读取的内容(一行日志、一张工单、一段工具描述)就可能改变它的行为方向。本文主张,模型绝不能被视为安全边界,因此 Kubernetes 上的智能体安全性是一个基础设施问题:每一项保证都必须在假设智能体已被提示注入完全攻陷的前提下依然成立。我们的贡献包括:(i) 面向在 Kubernetes 之上及之内运行的智能体的威胁模型与十类威胁分类体系,并与新近出现的 OWASP 针对智能体应用的指南保持一致;(ii) 九项设计原则,其核心是在工具边界处实现完全中介,并打破不可信输入、敏感访问与外部外联三者的组合;(iii) 一个七层纵深防御框架,将每一项原则映射到 Kubernetes 原生或广泛采用的机制:工作负载身份、RBAC 与 ValidatingAdmissionPolicy、通过 SIG Apps Agent Sandbox 项目实现的 gVisor/Kata 沙箱、感知 FQDN 的外联策略、针对工具参数以策略即代码方式运行的智能体/MCP 网关,以及 eBPF 运行时强制执行;(iv) 一个参考架构,包含具体的策略制品以及针对 Amazon EKS、Azure Kubernetes Service 和 Google Kubernetes Engine 的逐层绑定;以及 (v) 一项定性评估,包括威胁—控制覆盖矩阵和四个攻击过程演练,并提出了一套实证方法论。我们未报告任何实测的攻击成功率或开销数值;相反,我们指出了残余风险以及验证该框架所需的测量方法。
cs.CR / 66 / 2610.02955
Digital Twin-Assisted Mapping of ICS Telemetry to ATT&CK for ICS with Evidence-Driven Dependency Reasoning
数字孪生辅助的 ICS 遥测到 ATT&CK for ICS 映射:基于证据驱动的依赖推理
Konstantinos E. Kampourakis, Vyron Kampourakis, Vasileios Gkioulos, Sokratis Katsikas
cs.CR
large language model
大语言模型相关
Abstract
Reconstructing adversarial behavior from Industrial Control System (ICS) telemetry is difficult because process observations reveal physical changes more directly than the actions that produced them. This paper presents a Digital Twin (DT)-assisted framework that extracts synchronized state changes, converts them into evidence-preserving descriptions, maps them to ATT&CK for ICS through retrieval-augmented Large Language Model (LLM) reasoning, and constructs a typed dependency graph. Evaluation comprises a four-configuration ablation on nine held-out SWaT scenarios containing ten telemetry-evaluable ground-truth episodes, three independent generations of the principal mapping configurations, and external evaluation on BATADAL and WADI. Across the three SWaT generations, DT-enriched mapping produces fewer annotation-relative False Positive (FP) episodes in every run, with a mean per-run reduction of 34.1%. Both configurations obtain a mean recall of 0.633, although recall varies between 0.50 and 0.70 and DT enrichment does not improve F1 in every run. These observations are descriptive: the primary-run paired comparisons do not reach statistical significance, and the mapping advantage does not transfer to either external dataset. An exploratory positive-only dependency benchmark recovers eight of nine documented Co-occurrence relationships with DT context. A separate end-to-end benchmark containing one positive pair and 25 negative controls exposes propagation of mapping errors into unsupported edges. The findings identify both the potential and limitations of DT context for semantic security interpretation, without establishing reliable autonomous attribution, general causal reconstruction, or practical analyst benefit.
Chinese Translation
从工业控制系统(ICS)遥测中重建对抗行为是困难的,因为过程观测所揭示的物理变化比产生这些变化的动作更为直接。本文提出一个数字孪生(DT)辅助框架,该框架提取同步的状态变化,将其转换为保留证据的描述,通过检索增强的大语言模型(LLM)推理将其映射到 ATT&CK for ICS,并构建一个有类型的依赖图。评估包括:在九个留出的 SWaT 场景上进行的四配置消融,这些场景包含十个可基于遥测评估的真实标注事件段;对主要映射配置进行的三次独立生成;以及在 BATADAL 和 WADI 上的外部评估。在三次 SWaT 生成中,DT 增强的映射在每一次运行中都产生更少的相对于标注的假阳性(FP)事件段,平均每次运行减少 34.1%。两种配置均获得 0.633 的平均召回率,尽管召回率在 0.50 到 0.70 之间波动,且 DT 增强并非在每次运行中都提升 F1。这些观察结果是描述性的:主运行的配对比较未达到统计显著性,且映射优势未能迁移到任一外部数据集。一个探索性的仅正例依赖基准在具有 DT 上下文的情况下恢复了九个已记录的共现关系中的八个。一个单独的端到端基准包含一个正例对和 25 个阴性对照,揭示了映射错误向无支撑边的传播。这些发现既指出了 DT 上下文在语义安全解释方面的潜力,也指出了其局限,同时并未确立可靠的自主归因、一般性因果重建或对分析人员的实际收益。
cs.CR / 67 / 2610.03014
Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents
超越预定义汇聚点:面向LLM智能体的安全感知依赖分析
Hang Cui
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM)-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external tools. Existing analyses often use predefined sensitive operations as anchors, but operation identity alone is insufficient to determine security implications. We present AgentSecGraph, a security-aware static analysis framework that constructs a candidate-centered Security-Aware Agent Dependency Graph (Security-ADG) for each security-sensitive operation. It augments operation identity with agent relevance, source and dependency evidence, trust-boundary context, guard evidence, and external-effect semantics. We further introduce AgentSecBench, a corpus of 67 real-world LLM-agent repositories spanning 11 ecosystems and 37,542 source files. The current analyzer identifies 23,866 static security-sensitive operation candidates across 65 repositories and emits one Security-ADG artifact per candidate. Corpus-wide analysis recovers source-to-operation dependency evidence for 9,821 candidates (41.15%) and potential guard evidence for 3,075 (12.88%), completing in 50.8 minutes. Using a separate reproduction-backed evaluation layer, we establish 22 security-sensitive behaviors across 13 repositories: one confirmed vulnerability, one pending disclosure candidate, and 20 guarded behaviors. In nine held-out cases, Security-ADG preserves 91.1% of the reference context and all five observed guards, compared with 20.0% for a sink-only view and 40.0% for a simplified ADG. These results show that security-aware dependency and contextual evidence enable distinctions that cannot be recovered from sensitive-operation identity alone.
Chinese Translation
基于大语言模型(LLM)的智能体日益将模型生成的决策与安全敏感的软件能力连接起来,例如命令执行、文件系统访问、网络通信、浏览器控制和外部工具。现有分析通常使用预定义的敏感操作作为锚点,但仅凭操作身份不足以确定安全影响。我们提出 AgentSecGraph,一个安全感知的静态分析框架,它为每个安全敏感操作构建以候选为中心的安全感知智能体依赖图(Security-ADG)。它通过智能体相关性、来源与依赖证据、信任边界上下文、防护证据以及外部效应语义来增强操作身份。我们进一步提出 AgentSecBench,一个包含 67 个真实世界 LLM 智能体仓库、跨越 11 个生态系统和 37,542 个源文件的语料库。当前分析器在 65 个仓库中识别出 23,866 个静态安全敏感操作候选,并为每个候选生成一个 Security-ADG 制品。全语料库分析为 9,821 个候选(41.15%)恢复了从来源到操作的依赖证据,并为 3,075 个候选(12.88%)恢复了潜在防护证据,在 50.8 分钟内完成。使用一个独立的、以复现为支撑的评估层,我们在 13 个仓库中确立了 22 个安全敏感行为:一个已确认漏洞、一个待披露候选以及 20 个受防护行为。在九个留出案例中,Security-ADG 保留了 91.1% 的参考上下文以及全部五个观测到的防护,相比之下,仅汇聚点视图为 20.0%,简化 ADG 为 40.0%。这些结果表明,安全感知的依赖与上下文证据能够实现仅凭敏感操作身份无法恢复的区分。
cs.CR / 68 / 2610.03153
EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents
EvoRiskBench:面向工作区智能体运行时安全风险的演进基准
Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchmark evolution. We introduce EvoRiskBench, an evolving benchmark organized around the EP-Path-EF framework, which links an initial risk entry point to a one-hop technical effect through an agent-mediated risk path. The framework defines nine entry-point categories and five effect categories; a 20-participant study supports their interpretability and classification consistency on representative cases. Guided by this framework, an automated end-to-end workflow constructs and executes risk cases in isolated environments and independently verifies outcomes using runtime traces and environment states. The benchmark provides a reproducible dataset of 450 adversarial tasks across six scenarios. We evaluate nine model-harness configurations spanning three models (GPT-5.6 Sol, DeepSeek-V4-Pro-0813, and Claude Opus 5) and three harnesses (Claude Code, Codex, and OpenClaw). Our results reveal substantial vulnerabilities across systems. The most vulnerable configuration, Codex with DeepSeek-V4-Pro-0813, reaches a 68.44% attack success rate (ASR), indicating that configuration of workspace agent is insufficient to ensure secure autonomous execution. ASR varies more across models than harnesses, and harness differences depend on the model. The benchmark cases and evaluation platform will be released after completion of artifact safety and reproducibility checks.
Chinese Translation
工作区智能体将大语言模型与执行框架相结合,以执行有状态、多步骤的任务,这些任务会访问或修改外部资源。现有基准在其运行时安全风险的可执行覆盖方面仍存在空白,而不断演化的模型能力、执行框架、工具和威胁推动着基准的演化。我们提出 EvoRiskBench,一个围绕 EP-Path-EF 框架组织的演进基准,该框架通过由智能体介导的风险路径,将初始风险入口点与单跳技术效应联系起来。该框架定义了九类入口点和五类效应;一项有20名参与者的研究支持了它们在代表性案例上的可解释性和分类一致性。在该框架指导下,一个自动化端到端工作流在隔离环境中构建并执行风险案例,并使用运行时轨迹和环境状态独立验证结果。该基准提供了一个可复现的数据集,包含跨六个场景的450个对抗性任务。我们评估了九种模型-执行框架配置,涵盖三个模型(GPT-5.6 Sol、DeepSeek-V4-Pro-0813 和 Claude Opus 5)和三种执行框架(Claude Code、Codex 和 OpenClaw)。我们的结果揭示了各系统中存在大量漏洞。最脆弱的配置,即 Codex 与 DeepSeek-V4-Pro-0813 的组合,达到了68.44%的攻击成功率(ASR),这表明工作区智能体的配置不足以确保安全的自主执行。ASR 在不同模型之间的差异大于在不同执行框架之间的差异,并且执行框架的差异取决于模型。该基准用例和评估平台将在完成制品安全性和可复现性检查后发布。
cs.CR / 69 / 2610.03319
Defense-in-Depth at the Perception-Reasoning Interface of LLM-Centric Agentic UAV Swarms
以LLM为中心的智能体无人机蜂群感知-推理接口处的纵深防御
Mohammadhossein Homaei, Yousef Emami, Sajad Homayoun, Rahim Taheri, Hao Zhou, Miguel Gutierrez Gaitan, Bo Wei
cs.CR · cs.MA
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) increasingly support Uncrewed Aerial Vehicle (UAV) swarm operations such as data collection scheduling, where the model reads structured sensor reports and decides which sensors to visit. An adversary who quietly manipulates those reports can redirect the swarm without modifying the model weights or the UAV. Defenses for this interface have been proposed architecturally but rarely implemented or evaluated. We implement and evaluate defense-in-depth at the perception-reasoning interface of LLM-Centric Agentic UAV Swarms. Five layers check the provenance of a report, whether its values are physically admissible, whether they agree with what swarm geometry and service history predict, whether the resulting schedule starves any sensor, and, when these fail, hand control to a deterministic scheduler that ignores the suspect input. We test each layer against an adversary strong enough to defeat the layer before it. For each of the three input-side layers, we derive in closed form how far a report can be distorted before that layer reacts, fixing each boundary from deployment parameters before any attack data is collected; across thirty matched simulation runs, predicted and measured boundaries agree. Separating attack detection from response is a well-established principle, and we quantify the cost of neglecting this distinction at the perception-reasoning interface. When the system rejects a report, it replaces it with the most recent accepted report. This prevents the adversary from controlling the UAV schedule, but it also increases cumulative cost by 79% and 74% for the two detectors, respectively, compared with the undefended system. The safety check does not detect any attacks, but it nevertheless reduces the attack-induced cost by 37.5%.
Chinese Translation
大语言模型(LLM)日益支持无人驾驶航空器(UAV)蜂群作业,例如数据采集调度,其中模型读取结构化传感器报告并决定要访问哪些传感器。一个悄悄篡改这些报告的对手无需修改模型权重或无人机,就能改变蜂群的行动方向。针对该接口的防御已在架构层面被提出,但很少被实现或评估。我们实现并评估了以LLM为中心的智能体无人机蜂群感知-推理接口处的纵深防御。五个层分别检查一份报告的来源、其数值是否在物理上可接受、它们是否与蜂群几何结构和服务历史所预测的一致、由此产生的调度是否会使任何传感器得不到服务,以及当这些检查失败时,将控制权交给一个忽略可疑输入的确定性调度器。我们用强到足以攻破其前一层防御的对手来测试每一层。对于三个输入侧的层中的每一层,我们以闭式形式推导出在该层做出反应之前一份报告可以被扭曲到何种程度,并在收集任何攻击数据之前根据部署参数确定每个边界;在三十次匹配的仿真运行中,预测边界与实测边界一致。将攻击检测与响应分离是一条已被确立的原则,而我们量化了在感知-推理接口处忽视这一区分所付出的代价。当系统拒绝一份报告时,它会用最近一份被接受的报告来替换该报告。这阻止了对手控制无人机调度,但与无防御系统相比,它也分别使两个检测器的累计成本增加了79%和74%。安全检查没有检测到任何攻击,但它仍将攻击导致的成本降低了37.5%。
cs.CR / 70 / 2610.03434
Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems
Persona Guardrail:面向智能体系统的生产级防御框架
Bijeeta Pal, Sridhar Reddy Maddireddy, Muhaimin Bin Munir, Zoltan Puha, Max Zhurovich, Adi Raghavendra, Sean Tout
cs.CR
large language model
大语言模型相关
Abstract
Large language model-based agents are increasingly deployed to perform domain-specific tasks by interacting with enterprise knowledge, tools, and external services. Existing runtime guardrails primarily target prompt injection and other attack-specific behaviors under a black-box threat model, but provide limited guarantees that agents operate within their intended functionality. As a result, production agents remain vulnerable to malicious requests and out-of-domain queries that existing defenses often fail to distinguish. We present Persona Guardrail, a production-grade runtime defense framework that enforces explicit functional boundaries for customer-facing agentic AI systems through synchronous input and output validation driven by semantic allowlist and blocklist specifications. We also introduce PAGE (Persona-Aware Guardrail Evaluation), a benchmark for evaluating function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns. Compared with a generic LLM-based guardrail, Persona Guardrail improves overall accuracy from 85.7% to 95.9%, increases out-of-domain detection from 57.3% to 93.5%, and reduces the false-approved rate from 25.0% to 4.7%. Currently deployed in production, Persona Guardrail meets its latency budget while sustaining a very low false-block and false-allow rate under realistic production workloads. These results demonstrate that Persona Guardrail provides a practical, scalable, and production-ready foundation for securing agentic AI systems.
Chinese Translation
基于大语言模型的智能体正越来越多地被部署,通过与企业知识、工具和外部服务交互来执行特定领域任务。现有的运行时护栏主要针对黑盒威胁模型下的提示注入及其他特定攻击行为,但无法充分保证智能体在其预期功能范围内运行。因此,生产环境中的智能体仍然容易受到恶意请求和域外查询的影响,而现有防御往往无法区分这些请求和查询。我们提出 Persona Guardrail,一个生产级运行时防御框架,它通过由语义允许列表和阻止列表规范驱动的同步输入与输出验证,为面向客户的智能体 AI 系统强制执行明确的功能边界。我们还引入了 PAGE(Persona-Aware Guardrail Evaluation,角色感知护栏评估),这是一个用于评估特定功能护栏的基准,涵盖用户轮次和智能体轮次上的良性、对抗性和域外交互。与通用的基于 LLM 的护栏相比,Persona Guardrail 将总体准确率从 85.7% 提升至 95.9%,将域外检测率从 57.3% 提升至 93.5%,并将误批准率从 25.0% 降低至 4.7%。目前已部署在生产环境中,Persona Guardrail 在满足其延迟预算的同时,在真实生产工作负载下保持了极低的误拦截率和误放行率。这些结果表明,Persona Guardrail 为保护智能体 AI 系统提供了一个实用、可扩展且生产就绪的基础。
cs.CR / 71 / 2610.03518
PrivDev: Mapping Static-Analysis Data Types to DPV
PrivDev:将静态分析数据类型映射到 DPV
Simon Bernbeck, Ricardo Ramalho, Matheus Amendoeira, Juliana Alves Pereira
cs.CR · cs.SE
large language model
大语言模型相关
Abstract
Static-analysis scanners can identify personal-data types in source code, but they lack mechanisms to connect these findings to standardized privacy vocabularies. PrivDev maps 122 Bearer CLI data types to Data Privacy Vocabulary Personal Data (DPV-PD) categories and links them to potentially relevant GDPR provisions. Our approach combines deterministic mapping for 43 exact-label matches with a retrieval-grounded Large Language Model (LLM) to resolve the remaining 79 non-trivial mappings. The resulting RDF knowledge graph contains 118 ODRL policy resources that were structurally validated using SHACL. The artifact passed five complementary validation gates that cover structural correctness, query consistency, retrieval quality, LLM-based assessment, and human evaluation. In human evaluation, nine annotators produced 711 judgments, yielding a raw agreement of 0.72, Gwet's AC1 of 0.68, and Gwet's AC2 of 0.88. Our results indicate that the proposed mappings are plausible and reproducible, while also revealing ambiguities in scanner-defined data-type labels and coverage gaps in DPV-PD.
Chinese Translation
静态分析扫描器可以识别源代码中的个人数据类型,但缺乏将这些发现连接到标准化隐私词汇表的机制。PrivDev 将 122 个 Bearer CLI 数据类型映射到数据隐私词汇表个人数据(DPV-PD)类别,并将它们链接到可能相关的 GDPR 条款。我们的方法将针对 43 个精确标签匹配的确定性映射与一个基于检索的大型语言模型(LLM)相结合,以解决其余 79 个非平凡映射。生成的 RDF 知识图谱包含 118 个 ODRL 策略资源,这些资源使用 SHACL 进行了结构验证。该制品通过了五个互补的验证关口,涵盖结构正确性、查询一致性、检索质量、基于 LLM 的评估和人工评估。在人工评估中,九名标注者产生了 711 个判断,得到 0.72 的原始一致性、0.68 的 Gwet's AC1 和 0.88 的 Gwet's AC2。我们的结果表明,所提出的映射是合理且可复现的,同时也揭示了扫描器定义的数据类型标签中的歧义以及 DPV-PD 中的覆盖缺口。
cs.AI / 72 / 2610.02513
From Fragments to Global Maps: Learning Vectorized Map Aggregation with Large Language Models
从碎片到全局地图:利用大语言模型学习矢量化地图聚合
Ziwei Li, Yi-Tang Chen, Xiaoqi Wang, Wenbin He, Han-Wei Shen, Liu Ren
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Large-scale vectorized HD maps provide structured road information that is essential for perception, localization, and planning in autonomous driving. Constructing such maps requires aggregating noisy, fragmented, and overlapping local predictions collected along a vehicle trajectory into a coherent global map. Existing aggregation methods typically rely on hand-crafted rules for fragment association and refinement. However, a fixed set of thresholds cannot effectively handle variations in road structures and prediction errors, often requiring detector-specific tuning or manual adjustment. To address this limitation, we propose MapMergeLLM, a data-driven framework that formulates vectorized map aggregation as conditional sequence generation with a large language model. Given serialized local vectorized maps, our model directly predicts the aggregated global map polylines. To reduce dependence on any particular upstream detector, we train the model on synthetic local maps generated from clean vector maps using corruptions that simulate representative prediction errors. We further introduce a coordinate tokenizer with geometry-aware pretraining to precisely represent map coordinates. In addition, we propose a line-level association loss that explicitly supervises correspondences between local observations of the same map element. Experiments on Argoverse2 and nuScenes using multiple recent upstream detectors demonstrate that MapMergeLLM substantially outperforms heuristic and optimization-based aggregation baselines without detector-specific retraining.
Chinese Translation
大规模矢量化高精地图提供了结构化的道路信息,这对于自动驾驶中的感知、定位与规划至关重要。构建此类地图需要将沿车辆轨迹采集到的含有噪声、碎片化且相互重叠的局部预测聚合为一张连贯的全局地图。现有的聚合方法通常依赖人工设计的规则来进行碎片关联与精化。然而,一组固定的阈值无法有效应对道路结构的变化与预测误差,往往需要针对特定检测器进行调参或人工调整。为解决这一局限,我们提出了 MapMergeLLM,这是一个数据驱动的框架,它将矢量化地图聚合形式化为基于大语言模型的条件序列生成任务。给定序列化的局部矢量化地图,我们的模型直接预测聚合后的全局地图折线。为减少对任何特定上游检测器的依赖,我们使用由干净矢量地图经模拟代表性预测误差的损坏过程生成的合成局部地图来训练该模型。我们进一步引入了一种具有几何感知预训练的坐标分词器,以精确表示地图坐标。此外,我们提出了一种线级关联损失,用于显式监督同一地图元素的各局部观测之间的对应关系。在 Argoverse2 和 nuScenes 上使用多个近期上游检测器进行的实验表明,MapMergeLLM 在无需针对特定检测器重新训练的情况下,显著优于基于启发式和基于优化的聚合基线方法。
cs.AI / 73 / 2610.02567
DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering
DAGS:用于时间稳定生成式渲染的冻结图像 DiT 的解耦外观与几何引导
Karthik Mohan Kumar, Damian Andrysiak, Pedro Antonio Pena, Kunal Tyagi, Rama Harihara
cs.CV · cs.AI · cs.GR · cs.LG
diffusion
扩散模型相关
Abstract
Diffusion transformers (DiTs) generate high-fidelity images from text and image conditions, but their outputs carry large variance and their faithfulness to a desired target depends heavily on how the condition is supplied. We present DAGS, a lightweight, attention-free, disentangled appearance and geometry conditioning scheme that steers a frozen image DiT to produce high-fidelity, highly faithful, and independently controllable renders. Two small convolutional encoders compute conditioning features once per frame and inject them as a learned, per-layer, element-wise residual into the image tokens, avoiding the quadratic cost of stacking conditions through attention. Because control and temporal handling live outside the frozen backbone, we retain its vast pretrained prior and eliminate backbone-overfitting risk. We further add a small recurrent lighting stabilizer and a training-free temporal guidance term that, coupled with our conditioning, elevate a per-frame image model into a streaming renderer. DAGS produces controllable, high-quality renders at a fraction of the compute of path tracing; it is not real-time, trading compute for controllability and quality. On a matched 1-spp + G-buffer input, per-frame DAGS reconstructs +8.6 dB / +10.1 dB PSNR over the real-time denoiser Intel OIDN and the diffusion renderer RGB<->X while being 2.5-8x more temporally stable perceptually (temporal-LPIPS flicker).
Chinese Translation
扩散 Transformer(DiT)能够根据文本和图像条件生成高保真图像,但其输出带有很大的方差,并且其对期望目标的忠实度在很大程度上取决于条件是何种方式提供的。我们提出 DAGS,一种轻量级、无注意力、解耦的外观与几何条件化方案,它引导一个冻结的图像 DiT 生成高保真、高度忠实且可独立控制的渲染结果。两个小型卷积编码器每帧计算一次条件特征,并将其作为学习得到的、逐层的、逐元素的残差注入到图像 token 中,从而避免了通过注意力堆叠条件所带来的二次开销。由于控制和时序处理都位于冻结主干之外,我们得以保留其庞大的预训练先验,并消除主干过拟合的风险。我们进一步加入了一个小型循环光照稳定器以及一个无需训练的时间引导项,它们与我们的条件化相结合,将逐帧图像模型提升为一个流式渲染器。DAGS 以路径追踪计算量的一小部分即可生成可控的高质量渲染;它不是实时的,而是以计算换取可控性和质量。在匹配的 1-spp + G-buffer 输入下,逐帧 DAGS 相比实时降噪器 Intel OIDN 和扩散渲染器 RGB<->X 重建出 +8.6 dB / +10.1 dB 的 PSNR,同时在感知上(temporal-LPIPS 闪烁)具有高出 2.5-8 倍的时间稳定性。
cs.AI / 74 / 2610.02753
Correcting Guided Diffusion Trajectories with Spectral Alignment
用谱对齐校正引导扩散轨迹
Gihoon Kim, Taesup Kim
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
The practical success of conditional image generation hinges on fine-grained differences in condition alignment and visual fidelity. Classifier-free guidance (CFG) is central to this success, but its lack of an explicit criterion makes it difficult to assess whether the guided trajectory is progressing as intended. To address this gap, we show that spectral alignment provides a principled criterion for understanding guidance behavior and improving guided diffusion sampling through adaptive correction. Our analysis identifies the spectra of intermediate states as an indicator of consistency with the expected spectral evolution of the forward process. Based on this observation, we introduce Spectral Correction Guidance, a method that corrects deviations from an analytic reference spectrum during sampling. The proposed method is training-free and applicable across diffusion backbones and conditional generation tasks without modifying the underlying model. Experiments demonstrate consistent gains in preference-based metrics over baseline guidance methods in text-to-image generation and improved generation quality over CFG on ImageNet. These improvements persist across a range of guidance scales and with fewer denoising steps. Our analyses and ablations provide insight into guidance behavior and how the proposed method affects generation quality.
Chinese Translation
条件图像生成的实际成功取决于条件对齐与视觉保真度方面的细粒度差异。无分类器引导(CFG)是这一成功的核心,但它缺乏明确的标准,这使得人们难以评估引导轨迹是否正按预期推进。为填补这一空白,我们表明,谱对齐提供了一个有原则的标准,用于理解引导行为,并通过自适应校正来改进引导扩散采样。我们的分析将中间状态的谱识别为与前向过程预期谱演化之间一致性的一个指标。基于这一观察,我们提出了谱校正引导(Spectral Correction Guidance),这是一种在采样过程中校正对解析参考谱之偏离的方法。所提方法无需训练,并且无需修改底层模型即可适用于各种扩散骨干网络和条件生成任务。实验表明,在文本到图像生成中,该方法相对于基线引导方法在基于偏好的指标上取得了一致的增益,并且在 ImageNet 上相对于 CFG 提升了生成质量。这些改进在一系列引导尺度下以及在使用更少去噪步骤时依然保持。我们的分析和消融实验为理解引导行为以及所提方法如何影响生成质量提供了洞见。
cs.AI / 75 / 2610.02887
Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking
通过因果不变掩码揭示 MLLMs 中的认知不确定性
Haoyang Luo, Linwei Tao, Jie Gui, Xinghao Chen, Chang Xu, Jianyuan Guo, Minjing Dong
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model's sensitivity to non-causal correlations, establishing its ability to capture MLLM's limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50% with comparable performance.
Chinese Translation
多模态大语言模型(MLLMs)受到幻觉问题的困扰,这使得不确定性量化(UQ)成为确保可靠部署的关键需求。然而,现有方法难以检测由表面关联引起的不确定性,尤其是当查询相关信号较弱时。我们主要将这一问题归因于它们偏向由数据模糊性引起的偶然不确定性,而忽视了源自模型局限性的认知不确定性。为了进一步分解不确定性类型以实现全面的 UQ,我们提出了因果不变掩码(CIM),它度量原始预测与以因果聚焦视图为条件的预测之间的语义偏移。基于这一框架,我们引入语义散度作为 UQ 的核心度量,并提供理论证据表明它收敛于模型对非因果相关性的敏感度的方差,从而确立了其捕捉 MLLM 局限性的能力。为了加速 MLLMs 中的 UQ,我们进一步提出了期望嵌入漂移(EED),这是一种快速的几何代理度量,直接在超球嵌入空间中估计语义偏移。实验表明,我们的方法在各种基准上取得了最先进的性能,而所提出的 EED 在性能相当的情况下加速了近 50%。
cs.LG / 76 / 2610.03154
Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models
物理是否存在于激活之中?在视频扩散模型中定位物理量
Jonas Kneifl, Jakub Skalski, Bartłomiej Twardowski, Kamil Deja
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar motion patterns. We address this by probing internal representations of video Diffusion Transformers (DiTs) for simulator-derived ground-truth physical quantities spanning kinematic motion and rigid-body dynamics under gravity and contact. We find that these quantities are linearly decodable with high accuracy early in the denoising process, substantially outperforming a baseline decoded directly from the model's own noised latents, indicating that the relevant physical information is actively constructed during denoising rather than already present in the input. Additionally, we show that activations at on-object tokens carry the relevant physical information and that quantities defined over multiple frames are readable from single latent frames. Hence, information is sharply localized within the token sequence and is computed globally but stored locally. The probes further show partial extrapolation, transferring to scene variations and object configurations outside their training regime, so what they read is not simply a correlate of the scenes they were fit on. When fitted directly in the full-resolution activation space, the probing directions can serve as steering vectors to change the model's output.
Chinese Translation
视频生成模型能够生成极其逼真的序列,并越来越多地被提议作为世界模型,然而近期的基准测试揭示出它们在物理推理方面存在明显的缺陷。这引出了一个问题:这些模型究竟是内化了物理原理,还是仅仅在复现熟悉的运动模式。我们通过探查视频扩散 Transformer(DiT)的内部表征来回答这一问题,探查的目标是来自仿真器的真值物理量,涵盖运动学运动以及重力与接触作用下的刚体动力学。我们发现,这些量在去噪过程的早期即可被高精度地线性解码,显著优于直接从模型自身的加噪潜变量中解码的基线,这表明相关的物理信息是在去噪过程中被主动构建出来的,而非已经存在于输入之中。此外,我们表明,位于物体上的 token 处的激活携带了相关的物理信息,并且定义在多个帧上的量可以从单个潜变量帧中读出。因此,信息在 token 序列中被高度精确地局部化,并且是全局计算但局部存储的。这些探针进一步表现出部分外推能力,能够迁移到其训练范围之外的场景变化和物体配置上,因此它们所读取的并不仅仅是其所拟合场景的一个相关量。当直接在全分辨率激活空间中进行拟合时,这些探查方向可以作为引导向量来改变模型的输出。
cs.AI / 77 / 2610.03224
Uncertainty as a Proxy for Semantic Correctness in Diffusion-Based Medical Image Synthesis
不确定性作为基于扩散的医学图像合成中语义正确性的代理指标
Yuxuan Ou, Konstantinos Kamnitsas, OxAAA Study, AICT Consortium, Regent Lee, Vicente Grau
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Diffusion models can synthesise contrast-enhanced CT (CECT) from non-contrast CT (NCCT), avoiding contrast administration and its environmental and patient-access costs. However, visually realistic images are not necessarily anatomically correct, and the pixel-intensity and feature-space similarity metrics used to assess generation quality do not directly measure anatomical correctness. In this work, we investigate whether uncertainty can serve as a proxy for semantic correctness in diffusion-based medical image synthesis. We study NCCT-to-CECT synthesis using AortaDiff, a multitask diffusion framework that jointly generates CECT images and lumen segmentations. The segmentation output provides an explicit representation of the generated vascular anatomy, enabling segmentation-derived errors to be used as a quantitative measure of generation correctness. Six methods spanning weight (Ensemble, HyperDiff, BayesDiff), architecture-perturbation (MCDropout), generative-stochasticity (RDS) and input-perturbation (TTA) uncertainty are compared at the pixel, region and image levels, and for detection of clinically relevant out-of-distribution (OOD) cases. Uncertainty proves informative at all three spatial scales, remains informative on an external multi-centre dataset under distribution shift, and supports OOD detection. MCDropout stands out among the six: it ranks among the leading methods at every scale, generalizes well on the external dataset, and can be enabled at inference on any model already trained with dropout, so reliable uncertainty comes at no extra training cost. Uncertainty reliably flags severe failures but discriminates poorly among already high-quality images. These findings support uncertainty as a practical and computationally economical signal for quality filtering, reliability assessment and OOD detection in NCCT-to CECT synthesis.
Chinese Translation
扩散模型能够从非增强CT(NCCT)合成对比增强CT(CECT),从而避免造影剂的使用及其在环境和患者可及性方面带来的成本。然而,视觉上逼真的图像并不一定在解剖学上正确,而用于评估生成质量的像素强度与特征空间相似度指标并不能直接衡量解剖学正确性。在本工作中,我们探究不确定性是否能够作为基于扩散的医学图像合成中语义正确性的代理指标。我们使用AortaDiff研究NCCT到CECT的合成,AortaDiff是一个多任务扩散框架,可联合生成CECT图像与管腔分割。分割输出为所生成的血管解剖结构提供了显式表示,使得由分割导出的误差可用作生成正确性的定量度量。我们比较了六种方法,涵盖权重不确定性(Ensemble、HyperDiff、BayesDiff)、架构扰动不确定性(MCDropout)、生成随机性不确定性(RDS)和输入扰动不确定性(TTA),在像素、区域和图像三个层面上进行比较,并用于检测临床相关的分布外(OOD)病例。不确定性在全部三个空间尺度上都被证明具有信息量,在分布偏移下的外部多中心数据集上仍保持信息量,并支持OOD检测。MCDropout在六种方法中脱颖而出:它在每一个尺度上都位列领先方法之中,在外部数据集上泛化良好,并且可以在任何已使用dropout训练的模型上于推理时启用,因此可靠的不确定性无需额外的训练成本。不确定性能够可靠地标记出严重失败,但在已经高质量的图像之间区分能力较差。这些发现支持将不确定性作为一种实用且计算经济的信号,用于NCCT到CECT合成中的质量筛选、可靠性评估和OOD检测。
cs.AI / 78 / 2610.03261
Consecutive Posterior Fusion for Diffusive Recovery of Unobservable Image Structures
用于不可观测图像结构扩散恢复的连续后验融合
Elena Morotti, Davide Evangelista, Elena Loli Piccolomini
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Solving severely ill-posed imaging inverse problems requires recovering image structures that are unobservable or weakly constrained by the measurements. Diffusion models provide expressive learned priors for inferring such missing information, while posterior sampling incorporates measurement consistency along the reverse process. Standard diffusion posterior samplers, however, rely on instantaneous measurement-aware estimates, without explicitly exploiting information carried by previous posterior corrections. We introduce Consecutive Posterior Fusion Denoising Diffusion Null-Space Models (CPF-DDNM), an inference-time strategy that fuses consecutive measurement-aware estimates to improve the diffusive recovery of unobservable image structures, without requiring retraining or additional denoiser evaluations. We instantiate this principle within DDNM, whose range/null-space decomposition reveals that consecutive fusion preserves the measurement-determined component while acting exclusively on the prior-driven null-space estimate. We thus provide a geometric interpretation of CPF-DDNM and a local error analysis that characterizes the optimal time-dependent fusion coefficient, including the extrapolative regime. Experiments on sparse-view and simulated low-dose computed tomography, as well as medical image super-resolution, show consistent improvements over DDNM and competitive performance against diffusion-based inverse solvers.
Chinese Translation
求解严重病态的成像逆问题需要恢复那些不可观测或受测量弱约束的图像结构。扩散模型为推断此类缺失信息提供了具有表现力的学习先验,而后验采样则在反向过程中融入测量一致性。然而,标准扩散后验采样器依赖于瞬时的测量感知估计,而没有显式利用先前后验校正所携带的信息。我们引入了连续后验融合去噪扩散零空间模型(CPF-DDNM),这是一种推理时策略,它融合连续的测量感知估计,以改进不可观测图像结构的扩散恢复,且无需重新训练或额外的去噪器评估。我们在 DDNM 中实例化了这一原则,其值域/零空间分解表明,连续融合保留了由测量确定的成分,同时仅作用于由先验驱动的零空间估计。因此,我们提供了 CPF-DDNM 的几何解释,以及一种局部误差分析,该分析刻画了最优的随时间变化的融合系数,包括外推情形。在稀疏视角和模拟低剂量计算机断层扫描以及医学图像超分辨率上的实验表明,相较于 DDNM 有一致的改进,并且与基于扩散的逆问题求解器相比具有竞争力的性能。
cs.AI / 79 / 2610.03577
Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection
重新思考少步扩散 Transformer 中应缓存什么:求解器感知的目标选择
Shuo Yang, Lihao Fang, Yi Zhang, Haixiang Wang, Xincheng Ye, Shufan Chen, Jipeng Guo, Youqing Wang
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tensor computed at an earlier step. Most caching methods decide in advance which tensor to reuse. After distillation, adjacent sampling steps are farther apart. Reusing a tensor across this larger gap introduces more error, so choosing what to cache becomes especially important. We therefore introduce AutoTarget, a method that chooses the cached tensor for a given model, solver, and reuse schedule. AutoTarget uses a small set of runs without cache reuse to measure the error caused by reusing each candidate tensor, then selects the candidate with the lowest error. We also analyze how an error at one reuse step affects the final sample. For Euler sampling, we identify cache targets that produce the same trajectory and show why a stored solver update may not. Experiments on distilled image and video DiTs show that the best cache target changes with the model, image resolution, and solver. AutoTarget reduces DiT evaluations and retained cache storage. Generation quality remains close to the corresponding uncached run. On the tested PixArt-LCM and FLUX.1-schnell settings, its calibration ranking matches the ranking from held-out cached runs. To help others reproduce the method, we provide its core implementation on GitHub at https://github.com/wali1024-offical/AutoTarget.
Chinese Translation
扩散 Transformer(DiT)能够生成高质量的图像和视频,但生成每个样本都需要多次代价高昂的 DiT 前向传播。加速 DiT 采样的两种常见方法是步数蒸馏,它减少采样步数;以及缓存,它通过复用较早步骤计算得到的张量来跳过一些 DiT 评估。大多数缓存方法会预先决定要复用哪个张量。在蒸馏之后,相邻采样步骤之间的间隔更大。在更大的间隔上复用张量会引入更多误差,因此选择缓存什么变得尤为重要。因此,我们提出了 AutoTarget,一种针对给定模型、求解器和复用调度选择被缓存张量的方法。AutoTarget 使用一小组不进行缓存复用的运行来测量复用每个候选张量所造成的误差,然后选择误差最低的候选。我们还分析了一个复用步骤处的误差如何影响最终样本。对于 Euler 采样,我们识别出能够产生相同轨迹的缓存目标,并说明为什么存储的求解器更新可能无法做到这一点。在蒸馏的图像和视频 DiT 上的实验表明,最佳缓存目标会随模型、图像分辨率和求解器而变化。AutoTarget 减少了 DiT 评估次数和保留的缓存存储量。生成质量仍然接近相应的未缓存运行。在测试的 PixArt-LCM 和 FLUX.1-schnell 设置上,其校准排名与留出的缓存运行的排名相匹配。为了帮助他人复现该方法,我们在 GitHub 上提供了其核心实现,地址为 https://github.com/wali1024-offical/AutoTarget。
cs.AI / 80 / 2610.02369
Automating the Application of HCI Principles: Skills for On-Demand UI Construction, the Human-AI Space to Think, and the Future of HCI
自动化 HCI 原则的应用:按需 UI 构建的技能、人机协同的思考空间,以及 HCI 的未来
Nathan Conklin, Miranda Capra, Chris North
cs.HC · cs.AI
large language model
大语言模型相关
Abstract
Human-computer interaction (HCI) is in the middle of a transition: large language models can now generate functional user interfaces (UIs) on demand from natural-language task descriptions. A user explains what they are trying to accomplish, and the system materializes a working interface to support it. This capability already exists in systems such as Claude and ChatGPT and continues to grow in fidelity as the underlying models improve. The next step along this trajectory is to move from interfaces that are merely generated to interfaces that are generated well. We propose a framework in which the dialogue between user and artificial intelligence (AI) becomes a Space to Think: a shared, structured cognitive workspace in which task decomposition produces an on-demand user interface as an extension of the user's thinking rather than as a separate artifact. Within this paradigm, classical HCI design knowledge (Nielsen's heuristics, Norman's affordance prescriptions, Web Content Accessibility Guidelines (WCAG) success criteria, cognitive-load constraints, and mixed-initiative principles) is encoded as skills: machine-readable skill.md files that the generating agent loads at runtime as software engineering tools. Skills turn HCI design knowledge into declarative, inspectable, version-controlled, and editable artifacts owned by the HCI community itself so that accessibility, learnability, and consistency become properties of a generative process rather than properties of a finished product. We outline a research agenda depicting a future where the HCI field transitions from today's design and knowledge heuristic checklist towards a future where the craft becomes machine-readable, executable, and open.
Chinese Translation
人机交互(HCI)正处于一场转型之中:大语言模型如今能够根据自然语言的任务描述,按需生成功能性的用户界面(UI)。用户说明自己想要完成什么,系统便会将其物化为一个可用的界面来支持这一目标。这种能力已经存在于 Claude 和 ChatGPT 等系统之中,并且随着底层模型的不断改进,其保真度还在持续提升。沿着这一轨迹的下一步,是从仅仅被生成的界面,迈向被良好生成的界面。我们提出一个框架,其中用户与人工智能(AI)之间的对话成为一个思考空间(Space to Think):一个共享的、结构化的认知工作空间,在其中,任务分解所产生的是一个按需的用户界面,它作为用户思维的延伸,而非一个独立的制品。在这一范式内,经典 HCI 设计知识(Nielsen 的启发式原则、Norman 的可供性规定、Web 内容无障碍指南(WCAG)成功标准、认知负荷约束,以及混合主动原则)被编码为技能:即机器可读的 skill.md 文件,生成代理在运行时将其作为软件工程工具加载。技能将 HCI 设计知识转化为声明式的、可检查的、受版本控制的、可编辑的制品,并由 HCI 共同体自身所拥有,从而使可访问性、可学习性和一致性成为生成过程的属性,而非最终产品的属性。我们勾勒出一项研究议程,描绘了这样一幅未来图景:HCI 领域从当今的设计与知识启发式清单,转向一门技艺变得机器可读、可执行且开放的未来。
cs.LG / 81 / 2610.02324
Slow-Fast Multi-Teacher On-Policy Distillation for Capability Preservation
用于能力保持的慢-快多教师在线策略蒸馏
Xiaofei Yin, Tong Chu, Jiyuan Fu, Jun Lan, Shuheng Zhou, Huijia Zhu
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gradually drives the student away from its initialization model, and general capabilities decline as the displacement grows, resulting in capability interference. A direct remedy is constraining the student toward its initialization, but this suppresses the acquisition of domain expertise as well. We propose Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model, the current student updated directly by each teacher, with a slow model, an exponential moving average of the student. The slow model absorbs the learning signal gradually, serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, while retaining aligned and orthogonal components. Experiments across multiple model scales demonstrate that SF-MOPD effectively mitigates capability interference, enhances specialized multimodal capabilities, and reduces the average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
Chinese Translation
基础多模态大语言模型旨在支持跨越多样领域的广泛能力。多教师在线策略蒸馏(MOPD)为将领域特定专业知识整合到单个学生模型中提供了一个有效框架。然而,MOPD 训练会逐渐使学生模型偏离其初始化模型,并且随着偏移增大,通用能力下降,从而导致能力干扰。一种直接的补救措施是将学生模型约束向初始化模型,但这同样会抑制领域专业知识的获取。我们提出慢-快多教师在线策略蒸馏(SF-MOPD),它将一个快速模型——即由每个教师直接更新的当前学生模型——与一个慢速模型——即学生模型的指数移动平均——耦合起来。慢速模型逐渐吸收学习信号,作为一个移动的能力参考,将通用基础与已确认的领域专业知识融合起来。对于每个教师,SF-MOPD 在对数概率空间中计算教师引起的更新,并且只移除将快速模型进一步推离慢速模型的分量,同时保留对齐分量和正交分量。跨多个模型规模的实验表明,SF-MOPD 有效缓解能力干扰,增强专业化的多模态能力,并降低通用能力基准上的平均退化,持续优于原始 MOPD。
cs.LG / 82 / 2610.02353
Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation
每个用户都需要一个私有 LoRA 吗?将个性化与每用户适配解耦
Songyuan Sui, Srikanth Malla, Chiho Choi, Joon Hee Choi
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Personalized large language models often require a complete adaptation state for each user. However, this paradigm scales poorly as the user population grows. We revisit this design through the lens of personalization capacity allocation: how much adaptation capacity can be shared across users, how the shared capacity should be composed, and how much must remain user-specific. We answer them through three complementary empirical analyses. We find that independent user adapters contain substantial cross-user reusable structure, that the utility of reusable directions reflects both user relevance and variation across queries, and that user histories provide transferable signals for compact individual correction. Motivated by these findings, we propose LINEUP. It learns a bank of reusable low-rank personalization factors, composes them through user-conditioned recall and query-dependent calibration, and restricts target-user adaptation to a tiny user code over a shared correction space. This design decouples expressive personalization capacity from per-user trainable state. Each target user optimizes only eight scalars, while all shared components remain fixed. By comparison, the evaluated private-LoRA configuration uses 4.19 million per-user parameters. Our theoretical analysis gives a finite-step, finite-history risk bound and sufficient conditions for user-code refinement to improve on history initialization. Across six tasks spanning personalized classification, prediction, and generation, LINEUP leads on all 12 metrics, each averaged over three independent runs (e.g., reducing LaMP-3 RMSE by 11.4% relative to the strongest baseline). It maintains advantages under limited history. These results show that rich personalization can be supported primarily by reusable, conditionally composed shared capacity, while independent user adaptation remains confined to a tiny correction state.
Chinese Translation
个性化大语言模型通常需要为每个用户提供完整的适配状态。然而,随着用户规模增长,这种范式的扩展性很差。我们从个性化容量分配的角度重新审视这一设计:有多少适配容量可以在用户之间共享,共享容量应如何组成,以及有多少必须保持用户特定。我们通过三项互补的实证分析来回答这些问题。我们发现,独立的用户适配器包含大量跨用户可复用结构,可复用方向的效用既反映用户相关性,也反映跨查询的变化,并且用户历史为紧凑的个体校正提供了可迁移信号。受这些发现的启发,我们提出了 LINEUP。它学习一个可复用低秩个性化因子库,通过以用户为条件的召回和依赖查询的校准来组合它们,并将目标用户适配限制为共享校正空间上的一个微小用户代码。这种设计将富有表现力的个性化容量与每用户可训练状态解耦。每个目标用户仅优化八个标量,而所有共享组件保持固定。相比之下,所评估的私有 LoRA 配置使用 419 万每用户参数。我们的理论分析给出了有限步、有限历史的风险界,以及用户代码细化相较于历史初始化有所改进的充分条件。在涵盖个性化分类、预测和生成的六个任务中,LINEUP 在所有 12 个指标上领先,每个指标均为三次独立运行的平均值(例如,相对于最强基线将 LaMP-3 RMSE 降低 11.4%)。它在有限历史下保持优势。这些结果表明,丰富的个性化主要可由可复用的、条件组合的共享容量支持,而独立的用户适配仍限于一个微小的校正状态。
cs.LG / 83 / 2610.02355
Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
为什么自适应批处理有助于LLM预训练?一个来自无界方差的视角
Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi
cs.LG · math.OC
large language model
大语言模型相关
Abstract
Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence suggests that this assumption fails in many practical nonconvex problems. The Blum--Gladyshev (BG-$0$) noise model relaxes this assumption by allowing the variance to grow quadratically with the distance from initialization, suggesting that batch size schedulers can help by controlling the variance growth during training. However, this growth can be overly conservative in practice. We empirically investigate variance growth in LLM pretraining and observe that a generalized BG model with a tunable growth exponent provides a tighter description of practical noise behavior. Motivated by this observation, we introduce the generalized BG-$a$ noise model, which interpolates between bounded variance ($a=0$) and BG-$0$ noise ($a=2$). Under $L$-smoothness, we derive an information-theoretic lower bound with growth-dependent oracle complexity $Ω(ε^{-(4+a)})$ and establish a matching upper bound in $ε$-dependence by increasing the batch size as the iterates move away from initialization. Finally, we propose an adaptive batch scheduler that controls variance growth through dynamic batch size adjustments during training. In pretraining OLMo2 models of up to 1B parameters on C4, our scheduler achieves a lower validation loss than both small and large batch training under matched token budgets, while using less than 10\% of the iterations of small batch training.
Chinese Translation
在训练过程中增大批量大小是大语言模型(LLM)预训练中的一种常见做法,然而其成功背后的理论依据尚未得到很好的理解。对随机优化的分析通常假设随机梯度方差是一致有界的,但近期的证据表明,这一假设在许多实际的非凸问题中并不成立。Blum--Gladyshev(BG-$0$)噪声模型通过允许方差随与初始化点距离的增大而二次增长,放宽了这一假设,这表明批量大小调度器可以通过控制训练过程中的方差增长来发挥作用。然而,这种增长在实际中可能过于保守。我们对LLM预训练中的方差增长进行了实证研究,并观察到,一个具有可调增长指数的广义BG模型能够更紧致地刻画实际噪声行为。受这一观察的启发,我们引入了广义BG-$a$噪声模型,它在有界方差($a=0$)与BG-$0$噪声($a=2$)之间进行插值。在$L$-光滑性条件下,我们推导出一个信息论下界,其预言机复杂度为依赖于增长的$Ω(ε^{-(4+a)})$,并通过在迭代点远离初始化点时增大批量大小,建立了在$ε$-依赖意义上与之匹配的上界。最后,我们提出了一种自适应批量调度器,它通过在训练过程中动态调整批量大小来控制方差增长。在C4上预训练参数量最高达1B的OLMo2模型时,在匹配的token预算下,我们的调度器比小批量训练和大批量训练都取得了更低的验证损失,同时所用的迭代次数不到小批量训练的10\%。
cs.LG / 84 / 2610.02391
Hesitation Has a Geometry: Entropy-Trained Hyperbolic Probes for Sparse Activation Steering
犹豫具有几何:用于稀疏激活引导的熵训练双曲探针
Zeyong Zhang, Tung Sum Thomas Kwok, Tengfei Ma, Mengjia Xu
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
When a large language model solves a mathematical problem, its reasoning is largely hierarchical, and the solution often branches at a few tokens where the next-token entropy is high. Such tree-like structure embeds in hyperbolic space with far lower distortion than in Euclidean space. Activation steering, however, usually edits the hidden states of a pretrained model by adding one fixed Euclidean vector at every token, even though most tokens of a solution are already determined by the context. We propose Hyperbolic Entropy Steering (HEST), which embeds the hidden states in the Poincaré ball with a lightweight probe whose only label is the model's own next-token entropy. Where this entropy exceeds a threshold, HEST moves the embedded state along the geodesic of steepest descent of a readout of the probe and maps the change back to the hidden state. For the Busemann readout of a learned ideal point, we prove that a step of fixed length lowers it by the same amount at every state. On three instruction-tuned models from the Qwen2.5-Math and Llama-3.1 families, HEST with the Busemann readout improves greedy accuracy on MATH-500 and GSM8K in five of six settings, by up to 1.8 points, whereas a contrastive steering vector added at every token lowers accuracy. With a Euclidean probe trained in the same way, this gain disappears on Qwen2.5-Math-1.5B-Instruct. The gains are largest on problems where the model hesitates often, and accuracy on the remaining problems is almost unchanged.
Chinese Translation
当大型语言模型解决数学问题时,其推理在很大程度上是分层的,而解通常会在少数几个下一词元熵较高的词元处发生分支。这种树状结构嵌入双曲空间时,其失真远低于嵌入欧几里得空间。然而,激活引导通常通过在每个词元处添加一个固定的欧几里得向量来编辑预训练模型的隐藏状态,即使一个解的绝大多数词元已经由上下文确定。我们提出双曲熵引导(Hyperbolic Entropy Steering, HEST),它用一个轻量级探针将隐藏状态嵌入庞加莱球中,而该探针的唯一标签是模型自身的下一词元熵。在该熵超过阈值之处,HEST 沿着探针某读出量的最速下降测地线移动嵌入状态,并将该变化映射回隐藏状态。对于学习到的理想点的 Busemann 读出,我们证明固定长度的一步在每个状态下都会将其降低相同的量。在来自 Qwen2.5-Math 和 Llama-3.1 系列的三个指令微调模型上,使用 Busemann 读出的 HEST 在六种设置中的五种里提升了 MATH-500 和 GSM8K 上的贪心准确率,提升最高达 1.8 个百分点,而在每个词元处添加的对比式引导向量则会降低准确率。当使用以相同方式训练的欧几里得探针时,这一增益在 Qwen2.5-Math-1.5B-Instruct 上消失。这些增益在模型经常犹豫的问题上最大,而在其余问题上的准确率几乎不变。
cs.LG / 85 / 2610.02396
Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance
Inherit-MAS:通过工作流与执行继承实现多智能体系统的测试时演化
Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-executing unchanged requests can incur redundant computation. Inspired by the interplay of inheritance and selection in biological evolution, we introduce Inherit-MAS, which makes inheritance explicit at the workflow and execution levels. A meta-model first synthesizes a workflow of worker agents with declared roles, communication inputs, and tool permissions, and a separately prompted judge scores each executed candidate and diagnoses its deficiencies. In ordinary refinement rounds, \emph{workflow inheritance} starts from the latest completed candidate, may discard removable nodes judged unhelpful, and applies a validated edit to address the diagnosed deficiency. When the new candidate executes, \emph{execution inheritance} inherits eligible stored results only if the complete resolved request and execution context match, avoiding redundant model and tool calls. With GPT-4o-mini workers, Inherit-MAS achieves 55.4\% completion on WorkBench and 49.7\% joint F1 on HotpotQA FullWiki, outperforming EvoAgent, EvoMAS, and TacoMAS. With Qwen3-32B workers, it also exceeds these evolving-MAS baselines on both benchmarks. Compared with rerunning the same controller with execution inheritance disabled, execution inheritance reduces worker-token usage by 29.1\% on WorkBench and 34.6\% on HotpotQA, and total token usage by 5.3\% and 18.1\%.
Chinese Translation
由大语言模型构建的多智能体系统(MAS)协调专门的智能体来处理复杂任务,但有效的工作流难以事先设计。测试时演化利用执行反馈来改进工作流,但大范围修订可能干扰有用的组件,而重新执行未更改的请求会带来冗余计算。受生物演化中继承与选择相互作用的启发,我们提出 Inherit-MAS,它在工作流和执行两个层面将继承显式化。一个元模型首先综合出一个由工作智能体组成的工作流,这些工作智能体具有声明的角色、通信输入和工具权限,而一个单独提示的评判者会对每个执行后的候选方案评分并诊断其缺陷。在普通改进轮次中,\emph{工作流继承}从最新完成的候选方案开始,可能丢弃被判定为无帮助的可移除节点,并应用一个经过验证的编辑来解决所诊断出的缺陷。当新候选方案执行时,\emph{执行继承}仅当完整解析后的请求和执行上下文匹配时,才继承符合条件的已存储结果,从而避免冗余的模型和工具调用。在使用 GPT-4o-mini 工作智能体时,Inherit-MAS 在 WorkBench 上达到 55.4\% 的完成率,在 HotpotQA FullWiki 上达到 49.7\% 的联合 F1,优于 EvoAgent、EvoMAS 和 TacoMAS。在使用 Qwen3-32B 工作智能体时,它在两个基准上也超过了这些演化式 MAS 基线。与在禁用执行继承的情况下重新运行相同控制器相比,执行继承在 WorkBench 上将工作智能体 token 使用量减少 29.1\%,在 HotpotQA 上减少 34.6\%,并将总 token 使用量分别减少 5.3\% 和 18.1\%。
cs.LG / 86 / 2610.02438
Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval
你是在合成还是在回忆?评估大语言模型在算法代码检索上的表现
Nickil Maveli, Antonio Vergari, Shay B. Cohen
cs.LG · cs.AI · cs.CL · cs.PL
large language model
大语言模型相关
Abstract
Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textit{parametric code retrieval}: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to evaluate this capability in isolation, and assess 15 models (7B--34B parameters) in a zero-shot setting. We find substantial variation in retrieval accuracy across languages and input representations, even for widely documented algorithms and show that prompt augmentation with retrieved code snippets or structured algorithmic hints improve accuracy on complex algorithms, while SFT achieves broader language gains and GRPO achieves larger per-language gains on specific languages. Together, our results establish parametric code retrieval as a distinct, measurable capability and caution against deploying AI-generated algorithmic code without systematic validation.\footnote{Code and dataset are available at https://github.com/Nickil21/AlgoREval
Chinese Translation
大型语言模型(LLMs)在代码生成方面已表现出强大性能,其成功既取决于回忆相关算法知识,也取决于推理如何应用这些知识。然而,现有的 LLM 流程是不透明的,并未在这两个组成部分之间做出明确区分。我们认为,对于其规范实现广泛存在于预训练语料中的著名算法而言,代码生成更适合被衡量为 \textit{参数化代码检索}:从内化知识中复现一个指定名称的算法,而不是合成一个新算法。我们引入 AlgoREval,这是一个包含 599 个问题的基准,涵盖 14 个领域的 77 个经典算法、7 种编程语言和 4 种图输入表示,以单独评估这一能力,并在零样本设置下评估 15 个模型(7B--34B 参数)。我们发现,即使对于记录广泛的算法,检索准确率在不同语言和输入表示之间也存在显著差异,并表明,使用检索到的代码片段或结构化算法提示进行提示增强,能够提高复杂算法上的准确率,而 SFT 实现了更广泛的语言收益,GRPO 在特定语言上实现了更大的单语言收益。总之,我们的结果确立了参数化代码检索是一种独特且可衡量的能力,并警示不要在没有系统验证的情况下部署 AI 生成的算法代码。\footnote{代码和数据集可在 https://github.com/Nickil21/AlgoREval 获取}
cs.LG / 87 / 2610.02442
AI-driven Thermal-aware Data Center Capacity Planning
AI驱动的热感知数据中心容量规划
Yixing Li, Mark Fenton, Matthew Kaufeler, Ka Ming Leung, Xin Ai, Zhiyu Zeng
cs.LG
large language model
大语言模型相关
Abstract
The emerging of large language models (LLMs) has posed significant challenges to the thermal management of data center. Intense GPU computation for LLMs results in localized hotspots. Moreover, spiking thermal loads during training and inference bursts make real-time cooling response more difficult to predict and control. Thermal-aware capacity planning of data center requires massive expensive high-fidelity CFD simulations. AI models can perform real-time prediction for unseen designs. However, existing works either have large prediction error, or have over-simplified assumptions for data center operations. This work presents an AI-driven framework that can perform thermal-aware capacity planning for a real-world data center in seconds. The embedded AI model learns from numerous key parameters (rack power, server power, server placement, HVAC settings etc.), and provides temperature prediction within milliseconds. This AI model is tested against high-fidelity CFD simulations, and results show that for unseen data center designs, model can achieve high accuracy with 10000X speedup. Driven by the AI model, the authors design the thermal-aware capacity planning framework. This framework can help data center designers and operators instantaneously optimize both workload distribution and HVAC cooling efficiency.
Chinese Translation
大型语言模型(LLMs)的出现给数据中心的热管理带来了重大挑战。用于LLMs的密集GPU计算会导致局部热点。此外,训练和推理突发期间激增的热负载使实时冷却响应更难以预测和控制。数据中心的热感知容量规划需要大量昂贵的高保真CFD仿真。AI模型可以对未见过的设计进行实时预测。然而,现有工作要么具有较大的预测误差,要么对数据中心运行做出了过度简化的假设。这项工作提出了一个AI驱动的框架,可以在几秒内为真实世界数据中心执行热感知容量规划。嵌入的AI模型从众多关键参数(机架功率、服务器功率、服务器放置、HVAC设置等)中学习,并在毫秒内提供温度预测。该AI模型针对高保真CFD仿真进行了测试,结果表明,对于未见过的数据中心设计,模型可以实现高精度并具有10000倍的加速。在该AI模型的驱动下,作者设计了热感知容量规划框架。该框架可以帮助数据中心设计人员和运维人员即时同时优化工作负载分布和HVAC冷却效率。
cs.LG / 88 / 2610.02593
Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models
面向大型语言模型持续预训练的 Fisher 引导子模数据选择
Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan
cs.LG
large language model
大语言模型相关
Abstract
Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity, or mitigates forgetting by spending many extra general-domain replay tokens. Neither strategy directly asks how training on a candidate will move the model parameters. We show that loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to, while leaving low-Fisher coordinates largely untouched. This asymmetry exposes a parameter-space mechanism for catastrophic forgetting. Motivated by this observation, we propose a Fisher-aware CPT selector that decomposes each candidate's gradient into an anchor component, which measures perturbation along committed parameter directions, and a frontier component, which measures update capacity in unconstrained low-Fisher subspaces. We aggregate these signals with a log-determinant submodular objective and optimize it in a single pass using a scalable streaming data selection pipeline. On TinyLlama-1.1B and Llama-3.1-8B CPT over medical data, our selector improves target-domain quality while bounding forgetting on held-out pretraining benchmarks. Most importantly, it is substantially more token-efficient than forgetting-aware replay. 1B selected tokens already outperform the replay strategy trained with 10B tokens on both adaptation and forgetting, giving a 10x token-efficiency advantage.
Chinese Translation
数据选择已经成为大型语言模型训练中的一个核心瓶颈,因为网络规模语料库有噪声且 token 预算有限。在持续预训练(CPT)中,它变成了一个遗忘控制问题:选择不当的目标领域语料库可能会覆盖预训练检查点中编码的能力。现有 CPT 实践要么用诸如困惑度之类与参数无关的标量对候选进行评分,要么通过花费许多额外的通用领域回放 token 来缓解遗忘。这两种策略都没有直接询问在某个候选上训练将如何移动模型参数。我们表明,基于损失的选择会导致 CPT 后的 Fisher 对角线在预训练模型已经固守的高 Fisher 坐标上向下漂移,而低 Fisher 坐标则基本不受影响。这种不对称性揭示了一种灾难性遗忘的参数空间机制。受此观察启发,我们提出一种 Fisher 感知的 CPT 选择器,它将每个候选的梯度分解为一个锚定分量,该分量衡量沿已固守参数方向的扰动,以及一个前沿分量,该分量衡量在无约束低 Fisher 子空间中的更新能力。我们使用对数行列式子模目标聚合这些信号,并使用可扩展的流式数据选择流水线在单次遍历中对其进行优化。在医学数据上的 TinyLlama-1.1B 和 Llama-3.1-8B CPT 中,我们的选择器提高了目标领域质量,同时在留出的预训练基准上限制了遗忘。最重要的是,它在 token 效率上显著高于遗忘感知回放。1B 个选定 token 在适应和遗忘两方面已经优于用 10B 个 token 训练的回放策略,带来了 10 倍的 token 效率优势。
cs.LG / 89 / 2610.02614
What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives
后训练中丢失了什么?默认坍缩与跨多样视角的上下文内可引导性的丧失
Jessica Dierking, Itai Shapira, Niclas Boehmer
cs.LG
large language model
大语言模型相关
Abstract
AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.
Chinese Translation
服务于异质人群的 AI 模型必须依据适合每个用户和情境的原则行事。尽管后训练已被证明会缩窄大语言模型所表达的观点范围,但先前工作关注的是默认行为,而非适应上下文内信息的能力。我们表明,后训练还会降低模型在上下文内被引导至其未被训练偏好的视角的能力。在受控实验中,我们将模型微调至文化价值分歧中的某一方,并在整个训练过程中评估检查点。在日常使用中,被训练的一方变得越来越占主导,而识别并忠实体现对立观点的能力则下降。这些发现指向一种张力:在优先考虑单一价值集与保留服务多元利益相关者所需的技术能力之间存在张力。最后,我们提出并分析了一种替代目标,它在表达视角上服从规定分布的前提下最大化奖励,并将立场分布匹配作为一种实际实现加以呈现。
cs.LG / 90 / 2610.02632
Online Verification of Language Model Responses Under Cost Constraints
成本约束下语言模型响应的在线验证
Erfan Hajihashemi, Yanning Shen
cs.LG
large language model
大语言模型相关
Abstract
As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model outputs is often done by querying a costly ground-truth oracle, which is impractical to invoke at every step in an online setting. Prior work addresses this by querying a single weak verifier on every step, and using its score to decide whether the costly strong verifier needs to be queried as well, reserving strong verification for only a small fraction of the steps. However, a single fixed weak verifier may not perform consistently well as the subject matter or difficulty of incoming queries changes over time, and committing to one in advance risks either overly costly or inaccurate verification. We introduce OMVV (Online Multi-Verifier Verification), an algorithm that maintains a pool of $K$ candidate weak verifiers with differing cost and verification performance, and adaptively routes each round's decision to a verifier selected via an online score combiner and an exponential-weights routing policy. OMVV provides a distribution-free, finite-time guarantee on false-accept and false-reject rates across the full pool of verifiers, and further achieves sublinear regret against the best fixed verifier in hindsight under a combined cost and consistency objective. Experiments on reasoning dataset benchmarks show that OMVV achieves higher accuracy at lower verification cost than any single fixed verifier, across a range of operating budgets.
Chinese Translation
随着大语言模型越来越多地被部署用于多步推理,验证其输出的正确性对于在大规模下保持可靠性已变得至关重要。验证大语言模型输出的正确性通常通过查询一个代价高昂的真值预言机来完成,而在在线场景下,每一步都调用它是不切实际的。先前的工作通过如下方式解决这一问题:在每一步查询单个弱验证器,并利用其得分来决定是否还需要查询代价高昂的强验证器,从而仅将强验证保留给一小部分步骤。然而,随着到来的查询的主题或难度随时间变化,单个固定的弱验证器可能无法始终表现良好,而事先固定选择其中一个,则会带来验证代价过高或验证不准确的风险。我们提出 OMVV(在线多验证器验证,Online Multi-Verifier Verification),这是一种维护由 $K$ 个在代价和验证性能上各不相同的候选弱验证器组成的池的算法,并通过在线得分组合器和指数权重路由策略选择验证器,自适应地将每一轮的决策路由到所选验证器。OMVV 在全部验证器池上对错误接受率和错误拒绝率提供了无分布假设的有限时间保证,并且在代价与一致性的联合目标下,相对于事后最优的固定验证器进一步实现了次线性遗憾。在推理数据集基准上的实验表明,在一系列运行预算下,OMVV 相比任何单一固定验证器都能以更低的验证代价实现更高的准确率。
cs.LG / 91 / 2610.02657
Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs
上下文塔转换保持生成能力,而冻结保留知识:MoE 大语言模型的低预算 AR 到扩散转换
Wentao Lu, Jesse Clark, Tianyu Zhu
cs.LG
diffusion
扩散模型相关
Abstract
Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent's GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion's scores in both directions across tasks, while its AR parent's scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent's generation performance than in-place conversion.
Chinese Translation
将预训练的自回归(AR)模型转换为扩散语言模型(dLLM),能够在不预训练新模型的情况下实现并行生成。已发表的转换方法在训练数据上相差约三个数量级,并且尚未在统一协议下进行比较。我们在固定语料、监督 token 预算、可训练参数集合和评测框架的条件下,比较了对同一个 30B 混合专家(MoE)父模型所做的两种转换,每种转换各自采用其自身的训练配方。原地模型使用去噪损失与表示对齐损失来更新父模型权重的一个子集;而冻结塔模型则改为通过交叉注意力,以父模型的一个冻结因果副本为条件。在 1B 训练 token 下,冻结塔模型在 HumanEval pass@10 上得分为 71.60,而原地模型为 6.19,提升了 11.6 倍。在相同预算下,它还保持了父模型 GSM8K 得分的 95% 以及其 MMLU-Pro 得分的 99%。一项稠密父模型实验复现了这一 HumanEval 上的差距。在约 500M token 的双塔设计中,冻结上下文塔比训练它保留了显著更多的 MMLU-Pro 性能,而两者给出的观测 HumanEval 得分相近。我们的理论分析证明,在硬注意力掩码以及每轮从左到右确定一个位置的承诺机制下,这两类转换都包含针对 AR 父模型的精确采样器。在共享损失下,冻结会消除经由上下文状态的梯度贡献。此外,评测协议会在不同任务上以两个方向显著影响一个已发表的 500B token 转换的得分,而其 AR 父模型的得分变化不到三个点,因此比较 dLLM 需要统一的协议。这些结果表明,在所测试的低预算情形下,冻结塔配置比原地转换保留了显著更多的父模型生成性能。
cs.LG / 92 / 2610.02695
Test-time Calibration Learning for Large Language Model Reasoning
用于大语言模型推理的测试时校准学习
Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predictions and supporting reliable decision-making in real-world deployment. Recent studies incorporate calibration learning into reinforcement learning (RL), jointly optimizing answer correctness and verbalized confidence using ground-truth correctness supervision. However, their reliance on labeled data limits their applicability in practical test-time settings, where ground-truth labels are unavailable and calibration may need to adapt to newly encountered target tasks. To address this challenge, we propose Test-Time Calibration Learning (TTCL), a label-free framework that jointly adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data. Specifically, TTCL derives self-supervision signals for both correctness and calibration from multiple model-generated responses, enabling calibration learning at test time without ground-truth labels. Theoretical analysis further establishes TTCL as a bounded surrogate for the ideal calibration objective. Extensive experiments on mathematical reasoning and factual question answering demonstrate that TTCL consistently improves both accuracy and calibration across diverse models and tasks. On base models, TTCL achieves an average relative accuracy improvement of +40.13% and an ECE reduction of +70.80% across eight benchmarks. Moreover, TTCL can further improve both accuracy and calibration for already calibrated models under domain shift, particularly when source-domain calibration transfers poorly to target tasks. In the math-to-factQA setting, TTCL achieves an average relative accuracy gain of +20.35% and reduces ECE by +53.83%. The source code is released at https://github.com/tmlr-group/TTCL.
Chinese Translation
可靠的大语言模型(LLMs)不仅必须产生准确的答案,还必须表达能够忠实反映其答案正确的概率的置信度。这种校准对于识别不确定的预测以及在真实世界部署中支持可靠决策至关重要。近期研究将校准学习纳入强化学习(RL),利用真实答案正确性监督来联合优化答案正确性和言语化置信度。然而,它们对标注数据的依赖限制了其在实际测试时设置中的适用性;在这些设置中,真实标签不可用,并且校准可能需要适应新遇到的目标任务。为应对这一挑战,我们提出测试时校准学习(TTCL),一个无标签框架,可直接在未标注的目标任务数据上联合适配推理准确性和言语化置信度。具体而言,TTCL 从多个模型生成的响应中为正确性和校准二者推导自监督信号,从而无需真实标签即可在测试时进行校准学习。理论分析进一步表明,TTCL 是理想校准目标的一个有界代理。在数学推理和事实问答上的大量实验表明,TTCL 在不同模型和任务上持续提升准确性和校准。在基础模型上,TTCL 在八个基准上实现了平均相对准确率提升 +40.13% 和 ECE 降低 +70.80%。此外,在领域偏移下,TTCL 可以进一步提升已经校准过的模型的准确性和校准,尤其是当源领域校准向目标任务迁移效果不佳时。在 math-to-factQA 设置中,TTCL 实现了平均相对准确率增益 +20.35%,并将 ECE 降低 +53.83%。源代码发布在 https://github.com/tmlr-group/TTCL。
cs.LG / 93 / 2610.02722
Structural-Functional Brain Connectivity Generation via Multimodal Hypergraph-based Flow Matching
基于多模态超图的流匹配生成结构-功能脑连接
Chyong Yi Poh, Hwa Hui Tew, Junn Yong Loo, Raphaël C. -W. Phan, Fuad Noman, Pew-Thian Yap, Chee-Ming Ting
cs.LG
diffusion
扩散模型相关
Abstract
Structural connectivity (SC) and functional connectivity (FC) provide complementary information on interactions between brain regions and are widely used in neuroimaging studies of neuropsychiatric disorders. Generative modelling can alleviate the scarcity of large-scale paired SC-FC data, but existing approaches typically use pairwise graphs that capture only dyadic interactions and often generate SC and FC independently, limiting preservation of higher-order structure-function relationships. We propose a Multimodal Hypergraph Flow Matching (MHG-FM) framework for joint SC-FC connectivity generation and cross-modal translation. MHG-FM constructs modality-specific hypergraphs, learns higher-order representations with Hypergraph Neural Network (HGNN) encoders, and performs bidirectional cross-modal fusion using Dual Cross-Attention (DCA). A variational autoencoder maps the fused representations to a compact latent space, where conditional flow matching enables connectivity synthesis and multimodal translation via latent transport. Experiments on the Human Connectome Project Young Adult (HCP-YA) dataset show that MHG-FM outperforms several state-of-the-art baselines in reconstruction quality, topology preservation, distributional similarity, and SC-FC coupling, while achieving approximately 8x faster sampling than a matched diffusion backbone.
Chinese Translation
结构连接(SC)和功能连接(FC)提供了关于脑区之间相互作用的互补信息,并广泛应用于神经精神疾病的神经影像学研究。生成式建模可以缓解大规模配对 SC-FC 数据的稀缺性,但现有方法通常使用仅能捕捉成对交互的成对图,并且往往独立生成 SC 和 FC,从而限制了高阶结构-功能关系的保持。我们提出一种多模态超图流匹配(MHG-FM)框架,用于联合 SC-FC 连接生成和跨模态转换。MHG-FM 构建模态特定的超图,使用超图神经网络(HGNN)编码器学习高阶表示,并使用双重交叉注意力(DCA)执行双向跨模态融合。一个变分自编码器将融合后的表示映射到紧凑的潜在空间,其中条件流匹配通过潜在传输实现连接合成和多模态转换。在人类连接组计划青年成人(HCP-YA)数据集上的实验表明,MHG-FM 在重建质量、拓扑保持、分布相似性和 SC-FC 耦合方面优于若干最先进的基线,同时相比匹配的扩散骨干模型实现了约 8 倍的更快采样。
cs.LG / 94 / 2610.02754
Jumping up and down: Denoiser diffusion models for discrete ordinal data
上下跳跃:用于离散有序数据的去噪器扩散模型
Yair Shenfeld, Ricardo Baptista, Stefano Peluchetti
cs.LG · stat.ME
diffusion
扩散模型相关
Abstract
Diffusion models are highly developed in continuous spaces for image and video domains. Recently, major advances have been made for discrete diffusion models for categorical data, specifically in the language domain. In contrast, diffusion models for discrete integer-valued data are less developed, despite the prevalence of this modality, ranging from images and music to gene counts. We introduce Jumping Up and Down (JUD)---a new family of denoiser-based diffusion models for discrete ordinal data. This is the first family of diffusion models for ordinal data which centers around training denoisers, which at the same time allows for bi-directional (up and down) perturbations of the data. The simplicity of the training objective, combined with the flexibility of bi-directional perturbations, leads us to obtain competitive results across different data modalities.
Chinese Translation
扩散模型在图像和视频领域的连续空间中已经高度发展。最近,针对分类数据的离散扩散模型取得了重大进展,特别是在语言领域。相比之下,用于离散整数值数据的扩散模型发展较少,尽管这种模态十分普遍,涵盖从图像、音乐到基因计数等。我们提出了 Jumping Up and Down (JUD)——一个用于离散有序数据的新型基于去噪器的扩散模型家族。这是首个以训练去噪器为核心的用于有序数据的扩散模型家族,同时允许对数据进行双向(向上和向下)扰动。训练目标的简单性,结合双向扰动的灵活性,使我们在不同数据模态上取得了有竞争力的结果。
cs.LG / 95 / 2610.02774
LatticeSMC: Where to Spend Inference-Time Compute in Chunked Sequence Generators
LatticeSMC:在分块序列生成器中应在何处花费推理时计算
Xuanchen Wang, Heng Wang, Weidong Cai
cs.LG
diffusion
扩散模型相关
Abstract
Long-form generators for music, motion and video produce sequences chunk by chunk, with each chunk generated by iterative denoising while rewards are defined over the full sequence. Existing inference-time steering methods typically act on one axis at a time: best-of-N at the end, Feynman-Kac steering across denoising steps, or streaming pruning across chunks, and are often compared under unmatched compute or different return rules. We introduce budget-matched chunked steering and propose LatticeSMC, a sampler derived from a Feynman-Kac model on the two-dimensional lattice of chunk index and denoising step. Two telescoping results make its design exact: for chunk-additive rewards, the two axes induce identical weights, so resampling should occur where lookahead is cheapest; for terminal rewards, any prefix score defines an exact intermediate potential, making prefix-evaluable rewards twists with no estimation or extra denoiser calls. LatticeSMC resamples on these potentials at chunk boundaries and, when scoring is free, within chunks, returning either a weighted draw or the best particle. Under matched compute, on music-to-dance diffusion and 40-second text-to-music generation, it raises beat alignment from 0.234 to 0.441 (best-of-N: 0.354) and prompt adherence from 0.470 to 0.560 at 32 particles, while preserving held-out quality. It also retains its advantage on long-range rewards and is preferred by human raters in 60-77 percent of pairwise comparisons. Finally, we show that commitment strength should follow the information in the current potential, while the value of lookahead is predicted by the within-set predictability of future reward.
Chinese Translation
用于音乐、动作和视频的长序列生成器逐块生成序列,其中每个块通过迭代去噪生成,而奖励定义在整个序列上。现有的推理时引导方法通常一次只作用于一个轴:末尾的 best-of-N、跨去噪步骤的 Feynman-Kac 引导,或跨块的流式剪枝,并且常常在未匹配的计算量或不同的回报规则下进行比较。我们引入预算匹配的分块引导,并提出 LatticeSMC,一种从块索引与去噪步骤的二维格上的 Feynman-Kac 模型导出的采样器。两个叠缩结果使其设计是精确的:对于块加性奖励,两个轴诱导出相同的权重,因此重采样应发生在前瞻成本最低之处;对于终端奖励,任何前缀分数都定义了一个精确的中间势,使得可前缀评估的奖励成为扭曲项,而无需估计或额外的去噪器调用。LatticeSMC 在块边界处基于这些势进行重采样,并且当评分免费时,在块内进行重采样,返回加权抽取或最佳粒子。在匹配计算量下,在音乐到舞蹈扩散和 40 秒文本到音乐生成上,它将节拍对齐从 0.234 提升到 0.441(best-of-N:0.354),并在 32 个粒子时将提示遵循度从 0.470 提升到 0.560,同时保持留出集质量。它还在长程奖励上保持其优势,并在 60-77% 的成对比较中更受人类评分者偏好。最后,我们表明,承诺强度应跟随当前势中的信息,而前瞻的价值由未来奖励的集合内可预测性预测。
cs.LG / 96 / 2610.02780
Controlling Polar Exposure to Delay Memorization in Diffusion Models
控制极坐标暴露以延迟扩散模型中的记忆化
Xuanchen Wang, Heng Wang, Weidong Cai
cs.LG
diffusion
扩散模型相关
Abstract
Diffusion models can reach useful sample quality before copying training examples, but fast optimization can compress this generalization window by accelerating sample-specific fitting. We investigate this effect through update geometry and propose Quality-Gated De-whitening (QGD), a controller that retains a fast polar-update prefix and progressively restores fixed-gain momentum. Our random-feature analysis separates covariance-controlled, curvature-equalized and amplitude-controlled memorization clocks. Under aligned spectral assumptions, it establishes a finite-exposure condition under which a fixed-gain tail recovers a delay proportional to dataset size. QGD implements this principle with a confirmed quality gate, a bounded decay envelope and causal copy feedback. Immediate switching is the conservative limit; gradual control balances delayed copying against continued quality improvement. We pair QGD with Copy-Budgeted Selection (CBS), which applies simultaneous binomial calibration to a frozen checkpoint family, followed by a fresh evaluation of the released checkpoint. On 2,000-image CIFAR-10 subsets, QGD preserves the polar baseline's quality-arrival time while expanding its useful interval by 8.32x and reducing common-checkpoint copying by 75.9%. With identical calibration and independent quality evaluation, QGD achieves FID 75.56 versus 79.37 for SGD with the same selector. Exposure-matched controls, independent detector audits and transfer to flow matching and dance generation support adaptive exposure control as a practical way to improve the quality-copying tradeoff.
Chinese Translation
扩散模型能够在复制训练样本之前达到有用的样本质量,但快速优化会通过加速样本特定拟合来压缩这一泛化窗口。我们通过更新几何研究这一效应,并提出质量门控去白化(Quality-Gated De-whitening, QGD),一种控制器,它保留快速极坐标更新前缀,并逐步恢复固定增益动量。我们的随机特征分析将协方差控制、曲率均衡和幅度控制的记忆化时钟区分开来。在对齐谱假设下,它建立了一个有限暴露条件,在该条件下,固定增益尾部恢复出与数据集规模成正比的延迟。QGD 通过一个经过确认的质量门、一个有界衰减包络和因果复制反馈来实现这一原则。立即切换是保守极限;渐进控制则在延迟复制与持续质量改进之间取得平衡。我们将 QGD 与复制预算选择(Copy-Budgeted Selection, CBS)配对,后者对冻结的检查点族应用同时二项校准,随后对发布的检查点进行一次全新评估。在 2,000 张图像的 CIFAR-10 子集上,QGD 保持了极坐标基线的质量到达时间,同时将其有用区间扩大了 8.32 倍,并将公共检查点复制减少了 75.9%。在相同校准和独立质量评估下,QGD 达到 FID 75.56,而使用相同选择器的 SGD 为 79.37。暴露匹配的对照、独立检测器审计以及向流匹配和舞蹈生成的迁移,支持自适应暴露控制作为一种改善质量-复制权衡的实用方法。
cs.LG / 97 / 2610.02835
All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR
只工作不玩耍,杰克也会变迟钝:理解并防止 RLVR 中的灾难性策略崩溃
Qiyuan Huang, Tianshi Xu, Meng Li
cs.LG
large language model
大语言模型相关
Abstract
During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at https://github.com/Ayanami-0123/Open-Mesh-Learning.
Chinese Translation
在使用可验证奖励强化学习(RLVR)对大型语言模型(LLMs)进行后训练期间,GRPO 风格算法可能表现出严重的后期崩溃。基于提示的探测表明,这并不是良性的策略剪枝,而是有效策略容量的有害收缩,使不同的推理策略日益难以被访问。为了刻画这一现象,我们通过轨迹级策略更新交互来定义策略,并开发了一个结合优化动力学与信息论的统一理论框架。我们证明,主要的 RLVR 目标会逐步将概率质量集中到单一策略上,而维持非平凡的任务准确率则需要一个最低策略容量。这两个结果之间的冲突为灾难性崩溃提供了一种机制性解释。我们进一步推导出镜像纠缠指数(MEI)作为轻量级在线预警信号。为了防止崩溃,我们提出 \textbf{Mesh Learning},它呈现多个推理策略,并防止任何单一策略主导优化。在 AIME26、AIME25、MATH-500、GPQA 和 LiveCodeBench 上,Mesh Learning 在 Qwen 和 Phi 模型系列上持续优于强基线,分别取得高达 13.4 个百分点和 11.5 个百分点的提升。这些结果确立了策略保持作为稳定 RLVR 的一项关键原则。代码可在 https://github.com/Ayanami-0123/Open-Mesh-Learning 获取。
cs.LG / 98 / 2610.02865
On Unlearning for Time-series Forecasting
关于时间序列预测中的遗忘
Zeyu Shi, Yanhui Luo, Ziming Hong, Chongyang Gao, Kezhen Chen, Shanshan Ye, Lixu Wang
cs.LG
diffusion
扩散模型相关
Abstract
Time-series forecasting is widely used in sensitive domains. Models in these settings are often trained on longitudinal user- or entity-level records, which may later require removal because they contain sensitive or proprietary information or have been corrupted by sensor failures. To address such deletion requests without costly retraining, machine unlearning has been widely studied as a practical mechanism for privacy protection and data governance. However, the application of machine unlearning to time series prediction has not yet been well realized; this is mainly due to the following unique challenges: Gradient-based unlearning can be unstable because a deleted observation participates in multiple causally connected forecasting windows, causing parameter updates to propagate beyond the requested interval and degrade retained forecasting utility. Label-guided updating offers a more controlled alternative, but continuous and context-dependent forecasts lack a suitable replacement target, while the exact-retrained output is unavailable during unlearning. Moreover, the remaining support for a deleted temporal pattern is highly non-uniform. Some affected windows retain structurally similar counterparts in the retained data, whereas others become underrepresented or isolated. We present RDTU, a Residual Diffusion framework for time-series unlearning. RDTU first uses a retained-set neural tangent kernel predictor to obtain a deletion-compatible base forecast. Then it quantifies the global and local structural support of each affected window using the volume contribution of the retained-reference data. Then a diffusion model generates a residual correction that estimates the counterfactual forecast, yielding a pseudo-label field that guides a lightweight model update. Experiments show that RDTU consistently produces unlearned models that most closely match exact retraining.
Chinese Translation
时间序列预测被广泛应用于敏感领域。这些场景中的模型通常基于纵向的用户级或实体级记录进行训练,而这些记录之后可能因包含敏感或专有信息,或已因传感器故障而损坏,而需要被移除。为了在不进行高成本重新训练的情况下处理此类删除请求,机器遗忘作为一种用于隐私保护和数据治理的实用机制已被广泛研究。然而,将机器遗忘应用于时间序列预测尚未得到很好的实现;这主要是由于以下独特挑战:基于梯度的遗忘可能不稳定,因为一个被删除的观测值会参与多个因果关联的预测窗口,导致参数更新传播到所请求区间之外,并降低保留下来的预测效用。标签引导的更新提供了一种更可控的替代方案,但连续且依赖上下文的预测缺乏合适的替代目标,而在遗忘过程中也无法获得精确重新训练后的输出。此外,对于被删除的时间模式,剩余支持是高度非均匀的。一些受影响的窗口在保留数据中仍保留结构相似的对应窗口,而另一些则变得表示不足或被孤立。我们提出 RDTU,一个用于时间序列遗忘的残差扩散框架。RDTU 首先使用保留集神经正切核预测器来获得与删除兼容的基础预测。然后,它使用保留参考数据的体积贡献来量化每个受影响窗口的全局和局部结构支持。然后,一个扩散模型生成残差校正,以估计反事实预测,从而产生一个伪标签场,指导轻量级模型更新。实验表明,RDTU 持续产生与精确重新训练最接近的遗忘后模型。
cs.LG / 99 / 2610.02906
Tangent Schrödinger Bridge Matching: Learning Stochastic Transport with Mechanistic Sensitivities
切向薛定谔桥匹配:学习具有机制敏感性的随机输运
Jowaria Khan, Elizabeth Bondi-Kelly
cs.LG
diffusion
扩散模型相关
Abstract
Predicting how stochastic systems respond to changes in viscosity, reaction rates, or external forces requires costly simulations, motivating reusable learned models. Yet matching observed outcome distributions does not ensure accurate intervention responses. We introduce Tangent Schrödinger Bridge Matching (Tangent-SBM), which learns stochastic transports from endpoint observations and mechanistic sensitivities. It propagates parameter derivatives alongside trajectories and supervises them against supplied targets. For average-response targets, single-rollout squared error also penalizes response variability; our objective uses two independent rollouts to match the mean without this additional penalty. We establish conditions under which sensitivity accuracy bounds finite-change prediction error and decision regret. Across Gaussian, stochastic double-well, PDEBench reaction--diffusion, and stochastic Navier--Stokes systems, Tangent-SBM improves sensitivity and finite-change prediction over matched conditional-bridge baselines while maintaining comparable endpoint and distributional accuracy. Controls examine target correctness, response objectives, and simulator-budget allocation. To test decision usefulness, we evaluate calibrated viscosity selection in Navier--Stokes: Tangent-SBM reduces tracking error relative to taking no action on every evaluated task.
Chinese Translation
预测随机系统如何响应粘度、反应速率或外力的变化,需要代价高昂的模拟,这促使人们构建可复用的学习模型。然而,匹配观测到的结果分布并不能确保准确的干预响应。我们提出切向薛定谔桥匹配(Tangent-SBM),它从端点观测和机制敏感性中学习随机输运。它将参数导数与轨迹一同传播,并依据给定的目标对其加以监督。对于平均响应目标,单次 rollout 的平方误差还会惩罚响应变异性;我们的目标函数使用两次独立的 rollout 来匹配均值,从而避免这一额外惩罚。我们建立了相应条件,在这些条件下,敏感性精度可为有限变化预测误差和决策遗憾给出界。在高斯系统、随机双势阱系统、PDEBench 反应--扩散系统以及随机纳维--斯托克斯系统中,Tangent-SBM 在保持相当的端点精度与分布精度的同时,相较于相匹配的条件桥基线提升了敏感性预测与有限变化预测。对照实验考察了目标正确性、响应目标以及模拟器预算分配。为检验决策有用性,我们评估了纳维--斯托克斯中的校准粘度选择:在每一个被评估的任务上,Tangent-SBM 相对于不采取任何行动都降低了跟踪误差。
cs.LG / 100 / 2610.02954
Hyperparameter selection for equation learning with biologically-informed neural networks
基于生物信息神经网络的方程学习中的超参数选择
William Lavery, Jodie A. Cochrane, John T. Nardini, Sara Hamis
cs.LG
diffusion
扩散模型相关
Abstract
Biologically-informed neural networks (BINNs) have emerged as a flexible subclass of physics-informed neural networks (PINNs) for learning terms in partial differential equations from data. BINNs are particularly suited for biological systems, where the governing equations are highly nonlinear and only partially known a priori, and where data observations are often sparse, noisy, and incomplete. However, applying BINNs effectively in practice depends critically on hyperparameter selection, which remains a central challenge in equation-learning frameworks. Hyperparameters are often chosen heuristically and only cursorily documented, which limits the reproducibility of results and the transferability of methods. We present a diagnostic workflow for hyperparameter selection that can be used when the ground-truth equations are not known. The workflow is guided by three main questions: (1) Are the benefits of greater network capacity worth the cost? (2) Do more training epochs keep reducing the validation loss? (3) Do the learned terms stop changing as network capacity and training increase? We apply our workflow to synthetic systems of varying complexity with known ground truth, spanning diffusion and growth right-hand side terms and data ranging from 1D+t to 2D+t. We demonstrate that the validation loss generally follows the true error in the learned terms and distil practical rules of thumb for selecting hyperparameters in the BINN architecture. By providing a structured workflow, practical guidelines, and suggested starting values for hyperparameter selection, this work lowers the barrier to BINN-based equation learning.
Chinese Translation
生物信息神经网络(BINNs)已经作为物理信息神经网络(PINNs)的一个灵活子类出现,用于从数据中学习偏微分方程中的项。BINNs 特别适合生物系统,在其中控制方程高度非线性且仅部分先验已知,并且数据观测往往稀疏、含噪且不完整。然而,在实践中有效应用 BINNs 关键取决于超参数选择,而这仍然是方程学习框架中的一个核心挑战。超参数通常被启发式地选择,并且仅被粗略记录,这限制了结果的可重复性和方法的可迁移性。我们提出了一种用于超参数选择的诊断工作流程,可在真实方程未知时使用。该工作流程由三个主要问题指导:(1)更大的网络容量带来的收益是否值得其代价?(2)更多的训练轮次是否会持续降低验证损失?(3)随着网络容量和训练的增加,学习到的项是否会停止变化?我们将该工作流程应用于具有已知真实值的不同复杂度合成系统,涵盖扩散和增长右端项,以及从 1D+t 到 2D+t 的数据。我们证明,验证损失通常跟随学习到的项中的真实误差,并提炼出在 BINN 架构中选择超参数的实用经验规则。通过提供结构化的工作流程、实用指南以及用于超参数选择的建议起始值,这项工作降低了基于 BINN 的方程学习的门槛。
cs.LG / 101 / 2610.03034
Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling
用于快速随机扩散采样的自适应二阶求解器
Ella Kemperman, Luca Ambrogioni
cs.LG · cs.CL · cs.CV
diffusion
扩散模型相关
Abstract
Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization important. We adapt proportional-integral (PI) step-size control to diffusion, using our diffusion noise-normalised error estimator. Unlike existing adaptive methods in diffusion that respond only to the current error, the PI solver also incorporates the previous error, yielding smoother step adaptation. We further show that these per-sample trajectories exhibit shared structure and can be aggregated into a fixed schedule that retains much of the benefit of adaptive sampling. We evaluate both approaches on natural-image and language datasets, in terms of quality, measured by FID at a matched number of neural network evaluations (NFE), comparing them with widely used stochastic solvers and schedules. For images, our fixed discretization outperforms the commonly used EDM schedule in terms of sample quality when used with the stochastic Heun sampler, and with the EDM-churn sampler at low NFE. Additionally, our PI adaptive solver obtains better FID than most stochastic and adaptive baselines, although it does not beat the EDM-churn sampler at low NFE. Moreover, we find our solver outperforms both the EDM and the entropy schedule on language diffusion at low-to-medium NFE in terms of perplexity, with the drawback of lower token entropy. Lastly, we find that the benefit of per-sample adaptivity is problem-dependent. It is highly beneficial in 1D toy examples, while only marginal for image and language data, where the average schedule sometimes even outperforms the PI-adaptive solver. Code is available at https://github.com/ellakemperman/adaptive-second-order-diffusion-solvers
Chinese Translation
扩散模型依赖于需要时间离散化的数值求解器,而时间离散化对采样成本与质量之间的权衡有很大影响。然而,逆向过程的计算难度会沿采样轨迹以及在不同数据分布之间变化,这使得离散化的选择十分重要。我们将比例-积分(PI)步长控制适配到扩散中,使用我们的扩散噪声归一化误差估计器。与扩散中仅对当前误差作出响应的现有自适应方法不同,PI 求解器还纳入了先前的误差,从而产生更平滑的步长自适应。我们进一步表明,这些逐样本轨迹表现出共享结构,并且可以聚合为一个固定调度,该调度保留了自适应采样的大部分优势。我们在自然图像和语言数据集上评估这两种方法,以质量为标准,通过在匹配的神经网络评估次数(NFE)下测得的 FID 来衡量,并将它们与广泛使用的随机求解器和调度进行比较。对于图像,在样本质量方面,当与随机 Heun 采样器一起使用时,以及当在低 NFE 下与 EDM-churn 采样器一起使用时,我们的固定离散化都优于常用的 EDM 调度。此外,我们的 PI 自适应求解器相比大多数随机和自适应基线获得了更好的 FID,尽管在低 NFE 下它未能超过 EDM-churn 采样器。此外,我们发现,在低到中等 NFE 下,我们的求解器在语言扩散上以困惑度为标准优于 EDM 调度和熵调度,但其缺点是 token 熵更低。最后,我们发现逐样本自适应的益处依赖于问题。它在 1D 玩具示例中非常有益,而对于图像和语言数据则仅带来边际收益,在这些数据上,平均调度有时甚至优于 PI 自适应求解器。代码可在 https://github.com/ellakemperman/adaptive-second-order-diffusion-solvers 获取。
cs.LG / 102 / 2610.03036
WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
WebFovea:当模型是对的但点击是错的——面向在线网站上基于视觉的网页智能体的可靠往返
Jiangang Han
cs.LG · cs.AI · cs.CV
large language model
大语言模型相关
Abstract
We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models.
Chinese Translation
我们提出了 WebFovea,一个基于视觉的网页智能体,它在 WebRetriever Challenge 2026 中以 100 分中 57.0 分的最终得分获得第 2 名。该挑战在 WebRetriever 基准(arXiv:2607.06118)的 Protocol III 上端到端评估智能体:从在线网站上的入口 URL 开始,智能体必须操作该网站自身的界面,并返回一个可验证的答案。一个能力强的多模态大语言模型(LLM)对此是必要的,但并不充分。模型的决策通过 harness 到达浏览器,harness 是模型与页面之间的代码。在每一步,四件事都必须正确完成:模型的回复必须被解析为预期的动作,动作必须在页面上生效,结果必须被准确地报告回来,并且必须向模型展示其所需的信息。在真实网站上,我们观察到的许多失败发生在四个阶段之一,而不是发生在模型的推理中。坐标空间不匹配使每次点击都落在其预期坐标的 3/4 处;对原生下拉菜单、iframe 内部以及文本框中的操作静默失败;自生成的聊天模板 token 污染了 4.9% 的任务回合。WebFovea 强化每个阶段,并用护栏包围整个循环,使智能体保持在规则和其预算之内。四阶段视角并不依赖于模型,尽管某些单独的修复确实依赖于模型。因为我们在所有四次提交中使用了同一个模型,我们官方隐藏集得分从 31.0 到 57.0 的上升反映了 harness 的改动,只不过存在在线站点上的运行间方差。我们描述了设计、每个组件的证据(包括负面结果)、失败分析、局限性,以及一个路线图,其中包括将不同步骤路由到不同模型。
cs.LG / 103 / 2610.03071
Learn Feasibility Once, Optimize All Objectives: Derivative-Free Diffusion Models for Chance-Constrained Programming
一次学习可行性,优化所有目标:面向机会约束规划的无导数扩散模型
Ziwen Liu, Yan Liu, Congying Han, Tiande Guo, Yao Yan, Weichen Zhao
cs.LG · math.OC
diffusion
扩散模型相关
Abstract
Chance-constrained programs (CCPs) optimize decisions under uncertainty by limiting the probability of constraint violation. Despite advances in traditional and learning-based approaches, optimizing non-convex or non-smooth objectives and adapting to different objectives under fixed chance constraints remain challenging. In this paper, we propose a \textbf{D}erivative-free \textbf{D}iffusion-based framework that \textbf{D}isentangles constraint modeling from objective optimization, termed \textbf{D$^3$Opt}. We learn the chance-feasible structure once, independently of any particular objective, by training a risk-conditioned diffusion model solely on constraint-filtered decisions and freezing it as a reusable prior for post-specified objectives. At inference time, we propose an annealed, particle-based Feynman--Kac correction along the frozen reverse diffusion process to optimize post-specified objectives using only function evaluations. This enables derivative-free optimization of non-convex and non-smooth objectives without objective-specific retraining. We prove that the correction preserves feasibility when this property holds for the frozen prior, and derive an optimization-error bound separating learned-prior coverage, finite-particle approximation, and finite-temperature effects. Experiments on linear Gaussian CCPs, objective-transfer tasks, and chance-constrained economic dispatch demonstrate effective optimization across smooth and non-smooth objectives, including non-convex cases, and objective generalization under fixed chance constraints without retraining.
Chinese Translation
机会约束规划(CCPs)通过限制约束违反的概率,在不确定性下优化决策。尽管传统方法和基于学习的方法取得了进展,但在固定机会约束下优化非凸或非光滑目标以及适应不同目标仍然具有挑战性。在本文中,我们提出了一个无导数、基于扩散的框架,该框架将约束建模与目标优化解耦,称为 \textbf{D$^3$Opt}。我们仅通过在约束过滤后的决策上训练一个风险条件扩散模型,并将其冻结为用于后指定目标的可复用先验,一次学习机会可行结构,且独立于任何特定目标。在推理时,我们提出一种退火的、基于粒子的 Feynman--Kac 校正,沿冻结的反向扩散过程进行,仅使用函数评估来优化后指定目标。这使得无需针对特定目标重新训练,即可对非凸和非光滑目标进行无导数优化。我们证明,当这一性质对冻结先验成立时,该校正保持可行性,并推导出一个优化误差界,将学习先验覆盖、有限粒子近似和有限温度效应分离开来。在线性高斯 CCP、目标迁移任务和机会约束经济调度上的实验表明,该方法能够有效地优化光滑和非光滑目标(包括非凸情形),并在固定机会约束下无需重新训练即可实现目标泛化。
cs.LG / 104 / 2610.03199
Predicting and Repairing Merge Collapse in Large Language Models
预测与修复大型语言模型中的合并崩溃
Jungseob Lee, Seungyoon Lee, Sugyeong Eo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists' task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of interference. Under a working noise model, the disturbance that a merge injects grows with the merge coefficient and with interference, yielding a pre-merge score. In our experiments on twenty-two merge configurations from four model families, only destructive merges exceed a threshold on this score. We find that statistics of sign conflict between specialists, a common target of existing merge operators, are anti-predictive. We then predicted the outcomes of fourteen merges before evaluating them, and twelve predictions were correct, including the destructive outcome of a specialist pair pushed past the threshold by continued pretraining. To address this collapse, we introduce PRISM, an operator that averages the task vectors first and then soft-thresholds each layer at a level set by the layer's interference. Without data or tuning, PRISM keeps all five destructive merges above the threshold within evaluation noise of the base model, where plain averaging falls at least 14.4 points below it or collapses entirely. We apply PRISM only above the threshold and keep the plain average for merges below it, which include all fifteen harmless ones. Code is available at https://github.com/js-lee-AI/PRISM.
Chinese Translation
从共享基座微调得到的大型语言模型可以通过对其任务向量取平均来合并,但有些合并会崩溃到远低于基座模型的水平,而常见的合并算子在评估之前不会给出任何预警。我们表明,专家任务向量的一个统计量既能预测这种崩溃,也能校准对它的修复。取平均所移除的功率等于任务向量在各位专家之间的方差,我们将这一量称为干扰。在一个可用的噪声模型下,合并所注入的扰动随合并系数以及干扰一同增长,从而给出一个合并前评分。在我们对来自四个模型家族的二十二种合并配置所做的实验中,只有破坏性的合并才会在该评分上超过阈值。我们发现,专家之间符号冲突的统计量——现有合并算子的一个常见目标——具有反向预测性。随后,我们在评估之前预测了十四次合并的结果,其中十二次预测正确,包括一对专家因继续预训练而被推过阈值后所产生的破坏性结果。为解决这种崩溃,我们提出了 PRISM,这是一种先对任务向量取平均、然后按照由该层干扰所设定的水平对每一层进行软阈值化的算子。在无需数据或调参的情况下,PRISM 使全部五次破坏性合并都保持在阈值之上,且处于基座模型的评估噪声范围之内,而朴素平均则至少低于该阈值 14.4 个点,或完全崩溃。我们仅在超过阈值时应用 PRISM,而对于低于阈值的合并保留朴素平均,这些合并包括全部十五次无害的合并。代码可在 https://github.com/js-lee-AI/PRISM 获取。
cs.LG / 105 / 2610.03226
D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?
D2K-Bench:LLM 智能体能否将专家设计转化为高效的 GPU 内核?
Daifeng Li, Huiqiang Jiang, Chengruidong Zhang, Wei Wu, Xudong Guo, Jianhong Tu, Jianwei Zhang, Binhang Yuan, Dayiheng Liu
cs.LG · cs.AI · cs.DC
large language model
大语言模型相关
Abstract
GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from $1.69\times$ to $2.49\times$. Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.
Chinese Translation
由大型语言模型(LLM)智能体生成的 GPU 内核可能仍不如专家实现高效,但仅凭运行时并不能揭示这一差距与设计发现和实现之间的关系。我们提出 D2K-Bench,这是一个包含 26 个任务和 85 个工作负载的诊断基准,用于衡量智能体将专家设计指导转化为高效 GPU 内核的有效程度。该指导涵盖 L1:高层算法洞见,L2:数据流设计,以及 L3:底层优化技巧,包括这些层级之间的依赖关系。有指导和无指导的成对运行共享任务描述、工作负载、工具、硬件以及 350 轮预算。互补性评估考察独立提出的设计以及生成代码中实现的设计属性。在 NVIDIA B200 GPU 上的五个模型中,指导将 130 个模型-任务对上的正确率从 93.1% 提高到 98.5%,并将所有 26 个任务上的 Performance Score 从 1.46 提高到 1.95。对于在两次运行中所有 26 个任务上都提交正确结果的三款前沿模型(GPT-6-Astra、Claude-Opus-4.8 和 GPT-5.6-Sol),几何平均加速比从 $1.69\times$ 提高到 $2.49\times$。在全部五个模型中,平均综合实现得分从 100 分制中的 57 分提高到 70 分。这些结果显示了专家设计指导的价值,同时识别出仍未实现的设计属性。
cs.LG / 106 / 2610.03289
Architecture-Dependent Fusion Pathways in MLLMs
MLLMs 中依赖架构的融合路径
Hebao Zhu, Dongxia Wu
cs.LG
large language model
大语言模型相关
Abstract
Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.
Chinese Translation
多模态大语言模型(MLLMs)在各类视觉-语言任务上取得了强劲的性能,然而,视觉与文本信息如何跨层融合的内部机制仍未得到充分理解。我们研究了来自两种架构范式的代表性 MLLMs:拼接式架构与原生多模态架构。我们进行了三项逐步关联的分析:对齐解耦识别出是哪种模态发生了变化,注意力路由与熵刻画了跨模态信息如何分布,而内在维度则考察融合如何重塑特征空间。另外,我们单独开展因果干预实验,以此作为对所得解释的验证。作为一项补充分析,我们使用视觉 CKA 来检验柏拉图表示假说。这些分析共同揭示了两种截然不同的融合路径:拼接式模型遵循文本优先、视觉在后的路径,而原生模型则表现出更早的视觉-文本协同适应与特征空间重组。这项工作为理解多模态融合提供了一种机制性视角,并支持对多模态表示进行架构感知的诊断。
cs.LG / 107 / 2610.03303
S$^{2}$-PINN: Stochastic Separable Physics-Informed Neural Networks
S$^{2}$-PINN:随机可分离物理信息神经网络
Zhendong Li, Akwum Onwunta
cs.LG · math-ph
diffusion
扩散模型相关
Abstract
Uncertainty quantification (UQ) for random partial differential equations (PDEs) is ubiquitous in computational science and engineering. However, classical spectral solvers for this class of problems face the curse of dimensionality, and existing neural solvers often ignore the stochastic structure that makes moments and calibration tractable. We introduce a stochastic separable physics-informed neural network, dubbed S$^{2}$-PINN, that represents the solution $u(t,\mathbf{x},\mathbf{Z})$ of a random PDE with a learnable Gaussian spatial dictionary, Fourier temporal features, and a generalized polynomial chaos (gPC) stochastic basis, coupled by a low-rank Canonical Polyadic (CP) tensor decomposition core. The method is trained with a hybrid strong-form and gPC-projected residual loss. Our theoretical analysis establishes that the separable class is dense in $L^2$ under mild conditions, and the projected residual corresponds exactly to a stochastic Galerkin constraint. Furthermore, we show that mini-batch projection coefficients are logarithmically dependent on the number of gPC modes, and that the orthogonality penalty controls the conditioning of the learned spatial dictionary. Using four manufactured random PDE benchmarks, we show that S$^{2}$-PINN outperforms nine baselines in terms of mean and variance accuracy, as well as calibration, while using significantly fewer parameters. Further evaluations on non-manufactured Poisson and Darcy problems, a stochastic Navier--Stokes problem, a diffusion scaling study of higher random dimensions, and two stochastic inverse problems reveal the generalization capabilities of the proposed structure. Together, these results support stochastic separability as an effective design principle for physics-informed neural UQ. The code for the experiments can be found in https://github.com/DMax1314/s2pinn
Chinese Translation
随机偏微分方程(PDEs)的不确定性量化(UQ)在计算科学与工程中无处不在。然而,针对这类问题的经典谱求解器面临维数灾难,而现有的神经求解器往往忽略使矩和校准变得可处理的随机结构。我们提出一种随机可分离物理信息神经网络,称为 S$^{2}$-PINN,它使用可学习的高斯空间字典、傅里叶时间特征和广义多项式混沌(gPC)随机基来表示随机 PDE 的解 $u(t,\mathbf{x},\mathbf{Z})$,并通过低秩 Canonical Polyadic(CP)张量分解核心进行耦合。该方法使用混合的强形式与 gPC 投影残差损失进行训练。我们的理论分析证明,在温和条件下,该可分离类在 $L^2$ 中是稠密的,并且投影残差恰好对应于随机 Galerkin 约束。此外,我们表明小批量投影系数对数依赖于 gPC 模态的数量,并且正交性惩罚控制所学空间字典的条件数。使用四个人造随机 PDE 基准问题,我们表明 S$^{2}$-PINN 在均值和方差精度以及校准方面优于九个基线,同时使用的参数显著更少。对非人造 Poisson 和 Darcy 问题、一个随机 Navier--Stokes 问题、更高随机维度的扩散缩放研究以及两个随机反问题的进一步评估,揭示了所提出结构的泛化能力。总之,这些结果支持将随机可分离性作为物理信息神经 UQ 的一种有效设计原则。实验代码可在 https://github.com/DMax1314/s2pinn 中找到。
cs.LG / 108 / 2610.03361
Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
追随赢家:面向无 Critic 的 RFT 的基于交叉熵方法的保守策略改进
Joery Ariën de Vries, Neil David Lawrence, Zhenwen Dai
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textit{Follow the Winners} (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW through a control-as-inference lens, which also recovers GRPO and DPO as specific modelling choices, identifying GRPO as risk-neutral while DPO and FTW share a bounded risk-seeking offset that FTW controls. We identify this offset as an inherent trade-off of variance reduction through ordinal filters on samples, whereas a critic model induces a different trade-off between bias and variance. Scaled to agentic LLM post-training, FTW matches GRPO and PPO on Sokoban and Search-R1 baselines, showing a viable trade-off from a value model or group rollouts to CPU memory.
Chinese Translation
无 Critic 的强化微调(RFT)用于智能体式大语言模型时,通常通过 GRPO 风格的方法完成,这些方法在重复 rollout 上计算组基线,以降低目标方差。然而,这种设置不适合在有状态环境中行动的智能体,例如在线服务或安全沙箱,在这些环境中,重复 rollout 不切实际且难以获得,而激进的更新会固化长且稀疏验证轨迹中的噪声。我们提出 \textit{Follow the Winners} (FTW),一种无 Critic 的策略学习算法,它将交叉熵方法适配到 RFT,用对回放缓冲区样本的序数过滤器替代组 rollout,从而在回报的顺序统计量中产生多项式集中。我们通过控制即推断视角推导出 FTW,该视角也将 GRPO 和 DPO 恢复为特定建模选择,从而将 GRPO 认定为风险中性,而 DPO 和 FTW 共享一个由 FTW 控制的有界风险寻求偏移。我们将这一偏移识别为通过对样本使用序数过滤器进行方差降低的固有权衡,而 Critic 模型则会引发偏差与方差之间的另一种权衡。当扩展到智能体式 LLM 后训练时,FTW 在 Sokoban 和 Search-R1 基线上与 GRPO 和 PPO 相匹配,展示了从价值模型或组 rollout 到 CPU 内存的一种可行权衡。
cs.LG / 109 / 2610.03396
A Unified Framework for Bayesian Data Assimilation with Generative Models and Observation Interpolants
一个结合生成模型与观测插值函数的贝叶斯数据同化统一框架
Nikolaj T. Mücke, Benjamin Sanderse
cs.LG · cs.CE
diffusion
扩散模型相关
Abstract
Bayesian data assimilation combines model forecasts with noisy observations, but sampling high-dimensional, non-Gaussian posteriors remains challenging. We introduce an observation-interpolant framework that turns pretrained stochastic interpolant, flow matching, and diffusion models into posterior samplers without retraining. Conditioning the interpolant path on observations yields a shared likelihood-score correction to the drift or velocity, unifying stochastic and deterministic posterior sampling. The resulting SDEs and ODEs sample the exact posterior when the intermediate likelihood score is known. For practical computation, we approximate this score using a closed-form Gaussian surrogate with a bias-corrected mean and covariance inflated by the model's source covariance. Jacobian-free and ensemble-shared approximations make the method tractable in high dimensions. We evaluate the framework on linear-Gaussian dynamics, stochastic two-dimensional Navier-Stokes, and urban airflow with up to $O(10^4)$ degrees of freedom.
Chinese Translation
贝叶斯数据同化将模型预报与含噪观测相结合,但对高维、非高斯后验进行采样仍然具有挑战性。我们提出一个观测插值框架,该框架无需重新训练即可将预训练的随机插值、流匹配和扩散模型转化为后验采样器。将插值路径以观测为条件,会产生对漂移项或速度项的共享似然得分校正,从而统一了随机和确定性后验采样。当中间似然得分已知时,所得的随机微分方程(SDE)和常微分方程(ODE)能够采样精确后验。为了实际计算,我们使用一个闭式高斯代理来近似该得分,其具有偏差校正的均值以及由模型的源协方差膨胀的协方差。无需雅可比矩阵以及集合共享的近似使该方法在高维中可行。我们在线性高斯动力学、随机二维 Navier-Stokes 以及具有高达 $O(10^4)$ 自由度的城市气流上评估该框架。
cs.LG / 110 / 2610.03480
Metropolis-Hastings Dominates Importance Resampling for Policy Composition
Metropolis-Hastings 在策略组合中优于重要性重采样
Alexey Kurennoy, Ramil Yarullin, Fergal Reid
cs.LG
large language model
大语言模型相关
Abstract
Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted product of the policies' probabilities over complete responses, but standard implementations combine their next-token probabilities, generally introducing sampling bias. We analyze a known iterative correction based on independence Metropolis-Hastings (MH). Our main result shows that, for every rollout budget, MH produces an output distribution at least as close to the target as sampling-importance-resampling (SIR) with the same budget, as measured by every convex f-divergence. We also derive a lower bound on MH's improvement over the uncorrected decoder in a consensus objective measuring agreement with the supplied policies. We further characterize the correction's sampling error in two asymptotic regimes: when the reward-specific policies approach agreement, and when the log ratio between target and uncorrected-decoder probabilities fluctuates increasingly widely, as can happen for long responses. We complement our analysis with experiments in enumerable and LLM-scale settings.
Chinese Translation
对大型语言模型(LLM)进行后训练通常需要探索多个奖励之间的权衡,但针对每一种权衡重新训练代价高昂。解码时策略组合允许在推理时通过组合特定于奖励的策略来调整这些权衡。这种组合的目标是各策略在完整回复上的概率的加权乘积,但标准实现组合的是它们的下一词元概率,通常会引入采样偏差。我们分析了一种已知的基于独立 Metropolis-Hastings(MH)的迭代校正。我们的主要结果表明,对于每一个 rollout 预算,MH 产生的输出分布至少与相同预算下的采样-重要性-重采样(SIR)一样接近目标,且以每一种凸 f-散度度量。我们还推导出了 MH 相对于未校正解码器在一个衡量与所提供的策略的一致性的共识目标上的改进下界。我们进一步刻画了该校正的采样误差在两种渐近情形下的表现:当特定于奖励的策略趋于一致时,以及当目标概率与未校正解码器概率之间的对数比波动越来越大时(长回复中可能发生)。我们用可枚举设置和 LLM 规模设置中的实验补充了我们的分析。
cs.LG / 111 / 2610.03503
Getting Your Guidance Weights Right in diffusion and flow-matching posterior sampling
在扩散与流匹配后验采样中正确设置引导权重
Liam Moroy, Jean-François Giovannelli, Yoann Altmann, Steve McLaughlin, Frédéric Champagnat, Guillaume Bourmaud
cs.LG
diffusion
扩散模型相关
Abstract
Training-free posterior sampling methods, also known as Plug-and-Play methods, leverage pretrained unconditional diffusion or flow-matching models to solve inverse problems. Most existing approaches rely on guidance weights to balance, at each time step, prior information from the unconditional score or velocity network with measurement consistency, yet the tuning of these weights is often not discussed and is largely left to heuristics. We introduce a simple and principled offline strategy for automatically tuning these guidance weights. Our key observation is that, at each time step, the conditional denoising score-matching objective for diffusion models, or the conditional flow-matching objective for flow-matching models, is a least-squares objective. Therefore, when the conditional prediction is expressed as a weighted sum of the unconditional network output and a measurement-guidance term, optimizing over these weights reduces to a two-dimensional linear least-squares problem. The resulting time-dependent guidance weights can be optimized offline for a given measurement operator, noise level and sampler at the cost of a single minibatch of sampling trajectories, without retraining or fine-tuning the pretrained generative model. Instantiated with the standard Tweedie-based measurement-consistency term, our approach improves posterior sampling and achieves state-of-the-art reconstruction performance across diffusion- and flow-matching-based methods. Moreover, the optimized guidance weights enable diffusion samplers to reduce the number of sampling steps from 1000 to 50 with no significant degradation in reconstruction quality. Code will be made available.
Chinese Translation
无需训练的后验采样方法,也称为即插即用方法,利用预训练的无条件扩散或流匹配模型来求解逆问题。大多数现有方法依赖引导权重来在每个时间步平衡来自无条件得分或速度网络的先验信息与测量一致性,但这些权重的调优往往未被讨论,且在很大程度上被留给启发式方法。我们提出一种简单且有原则的离线策略,用于自动调优这些引导权重。我们的关键观察是,在每个时间步,扩散模型的条件去噪得分匹配目标,或流匹配模型的条件流匹配目标,都是一个最小二乘目标。因此,当条件预测被表示为无条件网络输出与测量引导项的加权和时,对这些权重的优化就化为一个二维线性最小二乘问题。由此得到的时间相关引导权重可以针对给定的测量算子、噪声水平和采样器进行离线优化,其代价仅为单个小批量的采样轨迹,而无需重新训练或微调预训练的生成模型。以标准的基于 Tweedie 的测量一致性项进行实例化后,我们的方法改进了后验采样,并在基于扩散和流匹配的方法中取得了最先进的重建性能。此外,优化后的引导权重使扩散采样器能够将采样步数从 1000 减少到 50,而重建质量没有显著下降。代码将公开提供。
cs.LG / 112 / 2610.03529
Divergence controls entropy in distillation
散度控制蒸馏中的熵
Nicolas Zucchet, Scott W. Linderman
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.
Chinese Translation
蒸馏已成为大型语言模型训练的一个核心原语,但其性质尚未被很好地理解。我们采取熵的视角,研究学生的熵如何依赖于定义蒸馏目标的数据和散度。我们证明,前向 KL 会将学生的熵抬高到超过教师的熵。由于交叉熵训练是一个特例,这产生了一个恒等式,我们在预训练和有监督微调中对其进行了定量验证。其他散度并不带有这样的保证:反向 KL 会压缩熵,直到学生与教师之间的差距变得过大;而在两者之间插值,会在训练早期平滑地改变熵,但在收敛时突然改变。在策略蒸馏的较低熵来自 token 级反向 KL,而不是来自在策略采样。因此,散度充当了一种隐式的熵正则化器,其作用在自蒸馏中最为清晰:由于以特权信息为条件会压缩熵,效果最好的散度超参数正是那些对其加以补偿的超参数。
cs.LG / 113 / 2610.03551
Objects Without Morphisms: What LLMs for Mathematics Do Not Represent
无态射的对象:用于数学的 LLM 不表征什么
Yanli Wang, Suijin Wang, Xiaopeng Yuan, Haohan Wang
cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that survives it while leaving untouched what the model represents. We examine that question where no external criterion exists: translating statements between the dialects of neighbouring subfields, where fidelity turns on the level of generality at which content is asserted. The source leaves that level implicit in its vocabulary, so a faithful translation must recover it from the relation between the theories. We introduce an instrument that codes truth, content and scope in separate blind queues, with a judge-free measure of whether a rewrite states the hypothesis implicit in its source, and establish its sensitivity with a planted-positive control. Across seven models from four families, translating towards the general framing widens the domain of quantification in 60.6% of rewrites and narrows it in none; translating towards the concrete framing narrows it in 28.3% and widens it in 0.3%. The hypothesis that would prevent it is stated in 21.6% of model rewrites and 4.2% of human statements. Capability does not govern the asymmetry: it appears in every model tested, and the most capable widens least. It replicates on the half of the benchmark held out by a pre-registered rule, and on statements written by mathematicians. Instructing a model to state every hypothesis it requires raises that rate but not its sensitivity to direction. We argue that these systems have acquired an object-level correspondence between subfield vocabularies without the constraint under which a translation between theories carries hypotheses to hypotheses.
Chinese Translation
大语言模型(LLM)在竞赛数学上达到专家级表现,主要是通过在它们周围布置的大量搜索:候选解被大量采样,并且仅当外部标准接受它们时才被保留。这样的程序改善了在其之后存活下来的结果,却不触及模型所表征的内容。我们在不存在外部标准的地方考察这一问题:在相邻子领域的方言之间翻译陈述,在这里,保真性取决于内容被断言时所处的一般性水平。源文本将该水平隐含在其词汇中,因此忠实的翻译必须从理论之间的关系中恢复该水平。我们引入一种工具,它在彼此独立的盲队列中编码真值、内容与范围,并用一种无需评判者的度量来衡量一次改写是否陈述了其源文本中隐含的假设,还用植入式阳性对照确立其灵敏度。在来自四个模型家族的七个模型中,朝向一般化框架翻译时,在 60.6% 的改写中拓宽了量化域,且没有一次使其收窄;朝向具体化框架翻译时,在 28.3% 中使其收窄,在 0.3% 中使其拓宽。本可防止这一点的假设在 21.6% 的模型改写和 4.2% 的人类陈述中被陈述。能力并不支配这种不对称:它出现在每个受测模型中,而能力最强的模型拓宽得最少。它在由预注册规则留出的基准的一半上得到复现,也在数学家撰写的陈述上得到复现。指示模型陈述它所需的每一个假设会提高该比率,但不会提高它对方向的敏感性。我们认为,这些系统已经获得了子领域词汇之间的对象级对应,却没有获得那种约束:在该约束下,理论之间的翻译会将假设传递为假设。
cs.LG / 114 / 2610.03641
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
IDRF:掩码离散扩散模型的逆蒸馏奖励微调
Vladislav Gromadskii, David Li, Samson Gourevitch, Yazid Janati, Eric Moulines, Maxim Panov, Alexander Korotin
cs.LG
diffusion
扩散模型相关
Abstract
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.
Chinese Translation
掩码离散扩散模型为自回归生成提供了一种有前景的替代方案,但迭代采样可能代价高昂,且难以处理的序列似然使奖励微调变得复杂。我们提出 IDRF,一个用于少步掩码离散扩散生成器奖励微调的框架。从标准的反向 KL 正则化目标出发,IDRF 用逆蒸馏正则化取代了难以处理的序列级 KL 惩罚。在最优辅助去噪器下,我们证明群体逆蒸馏损失是到参考分布的序列级 KL 散度的上界。IDRF 在不进行参考模型 rollout 的情况下优化该损失的基于轨迹的代理目标,因此学生模型保留其自身的少步采样器。我们将少步生成视为有限时域马尔可夫决策过程,并在学生模型的轨迹上以裁剪的策略梯度目标优化奖励。在 DNA、图像和文本生成中,IDRF 实现了高奖励,其去噪步数比参考模型最多少 $32\times$,同时缓解了奖励黑客并保持了样本质量。
cs.LG / 115 / 2610.03665
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD:面向掩码扩散语言模型的高效自蒸馏
Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan
cs.LG · cs.CL
diffusion
扩散模型相关
Abstract
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
Chinese Translation
掩码扩散语言模型(dLMs)为复杂推理提供了一种有前景的、可并行的自回归模型替代方案。然而,它们面临着一个独特的信用分配挑战,因为去噪过程中的少数几个承诺会急剧降低其余被掩码位置上的不确定性,并塑造了响应的大部分内容。大多数针对 dLM 的后训练方案并未利用这一信号来决定在哪些 token 上进行训练:它们通常是在最终文本上进行训练,或是将奖励分配给整个去噪步骤,而不是选择那些塑造了响应的单个承诺。我们提出 Pivot-SD,一个高效的离线自蒸馏框架,它仅对这些高影响力的承诺(关键支点,pivots)进行监督。Pivot-SD 使用一种信息增益指标来选择支点,该指标衡量在其余被掩码位置上不确定性的降低程度。来自成功轨迹的支点使用交叉熵进行训练,而来自失败轨迹的支点则使用定向非似然(targeted unlikelihood)进行训练,失败轨迹的其余部分保持不变。仅使用 200 个问题、每个问题四次采样(rollout),Pivot-SD 便在数学与代码基准上,使 LLaDA-8B-Instruct 相较于全序列 SFT 以及预算匹配的扩散 RL 基线均有所提升。
cs.LG / 116 / 2610.03702
LESSER: Post-Training Data Selection with Output-Layer Gradients
LESSER:基于输出层梯度的后训练数据选择
Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue
cs.LG
large language model
大语言模型相关
Abstract
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by $9.7\times$ for SFT and $3.0\times$ for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.
Chinese Translation
用于大型语言模型的后训练数据的选择会显著影响下游性能。基于梯度的数据选择是一种流行的方法,它根据训练数据的梯度与一个小型验证集的梯度的对齐程度来对训练数据进行排序。然而,使用全参数梯度进行排序需要对每个样本进行一次昂贵的反向传播,这使得计算在大型候选池上变得难以处理。这引出了一个自然的问题:我们能否以一小部分成本近似全梯度特征?方便的是,我们发现输出层梯度就足以进行有效的数据选择,而且只需要更便宜的前向传播。我们将其实现为 LESSER,一种用于选择方法的即插即用封装器,它在 SFT 上把特征提取的 FLOP 成本降低了 $9.7 imes$,在 RL 基准上降低了 $3.0 imes$,同时在下游任务上保持了全梯度的性能。从经验上看,我们发现即使输出层梯度和全梯度对单个样本的排序不同,它们也会选择梯度对齐的批次。
cs.NE / 117 / 2610.02361
SEDIMA: Cross-Run Hierarchical Insight Memory for Evolutionary Search Agents
SEDIMA:面向进化搜索智能体的跨运行分层洞察记忆
Amirhossein Abaskohi, Mahdi Mostajabdaveh, Zirui Zhou
cs.NE · cs.CL
large language model
大语言模型相关
Abstract
Large language model (LLM)-driven evolutionary search is a powerful paradigm for automated program and algorithm discovery, yet existing systems are largely memoryless: each run explores from scratch, so agents repeatedly rediscover the same improvements and re-encounter the same dead ends. We introduce SEDIMA, a persistent hierarchical insight memory for evolutionary search agents. SEDIMA distills raw traces into natural-language insights, clusters them by semantic similarity using attention-weighted centroids, and retrieves relevant guidance to condition future mutations, accumulating transferable knowledge across runs and problems rather than within a single trajectory. As a drop-in module that leaves the search operators unmodified, SEDIMA improves average final performance by 5.5% on AlgoTune and 6.6% on ALE-Bench LITE under a fixed budget of 100 evaluated candidates. Under OpenEvolve, SEDIMA requires 32.3% fewer iterations on average to reach baseline-best performance across the five evaluated backbones.
Chinese Translation
大语言模型(LLM)驱动的进化搜索是一种用于自动化程序和算法发现的有力范式,然而现有系统在很大程度上是无记忆的:每次运行都从头开始探索,因此智能体会反复重新发现相同的改进,并反复遇到相同的死胡同。我们提出 SEDIMA,一种用于进化搜索智能体的持久分层洞察记忆。SEDIMA 将原始轨迹提炼为自然语言洞察,使用注意力加权质心按语义相似度对其进行聚类,并检索相关指导以调节未来的变异,从而跨运行和问题累积可迁移知识,而不是在单个轨迹内累积知识。作为一个即插即用模块,且不修改搜索算子,SEDIMA 在 100 个被评估候选的固定预算下,将 AlgoTune 上的平均最终性能提高了 5.5%,将 ALE-Bench LITE 上的平均最终性能提高了 6.6%。在 OpenEvolve 下,在五个被评估的骨干模型上,SEDIMA 达到基线最佳性能平均所需的迭代次数减少了 32.3%。
cs.AI / 118 / 2610.02527
CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization
CriticHack:评估机器人策略优化下的视觉奖励
Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang
cs.RO · cs.AI · cs.LG
diffusion
扩散模型相关
Abstract
Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy. We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task. Starting from a supervised policy with no prior reward exposure, five training runs raise task success by 10.2 percentage points and wrong-object failures by 10.9 points on 512 evaluation seeds, whereas five runs trained on the simulator's task-completion signal raise success without amplifying wrong-object failures (difference 9.2 points, 95% CI 5.6 to 13.0). The amplification recurs from a policy previously optimized against learned rewards, under the policy's native diffusion sampler, at matched distance from the initial policy, and across constrained-policy experiments with two critics and two optimizers. A tilt model explains when it occurs: under KL-regularized optimization, an outcome becomes more frequent whenever its expected reward under the initial policy exceeds the population average. Robometer separates successes from failures well overall (AUROC .81) but scores wrong-object failures slightly above successes (AUROC .37), so optimization raises both. The same model predicts the outcome shifts across 26 constrained settings (Spearman .89), including those in which task success falls, and Robometer's own published success-termination recipe inherits the error. A frozen outcome verifier redirects the same optimization toward the requested task.
Chinese Translation
学习得到的视觉奖励模型正越来越多地被用于优化机器人策略,但一个奖励模型可能给作用在错误物体上的执行打出与完成任务的执行一样高的分数。我们表明,优化此类奖励会在奖励和任务成功率二者都上升的同时放大这些错误物体失败,因此从业者通常会监测的信号看起来是健康的。我们在一个抽屉任务上,针对 Robometer 微调扩散策略的每一个去噪器参数。从一个此前未接触过奖励的监督策略出发,五次训练运行在 512 个评估种子上将任务成功率提高 10.2 个百分点,并将错误物体失败提高 10.9 个百分点,而基于模拟器任务完成信号训练的五次运行则提高了成功率,却没有放大错误物体失败(差异为 9.2 个百分点,95% CI 为 5.6 至 13.0)。这种放大现象会复现:从一个先前针对学习奖励优化过的策略出发,在该策略原生的扩散采样器下,在与初始策略匹配的距离下,以及在带有两个 critic 和两个优化器的受约束策略实验中。一个倾斜模型解释了它何时发生:在 KL 正则化优化下,只要某个结果在初始策略下的期望奖励超过总体平均值,它就会变得更频繁。Robometer 总体上能很好地区分成功与失败(AUROC .81),但对错误物体失败的评分略高于成功(AUROC .37),因此优化会同时提高二者。同一个模型预测了 26 个受约束设置中的结果偏移(Spearman .89),包括任务成功率下降的那些设置,并且 Robometer 自己已发表的成功终止配方继承了这个错误。一个冻结的结果验证器会将同一优化重定向到所请求的任务。
cs.LG / 119 / 2610.03132
Safe Streaming Flow Planning by Aligning Sampling Dynamics with Execution Dynamics
通过将采样动力学与执行动力学对齐的安全流式流规划
Seunghwan Jang, Jeongyong Yang, Siddharth Ancha, SooJean Han
cs.RO · cs.LG
diffusion
扩散模型相关
Abstract
Generative planners based on diffusion/flow matching can learn to synthesize long-horizon trajectories from demonstrations. However, real-world deployment requires (i) enforcing safety constraints during execution and (ii) tight online replanning at fast execution rates. Prior safe diffusion/flow planners generate the agent's full trajectory at once, while repeatedly perturbing intermediate states to satisfy safety constraints. This approach is not only computationally intensive, but also introduces distribution shift since the learned sampling dynamics is distinct from the system's execution dynamics. We propose SafeStreamingFlow, a goal-conditioned planner that aligns flow sampling dynamics with execution dynamics by sequentially integrating a learned state vector field with hierarchical state prediction. Importantly, we need to enforce safety constraints only for the executed step via high order control barrier functions. Across navigation, racing, and locomotion benchmarks, SafeStreamingFlow reduces planning latency and improves safety compared to existing methods, while maintaining competitive goal-reaching success.
Chinese Translation
基于扩散/流匹配的生成式规划器可以学习从演示中合成长期轨迹。然而,现实世界的部署需要 (i) 在执行过程中施加安全约束,以及 (ii) 以快速执行速率进行紧密的在线重规划。先前的安全扩散/流规划器一次性生成智能体的完整轨迹,同时反复扰动中间状态以满足安全约束。这种方法不仅计算密集,而且会引入分布偏移,因为所学的采样动力学与系统的执行动力学不同。我们提出 SafeStreamingFlow,一种目标条件规划器,它通过将学习到的状态向量场与分层状态预测顺序积分,使流采样动力学与执行动力学对齐。重要的是,我们只需通过高阶控制屏障函数对所执行的步骤施加安全约束。在导航、竞速和运动基准上,与现有方法相比,SafeStreamingFlow 降低了规划延迟并提高了安全性,同时保持了具有竞争力的目标到达成功率。
cs.AI / 120 / 2610.03333
Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation
面向富接触操作的等变视觉-触觉扩散策略
Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou
cs.RO · cs.AI
diffusion
扩散模型相关
Abstract
Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: https://vista-paper.github.io/
Chinese Translation
富接触操作的模仿学习需要高质量专家数据,而这些数据获取成本高昂。这使得学习样本高效的策略成为一个关键问题。为了解决这个问题,我们提出了 VISTA,一种用于数据高效富接触模仿学习的工作空间级等变视觉-触觉扩散策略。VISTA 将视觉和触觉观测投影为球面标记,通过置换等变球面融合将触觉接触线索注入视觉球面方向,并利用末端执行器朝向旋转融合后的谐波表示。所得的表示作为等变扩散策略的条件,以预测空间一致的动作。在仿真和真实世界机器人设置中的大量实验表明,VISTA 在数据效率上显著优于强视觉-触觉模仿学习基线。项目网站:https://vista-paper.github.io/
cs.LG / 121 / 2610.03516
XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation
XGenAct:通过跨任务生成实现几何增强的世界动作模型
Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li
cs.RO · cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.
Chinese Translation
世界动作模型(WAMs)通过预测观测和动作如何随时间演变,推动了机器人控制的发展。尽管取得了这一进展,基于 RGB 和动作的未来预测并未明确解决机器人操作所需的空间理解。现有工作通常通过专用头或分支添加一组有限的空间预测任务,导致空间监督的范围和模型架构都变得碎片化。我们提出 XGenAct,一种世界动作模型,它通过确定性编解码器将 RGB 观测、机器人动作、度量深度、表面法线和功能角色分割表示为 RGB 视频。通过在训练期间对感知和动作任务进行采样,XGenAct 使用一个视频扩散变换器和一个目标函数,在这些空间上学习时序预测,而无需特定于模态的学习头。在留出的 RLBench 任务上,结构化感知训练相比仅 RGB 训练提高了平均闭环成功率,并且 XGenAct 在五项任务的外部比较中达到 52% 的成功率,而评估的最强基线为 26%。它也比那些先生成 RGB、然后应用冻结感知专家的评估流程更准确地预测未来深度和分割。
cs.AI / 122 / 2610.03390
DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
DriftTTS:通过分布匹配漂移实现无需蒸馏的少步文本到语音
Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam
cs.SD · cs.AI
diffusion
扩散模型相关
Abstract
Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git
Chinese Translation
少步神经文本到语音模型通常依赖于缩短的扩散或流匹配调度,或依赖于从预训练的多步教师模型中进行蒸馏。为了避免这些依赖,我们提出了 DriftTTS,一种少步梅尔频谱图生成器,其训练不依赖生成式教师、蒸馏或对抗判别。DriftTTS 在由原始梅尔频谱和在同一 LJSpeech 训练划分上预训练的冻结掩码自编码器编码器定义的梅尔域特征空间中,使用分布匹配漂移目标。同策略展开在其自身的中间状态上训练解码器,并支持推理达到训练时的展开深度。在 LJSpeech 上,NFE=4 的 DriftTTS 取得了 3.87 dB MCD 和 3.7% WER,而 Matcha-TTS 为 3.85 dB 和 3.4%。在完全配对的盲听测试中,DriftTTS 获得 4.18 MOS,而 Matcha-TTS 为 3.96,真实语音为 4.22。这些结果表明,在没有预训练生成式教师的情况下,能够实现具有竞争力的少步合成。代码可在 https://github.com/BASHLab/driftTTS.git 找到。
cs.SE / 123 / 2610.03010
Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows
工程化可持续智能体:面向开发者工作流的智能体式大语言模型的系统性比较
Merve Astekin, Yan Naing Tun, Arda Goknil, Erik Johannes Husom, Lwin Khin Shar, Hasan Sözer, Ratnadira Widyasari, Hui Song
cs.SE · cs.MA
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across five software engineering tasks: code generation, technical debt identification, code vulnerability detection, log parsing, and log analysis. For each task, we compare LLM configurations that range from a non-agentic single-query baseline to multi-agent workflows, using six open-weight LLMs, two prompt strategies, and three hardware platforms. We assess each configuration in terms of accuracy, inference latency, and energy consumption. Our results reveal substantial trade-offs between agentic complexity and energy efficiency: multi-agent designs consume on average 6.36$\times$ as much energy and run 6.07$\times$ as long as the non-agentic baseline, with worst-case slowdowns of up to 160$\times$ for individual task--hardware pairs. Accuracy gains from additional agents are limited and task-specific: multi-agent improves average vulnerability-detection accuracy, but lightweight non-agentic and single-agent configurations still dominate the Pareto front, accounting for 59 of 66 Pareto-optimal configurations. Model and prompt choice act as task-specific levers whose effective direction varies between tasks rather than as global defaults. We translate these findings into design guidelines for sustainable, task-aware LLM-based development tools.
Chinese Translation
大语言模型(LLM)正日益被用于软件工程,包括协调多个智能体的智能体式系统,但其带来了更高的计算成本和环境成本。在本文中,我们对跨五项软件工程任务的智能体式 LLM 系统进行了全面的实证研究:代码生成、技术债务识别、代码漏洞检测、日志解析和日志分析。对于每项任务,我们比较了从非智能体式的单次查询基线到多智能体工作流的各种 LLM 配置,使用了六个开放权重 LLM、两种提示策略和三种硬件平台。我们从准确率、推理延迟和能耗三个方面评估每种配置。我们的结果揭示了智能体式复杂性与能效之间的显著权衡:多智能体设计平均消耗的能耗是非智能体式基线的 6.36$\times$,运行时间是其 6.07$\times$,对于单个任务--硬件组合,最坏情况下的减速最高可达 160$\times$。增加智能体所带来的准确率提升有限且因任务而异:多智能体提高了平均漏洞检测准确率,但轻量级的非智能体式和单智能体配置仍然主导帕累托前沿,在 66 个帕累托最优配置中占据 59 个。模型和提示的选择充当任务特定的杠杆,其有效方向在不同任务之间有所变化,而非作为全局默认值。我们将这些发现转化为面向可持续、任务感知的基于 LLM 的开发工具的设计指南。
cs.SE / 124 / 2610.03080
MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code
MintEval:大语言模型是否实现了你所要求的交易策略?一个面向自然语言到策略代码的行为等价性基准
Siyu Wang, Yifan Wang, Yuecheng He
cs.SE · cs.CL · cs.LG · q-fin.TR
large language model
大语言模型相关
Abstract
Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
Chinese Translation
大语言模型正在从生成交易信号转向编写执行这些信号的代码。第二种角色的失败模式是静默的:生成的代码能够运行,回测能画出图,但交易者所描述的风险逻辑并不是实际被执行的逻辑。现有的代码基准通过单元测试检验功能正确性,金融基准检验预测能力;二者都没有衡量一个实现的行为是否与所要求的策略一致。我们提出 MintEval,这是一个基准,其中的参考策略由一组可组合的构建模块以程序化方式生成,被反向翻译为口语化的交易者指令,再由被测模型重新实现。生成程序与参考程序在相同的市场数据和交易摩擦下逐根 K 线执行,并基于它们的行为动作而非代码相似度或利润进行比较:alpha 被差分消去。MintEval v0 包含 800 个基于 BTCUSDT 15 分钟数据的任务,按一种由执行过程测得的、与描述长度解耦的状态跨度复杂度 tau 进行分层。低成本模型的平均 ActionMatch 至多为 0.544,至多能精确复现 0.087 的任务;在由 200 个任务组成的分层子集上,一个前沿模型(Claude Opus 5.5)达到 0.889,并精确复现 0.575,但在 0.275 的任务上仍然静默失败。当给定一份构建模块菜单时,模型几乎能完美识别出策略,然而在规格被正确读取的实现中,有 79.2% 在超过 10% 的活跃 K 线上出现偏离。一个近期策略生成基准的 LLM 评判器,按原文逐字应用时,接受了其中每一次静默失败。
cs.LG / 125 / 2610.03066
Explainable Molecular Structure Inference from GC--MS with Diffusion Models and LLM Reranking
使用扩散模型与 LLM 重排序从 GC--MS 进行可解释的分子结构推断
Changlin Liu, Tianyu Yi, Chengchun Liu, Boxuan Zhao, Fanyang Mo
physics.chem-ph · cs.LG · physics.comp-ph · physics.data-an
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
GC--EI--MS is an important technique for analyzing volatile and semivolatile compounds in complex samples. However, conventional methods rely heavily on reference spectral library matching, limiting their ability to identify compounds absent from these libraries and to infer complete molecular structures directly from fragmentation information. Here, we present DiffGCMS, a spectrum-conditioned discrete graph diffusion model for de novo structure elucidation from GC--EI--MS, and further develop a framework that integrates DiffGCMS with second-stage reasoning by a large language model (LLM). In the first stage, DiffGCMS generates candidate molecular structures from input spectra; in the second stage, the LLM uses mass spectral information to validate, repair, and rerank the candidates and provides interpretable analysis of fragment-ion peaks. This framework can generate plausible molecular structures for compounds absent from reference spectral libraries and provide traceable evidence supporting its decisions. On a test set comprising 13,696 spectra from NIST 20, the generative model achieved Acc@1 and Acc@10 of 6.01\% and 15.76\%, respectively. On the test subset containing molecules with no more than 10 heavy atoms, LLM-assisted molecular graph repair and reranking increased Acc@1 from 21.28\% to 21.95\%, Acc@10 from 46.91\% to 47.99\%, and candidate validity from 91.04\% to 100\%. These results demonstrate that spectrum-aware postprocessing can correct errors produced by the generative model while providing auditable and traceable explanations for the final ranking.
Chinese Translation
GC--EI--MS 是分析复杂样品中挥发性与半挥发性化合物的一项重要技术。然而,传统方法严重依赖参考谱库匹配,限制了其识别这些谱库中不存在的化合物以及直接从碎片化信息推断完整分子结构的能力。在此,我们提出了 DiffGCMS,一种用于从 GC--EI--MS 进行从头结构解析的、以谱图为条件的离散图扩散模型,并进一步开发了一个将 DiffGCMS 与大型语言模型(LLM)的第二阶段推理相结合的框架。在第一阶段,DiffGCMS 从输入谱图生成候选分子结构;在第二阶段,LLM 利用质谱信息对候选结构进行验证、修复和重排序,并对碎片离子峰提供可解释分析。该框架能够为参考谱库中不存在的化合物生成合理的分子结构,并提供支持其决策的可追溯证据。在包含来自 NIST 20 的 13,696 张谱图的测试集上,生成模型的 Acc@1 和 Acc@10 分别达到 6.01\% 和 15.76\%。在包含重原子数不超过 10 的分子测试子集上,LLM 辅助的分子图修复和重排序将 Acc@1 从 21.28\% 提升至 21.95\%,将 Acc@10 从 46.91\% 提升至 47.99\%,并将候选有效性从 91.04\% 提升至 100\%。这些结果表明,谱图感知后处理能够纠正生成模型产生的错误,同时为最终排序提供可审计且可追溯的解释。
cs.LG / 126 / 2610.02538
ENCORE: Exact Non-equilibrium COntrol with Replica Exchange for Diffusion Generation
ENCORE:用于扩散生成的、使用副本交换的精确非平衡控制
Jiahao Yu, Saifuddin Syed, José Miguel Hernández-Lobato, Jiajun He
stat.ML · cs.LG
diffusion
扩散模型相关
Abstract
Inference-time control steers a pretrained generative model towards a target distribution without retraining. We study tilted targets $π_0\propto G_0\,p_0$, where $p_0$ is the sampler output distribution and $G_0$ is an evaluable reweighting function. Existing approaches rely on sequential annealing with sequential Monte Carlo (SMC) or parallel annealing with replica exchange (RE). Sequential control is exact but needs large particle populations, whereas no exact parallel control method exists: existing RE corrections approximate an intractable time reversal and are biased. We propose Exact Non-equilibrium COntrol with Replica Exchange (ENCORE), the first exact parallel control method. Each replica stores its generation trajectory, so the upward move is a truncation and the intractable time reversal is never simulated. We prove target invariance and show that the resulting dynamics are those of non-equilibrium replica exchange with the exact time reversal as forward proposal. Under regularity conditions, our diffusion analysis shows that both sequential and parallel control become unstable under refinement of the time discretisation without guidance, whereas guided proposals remain stable and yield diagnostics for tuning the schedule and the computational budget. Across synthetic targets, Boltzmann sampling of biomolecules, and image generation, ENCORE achieves competitive accuracy and diversity, remains robust to sampler perturbations, and applies to distilled samplers where existing RE corrections are unavailable.
Chinese Translation
推理时控制引导预训练的生成模型朝向目标分布,而无需重新训练。我们研究倾斜目标 $π_0\propto G_0\,p_0$,其中 $p_0$ 是采样器输出分布,$G_0$ 是一个可求值的重加权函数。现有方法依赖于使用序贯蒙特卡洛(SMC)的序贯退火,或使用副本交换(RE)的并行退火。序贯控制是精确的,但需要大量粒子群体,而尚不存在精确的并行控制方法:现有的 RE 校正近似一个难以处理的时间反演,并且是有偏的。我们提出使用副本交换的精确非平衡控制(ENCORE),这是首个精确的并行控制方法。每个副本存储其生成轨迹,因此向上移动是一次截断,而难以处理的时间反演从未被模拟。我们证明了目标不变性,并表明所得动力学正是以精确的时间反演作为前向提议的非平衡副本交换的动力学。在正则性条件下,我们的扩散分析表明,在没有引导的情况下,随着时间离散化的细化,序贯控制和并行控制都会变得不稳定,而引导提议保持稳定,并产生用于调节调度和计算预算的诊断量。在合成目标、生物分子的玻尔兹曼采样和图像生成中,ENCORE 实现了有竞争力的准确性和多样性,对采样器扰动保持稳健,并适用于现有 RE 校正无法使用的蒸馏采样器。
cs.AI / 127 / 2610.02663
Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data
分数匹配扩散模型在内蕴低维数据上的泛化性质
Saptarshi Chakraborty, Quentin Berthet, Peter L. Bartlett
stat.ML · cs.AI · cs.LG · math.ST
diffusion
扩散模型相关
Abstract
Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in real data, such as natural images and molecular geometries. In this work, we study the statistical generalization of flow-matching models for learning an unknown distribution $P_{\mathrm{data}}$ from finitely many samples. We derive finite-sample error bounds on the learned generative distribution, measured in the Wasserstein-$p$ distance, for all $p\geq 1$. Specifically, given $n$ i.i.d. samples from $P_{\mathrm{data}}$, we show that, for every $d>d_p^\ast(P_{\mathrm{data}})$ and appropriately chosen network architectures and hyperparameters, the learned distribution $\widehat{P}^{\mathrm{FM}}$ satisfies $ \mathbb{W}_p(\widehat{P}^{\mathrm{FM}},P_{\mathrm{data}}) \lesssim n^{-1/d}+n^{-1/(2p)}\bigl(\log(1/ξ)\bigr)^{1/(2p)}$ with probability at least $1-ξ$, where $d_p^\ast(P_{\mathrm{data}})$ denotes the Wasserstein-$p$ dimension of the target measure. Our results demonstrate that flow matching naturally adapts to the intrinsic geometry of data and mitigates the curse of dimensionality, as the convergence exponent depends on the intrinsic rather than ambient dimension. These guarantees remain meaningful in high-dimensional regimes and provide a theoretical explanation for the empirical success of flow matching on structured data distributions under substantially more relaxed assumptions than those in existing analyses.
Chinese Translation
尽管流匹配模型在经验上取得了显著成功,但其统计泛化保证仍然不够完善。现有分析通常对估计的速度场施加限制性假设,并得到无法反映真实数据(如自然图像和分子几何结构)中常见的内蕴低维结构的收敛速率。在这项工作中,我们研究流匹配模型在从有限多个样本中学习未知分布 $P_{\mathrm{data}}$ 时的统计泛化。我们推导了学习到的生成分布的有限样本误差界,以 Wasserstein-$p$ 距离度量,对所有 $p\geq 1$ 成立。具体而言,给定来自 $P_{\mathrm{data}}$ 的 $n$ 个独立同分布样本,我们证明,对于每个 $d>d_p^\ast(P_{\mathrm{data}})$ 以及适当选择的网络架构和超参数,学习到的分布 $\widehat{P}^{\mathrm{FM}}$ 以至少 $1-ξ$ 的概率满足 $ \mathbb{W}_p(\widehat{P}^{\mathrm{FM}},P_{\mathrm{data}}) \lesssim n^{-1/d}+n^{-1/(2p)}\bigl(\log(1/ξ)\bigr)^{1/(2p)}$,其中 $d_p^\ast(P_{\mathrm{data}})$ 表示目标测度的 Wasserstein-$p$ 维数。我们的结果表明,流匹配自然地适应数据的内蕴几何,并缓解维数灾难,因为收敛指数取决于内蕴维度而非环境维度。这些保证在高维情形中仍然有意义,并在比现有分析宽松得多的假设下,为流匹配在结构化数据分布上的经验成功提供了理论解释。
cs.LG / 128 / 2610.03314
DAWIS: Data Assimilation with Windowed Inverse Sampling via Multitask Interpolants
DAWIS:通过多任务插值函数进行窗口化逆采样的数据同化
Erik Wikingsson, Martin Andrae, Tomas Landelius, Fredrik Lindsten
stat.ML · cs.LG · physics.ao-ph
diffusion
扩散模型相关
Abstract
Flow- and diffusion-based generative models have recently emerged as flexible and highly efficient forecasting models for dynamical systems. When combined with inference-time guidance, they offer a promising route to high-dimensional non-Gaussian data assimilation (DA), the problem of combining forecasts with observations to estimate latent system states. Existing filters, however, condition on a fixed history and assimilate only the most recent observation, leaving them unable to revise past states when new observations arrive. Estimates then stay tethered to a history that later observations may contradict, and errors accumulate over the assimilation run. To this end, we introduce **DAWIS**, a unified DA method covering filtering, fixed-lag smoothing, and block smoothing within a single framework. DAWIS replaces the single flow time of a state-level prior with a multitask stochastic interpolant over a window of consecutive states, assigning a separate flow time to each. An assimilation cycle inverts the window to a vector of per-state turning points and regenerates it under observation guidance, with the turning points controlling how strongly each state is held fixed, revised, or generated from scratch. The same construction can also absorb the forecast into the assimilation cycle, removing the need for a separate forecasting model. Experiments on challenging nonlinear systems show that DAWIS improves on both filtering and smoothing baselines under sparse, noisy, and nonlinear observations. The code for DAWIS is available at https://github.com/Erik-Wikingsson/DAWIS
Chinese Translation
基于流和扩散的生成模型近来已成为动力系统的灵活且高度高效的预测模型。当与推理时引导相结合时,它们为高维非高斯数据同化(DA)提供了一条有前景的途径,而该问题是结合预报与观测来估计潜在系统状态。然而,现有滤波器以固定历史为条件,并且仅同化最近的观测,这使它们在新的观测到来时无法修正过去的状态。于是估计值仍被束缚在一个可能被后续观测所矛盾的历史上,并且误差在同化运行过程中不断累积。为此,我们引入 **DAWIS**,一种统一的 DA 方法,在单一框架内涵盖滤波、固定滞后平滑和块平滑。DAWIS 将状态级先验的单一流时间替换为在一个连续状态窗口上的多任务随机插值函数,并为每个状态分配一个单独的流时间。一个同化循环将该窗口反演为一个由每个状态各自的转折点组成的向量,并在观测引导下重新生成该窗口,其中这些转折点控制每个状态被保持固定、修正或从头生成的程度有多强。同一构造还可以将预报吸收进同化循环,从而无需单独的预测模型。在具有挑战性的非线性系统上的实验表明,在稀疏、含噪和非线性观测下,DAWIS 优于滤波和平滑基线。DAWIS 的代码可在 https://github.com/Erik-Wikingsson/DAWIS 获取。
人工智能 (cs.AI)
120
cs.AI / 1 / 2610.02331
World Editing: Intervening on Executable Worlds at Increasing Depth
Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh, Ping Nie, Andy Zheng, Tat Hei Lai, Fei-Yueh Chen, Nikko Yu, Wei-Chieh Sun, Suzy Huang, Chiao-Wei Hsu, Chih-Chuan Huang, Chak-Wing Mak, Ho Yin Sam Ng, Edisy Kin Wai Chan, Min-Hung Chen, Ho Kei Cheng
cs.AI
Abstract
Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
cs.AI / 2 / 2610.02342
A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection
Ali Melih Kanca, Ilker Turker
cs.AI
Abstract
Natural Visibility Graph (NVG) based analysis characterizes network traffic through topological descriptors reflecting different structural properties. However, not all descriptors contribute equally to cyber-attack classification, and extracting a large metric set can increase computational cost. This study evaluates 21 NVG derived topological metrics and investigates whether a compact subset can preserve classification capability while improving computational efficiency. Four importance analysis methods SHAP, grouped Permutation Importance, Boruta, and Recursive Feature Elimination (RFE) are integrated through a Consensus Ranking strategy. Based on this ranking, Full21, Top15, Top10, Top7, Top5, and Top3 configurations are evaluated using the CICIDS2018 dataset, a CNN classifier, and stratified 5 fold cross validation. The three highest ranked metrics are avg_clustering_coeff_median, avg_clustering_coeff_std, and avg_clustering_coeff_mean. Top3 achieved the highest observed mean performance, with 97.148% accuracy, 97.055% weighted F1 score, and an MCC of 0.9675, compared with 95.999%, 95.521%, and 0.9549 for Full21, respectively. It also reduced total runtime from 14,961.39 s to 589.22 s (96.06%). These results indicate that importance guided metric reduction can provide a compact NVG representation with higher observed mean predictive performance and substantially lower computational cost under the evaluated setting.
cs.AI / 3 / 2610.02351
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
Ajay Vohra, Tao Chen, Neeti Narayan, Caron Zhang
cs.AI · cs.LG
Abstract
ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textsc{State} and certifies task completion. Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5--7.0 points for Qwen3-Coder-480B and 4.2--5.2 points for Claude Sonnet~4.5; gains diminish as Brain capability increases. Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. With Claude Opus~4.5, Pass@1 remains comparable to ReAct, while DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that completion control can trade earlier termination for stronger grounding. Overall, DeReAct improves weaker agents while retaining grounding benefits as models strengthen.
cs.AI / 4 / 2610.02378
THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS
Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu
cs.AI
Abstract
In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout (Oncorhynchus mykiss) in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient (AC) quantifying feeding intensity. Second, a Hierarchical Behavior Encoder (HBE) models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of explicit physical and implicit soft tokens. Finally, these tokens are integrated with environmental parameters, metadata, and expert rules to fine-tune an LLM via LoRA, followed by counterfactual multimodal Direct Preference Optimization (mDPO) to reinforce causal reasoning. Results show that AC exhibits a statistically significant monotonic positive correlation with expert-annotated feeding intensity (Spearman $ρ= 0.925$, $p < 0.001$). Ablations indicate that decision accuracy improves from 33.33% (text-only baseline) to 93.33% with dual-evidence tokens, confirming that continuous spatiotemporal tokens provide necessary physical grounding for LLMs. Compared with standard LoRA, counterfactual mDPO elevates decision accuracy from 93.33% to 96.67%, advances METEOR from 58.10% to 85.30%, reduces Self-BLEU-2 from 58.79% to 52.88%, and increases Distinct-3 from 6.68% to 7.81%, suppressing templating and actuation biases while reinforcing causal consistency and operational safety. Overall, by integrating continuous kinematics with LLM reasoning, this study provides a novel decision support paradigm for precision aquaculture.
cs.AI / 5 / 2610.02395
FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport
Felix X. -F. Ye, Yu Chin Fabian Lim, Naigang Wang, Davis Wertheimer
cs.AI · astro-ph.IM · math.NA
Abstract
Streaming GPU solvers for entropic optimal transport (EOT), such as FlashSinkhorn, avoid storing the dense kernel but still evaluate all $n\times m$ point pairs in every Sinkhorn iteration. We present \textbf{FlashSinkhorn~2} (FS2), a solver for squared-Euclidean cost on low-dimensional point clouds that solves large discrete EOT problems to a prescribed marginal residual on a single GPU by coupling two stages. A coarse stage solves on cell centroids, lifts the potentials to every point and, when a sampled marginal check rejects the lift, continues on the centroids, replacing most point-level updates. A block-sparse fine stage then removes the centroid error that coarse updates cannot. Its Morton-ordered blocks support screening and fused tensor-core execution, and a threshold set by the block masses bounds each omitted tile's contribution to every row and column. On synthetic benchmarks, FS2 reaches the target residual on all 32 problems and GeomLoss multiscale on 10. On one A100, FS2 solves discrete EOT between two $1.34\times10^8$-particle measures from a cosmological $N$-body simulation, at an entropic blur equal to the mean interparticle distance, to an all-particle marginal residual below 0.01 in under 2.5 hours. To our knowledge, it is the largest discrete EOT problem solved to this accuracy within hours. For reproducibility, we release an open-source implementation at https://github.com/ot-triton-lab/flash-sinkhorn
cs.AI / 6 / 2610.02405
When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
Xi Qin, Isabel Kurth, Xin Cui, Elin Park, Alexander Schaefer, Yaad Oren
cs.AI
Abstract
Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and context extension raise baseline solvability 5.6 times, but a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks. Adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, a strong evidence that the solvability band is model-specific. These findings demonstrate that meta-agent reliability requires solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria, not post-hoc diagnost.
cs.AI / 7 / 2610.02452
Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments
Armen Kasparian, Torri Jeske, Monibor Rahman, Chris Keith, James Maxwell, Thomas Britton, Malachi Schram, David Lawrence
cs.AI
Abstract
The operation of dynamically polarized targets in nuclear physics experiments relies on continuous tuning of the microwave frequency to compensate for radiation damage and evolving material properties, a task that is traditionally performed through manual trial-and-error by expert operators. This work presents a data-driven control framework that combines surrogate modeling with reinforcement learning to optimize the target polarization. Using operational data from the APOLLO cryogenic target system, we train and evaluate multilayer perceptron and Gaussian process regression models to predict polarization as a function of microwave frequency, beam current, and accumulated radiation dose. We show that Gaussian process-based models provide calibrated uncertainty estimates and reliably identify regions outside the training distribution, while MLPs exhibit limited sensitivity to distributional shift. To enable learning and control across multiple target samples, we introduce a Gaussian process approximation and embed the surrogate model within a standardized simulation environment. A reinforcement learning agent is trained using a lower-confidence-bound reward formulation that balances performance maximization against uncertainty. We are able to show an almost 2x improvement on the operators actions utilizing our RL agent.
cs.AI / 8 / 2610.02480
MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
Yuyang Cheng, Raghav Kaushik Ravi, Srivarshinee Sridhar, Sriparna Saha, Akash Ghosh, Chirag Agarwal
cs.AI
Abstract
Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools. To this end, we present MEA, a multi-agent framework that removes the explanation knowledge barrier entirely: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent is optimized end-to-end against faithfulness, transforming the outputs into natural language explanations grounded in model behavior across tabular, text, and vision modalities. Further, we introduce diverse question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. We find that frontier LLMs systematically produce unfaithful explanations. By optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with reward-driven optimization yielding faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. More broadly, our findings suggest that AI agents themselves can serve as a scalable, adaptable interface to ML explainability, opening a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field.
cs.AI / 9 / 2610.02492
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
Harry Lyu, Neil Thompson
cs.AI · cs.CL · cs.CY · cs.LG · econ.GN
Abstract
LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.
cs.AI / 10 / 2610.02496
"I just assumed that it would translate": examining MT risk awareness among healthcare staff with abbreviations as a use case
Eleanor Taylor-Stilgoe, Félix do Carmo, Constantin Orăsan
cs.AI
Abstract
In the UK, public healthcare staff report turning to machine translation (MT) - predominantly Google Translate (GT) - to communicate with patients across language barriers. Though intended to support their duty of care, potentially uninformed reliance on MT in such contexts could have serious consequences for patient safety. Research nonetheless remains limited on staff awareness of the possible risks posed by higher-stakes MT use in general and with patient medical records in particular, most existing literature instead examining its use in interpersonal situations or with patient-oriented documentation. Moreover, medical abbreviations are well-documented as increasing patient risk even monolingually, with outcomes from their misuse and/or misinterpretation ranging from temporary harm to the death of the patient. Abbreviations were therefore selected as a use case for identifying the potential risks posed by their translation with MT. Contextualised French and Spanish data examples drawn from authoritative clinical corpora and translated via GT were presented during semi-structured interviews to 21 healthcare staff participants in diverse roles and specialties. The results were then subject to qualitative analysis and cross-analysis.
cs.AI / 11 / 2610.02504
HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems
Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Yan Zhang
cs.AI
Abstract
Balancing electricity demand and supply is increasingly difficult due to the inherent intermittency of renewable power generation and the stochastic power consumption. Grid operators require fine-grained, decision-relevant insights into household energy consumption to manage peak loads and design responsive tariffs, but increased transparency at this level raises significant privacy concerns. Traditional methods for explainable AI (XAI) can reveal sensitive information, while standard privacy techniques often reduce the usefulness of explanations. To address this issue, we introduce HXAI, a hierarchical framework that preserves privacy while enabling reasonable explainable analysis for grid-level demand management. HXAI consists of two main components: (1) a local model that generates fine-grained explanations within a secure, private environment, and (2) a zonal model that aggregates these explanations to support grid-level analysis while enforcing privacy through flexible privacy-budget management. We explicitly limit cumulative privacy exposure under repeated operator queries and show that the proposed framework preserves decision-relevant information without compromising household privacy. Experiments on both simulated and real-world energy datasets demonstrate that HXAI provides useful insights for zonal load management while ensuring that appliance-level consumption remains local and is never transmitted to grid operators. Our results show that preserving the semantic structure of explanations, rather than minimizing numerical error, is the key to XAI under differential privacy. This framework provides a way to achieve both privacy and explainability in energy management.
cs.AI / 12 / 2610.02508
World Action Modeling with Progressive Visual Planning
Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, Amir Bar
cs.AI · cs.CV · cs.RO
Abstract
World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in https://sii-ferenas.github.io/ProWAM-page.
cs.AI / 13 / 2610.02525
Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
Ankur Samanta, Yonathan Efroni, Paul Sajda, Kaveh Hassani, Anirudh Goyal
cs.AI
Abstract
Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model's own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.
cs.AI / 14 / 2610.02542
How To Train Your World Model: Fine-tuning vs RAG for LM-based World Modeling
Dhananjay Ashok, Shantanu Agarwal, Vivek Datla, Jonathan May, Alfy Samuel
cs.AI
Abstract
World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based environments, fine-tuning a Language Model (LM) to serve as a WM has emerged as a dominant paradigm. However, despite the widespread success of non-parametric approaches such as Retrieval Augmented Generation (RAG), retrieval for LM-based world modelling remains underexplored. We conduct a systematic evaluation across five diverse environments spanning embodied, web navigation and social settings, comparing fine-tuning and RAG-based approaches for LM-based world modelling. Our study reveals that fine-tuning often outperforms RAG, with fine-tuned WMs enabling agents to obtain higher rewards on 15/20 settings. While both construction paradigms benefit from additional and more diverse exploration, RAG-based approaches prove more data-efficient, and fine-tuning approaches disproportionately benefit from scaling the amount of experience collected. With a focus on RAG-based WMs, we devise a procedure that uses counterfactual intervention to estimate the error rate of the retrieval stage, and show that retrievers consistently surface suboptimal transitions from the experience buffer. Hoping to address this failing, we study a variety of query reformulation strategies, demonstrating that a hierarchical approach outperforms the traditional retrieval pipeline. Finally, we compose our findings into a hybrid world modelling system that parametrically captures core environment dynamics, while learning to rely on retrieval from an actively maintained memory store. Our hybrid system consistently outperforms other methods across multiple environments and models, showcasing the robustness of the approach and the applicability of our findings.
cs.AI / 15 / 2610.02557
How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI Debate
Jiawei Li, Zhiyang Xun, Lijie Chen, Jonah Brown-Cohen
cs.AI · cs.CC · cs.GT · cs.LG
Abstract
As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and supervision of these systems has become increasingly urgent. One promising approach is AI debate, which seeks to leverage a debate between two powerful AIs to break complex questions down into simpler claims that can be easily judged directly. Theoretical work on debate has formalized this intuition in the language of computational complexity theory, where the goal is to design protocols (i.e., rules of the debate game) that provide rigorous guarantees on correctness for judging solutions to complex problems with limited supervision. Specifically, the current best protocol has been shown to work for all problems that have sufficiently stable decompositions into subproblems. In this paper, we design a new protocol for this same class of problems that improves on the prior work in several ways. First, correctness holds in a worst-case rather than an average-case sense. Second, being honest and correct is a dominant-strategy equilibrium for both debaters, rather than a Stackelberg equilibrium. Finally, we prove black-box lower bounds, showing that our new protocol is instance-wise optimal. That is, no protocol for this class of problems can outperform ours while making only black-box queries to human judgments. We obtain these results by relating the notion of stable problem decompositions to the concept of fractional block sensitivity from query complexity.
cs.AI / 16 / 2610.02568
Mitigating Social Sycophancy via Pluralistic Preference Optimization
Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou, Qian Hu, Rahul Gupta, Natasha Jaques, Emma Brunskill
cs.AI
Abstract
Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much more often than humans do, which can make people overconfident and less willing to repair their relationships after a conflict. Prior work on mitigating sycophancy has focused on factual settings where a response can be checked against a ground truth answer, while mitigations for social sycophancy (e.g., personal advice, where there is no ground truth) have relied on simple prompting and post-training methods with limited effectiveness. Our insight is that social sycophancy occurs in part because LMs overly center on the user and fail to consider the perspectives of other stakeholders impacted by the user's behavior. To address this problem we propose Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the LM identifies and simulates the relevant stakeholders, and is then trained to prefer and generate responses acceptable to all stakeholders. PlurPO uses only signals the model produces about its own outputs, without ground-truth labels. PlurPO substantially reduces social sycophancy across four datasets and four model families compared to prior methods. For example, on statements of intent to cause harm, where the users' actions should not be endorsed, PlurPO reduces the endorsement rate by 89% on average across four models. On general advice questions, where the target is to match the endorsement rate of human responses, it closes the gap by more than half, from 17.8% to 8.0% on average. The preference dataset constructed by PlurPO for an 8B model also effectively transfers to mitigating sycophancy in a larger (32B) model. Our results indicate that social sycophancy can be reduced by leveraging a model's own capabilities to simulate a plurality of relevant perspectives.
cs.AI / 17 / 2610.02576
Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models
Manan Roy Choudhury, Suparno Roy Chowdhury, Swastik Sahoo, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani, Mohamad Bassam Sonbol, Irbaz Bin Riaz, Vivek Gupta
cs.AI
Abstract
Systematic reviews condense clinical trials into evidence tables, yet clinicians can interrogate these tables only through database queries, and many questions concern attributes that the table does not record, such as a drug's target class or a harmonised endpoint. Here we introduce FD-SCoPE, a language-model framework that answers both kinds of question, exposes the query, the selected trials and the derivation rule behind every answer, and learns from expert corrections. On an oncology evidence table of 159 immune checkpoint inhibitor trial records, FD-SCoPE completed all 140 clinician-style tasks (alternatives, 90.7-97.9%). For questions needing derived attributes it retrieved 99.3% of relevant trial records at a positive predictive value of 89.8% and outperformed four alternative approaches (derived-value F1 77.7% versus 64.8-73.4%). Corrections on 299 questions, simulated from reference answers, raised F1 on 1,201 unseen questions from 77.9% to 84.9%. Language models coupled with executable queries, verified programs and expert feedback can give clinicians auditable access to trial evidence.
cs.AI / 18 / 2610.02586
Labels Override Definitions in Jev-Style Typed Decision Models
Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram
cs.AI
Abstract
A typed decision model answers a fixed question about an input by returning a probability for each of several caller-defined options. Each option carries a short label and a written definition, which is where a developer states the rule the model should apply. Jev introduced this interface for routing, moderation and triage, open implementations followed, and the same operation occurs whenever a language model is used as a classifier by scoring label strings. We study the open implementations, whose weights we can inspect and patch, and ask whether the probability follows the definitions or the labels. A preference for the label we call option-label bias. Across four open-weight typed decision models, three ways of reading an answer from a Qwen2.5 backbone, eleven classification tasks and PolicyBench, a synthetic routing suite we introduce in which the rule appears only in the definitions, the answer is mostly the labels. Deleting every definition leaves accuracy unchanged (laya-td: 0.8559 against 0.8487), although those definitions support 0.7971 on their own, and renaming the options to A and B raises accuracy by +0.1511 [+0.1377, +0.1646]. One system, von, is unaffected, and the two code bases differ in one expression: laya writes each option as "{label}: {definition}", while von writes only the definition. Changing that expression in both directions, with no weight changed, makes all three laya checkpoints exactly invariant (+0.0000 [+0.0000, +0.0000]) and creates the effect in von, whose accuracy falls from 0.8511 to 0.2281 when a label contradicts its definition. Earlier work attributed this failure to the constrained decision head these models use in place of a text decoder; our results locate it in the prompt rendering. We give a two-call test that tells a practitioner which case applies to their model, and measure what four mitigations are worth.
cs.AI / 19 / 2610.02588
Open-Endedness Bench: Measuring Epistemic Process from Agent Records
Chengyang Shi, Xianglin Ji, Jintao Huang, Jicheng Wang, Yifeng He, Jiachen Liu
cs.AI
Abstract
Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every step of this research, yet the runs are still judged by their outcome score. That score alone does not establish whether an agent's claims follow from executed experiments, and a reference answer may be unavailable. We evaluate the agent's epistemic process: how it forms hypotheses, tests them, and revises them in response to evidence. We introduce OEB (Open-Endedness Bench), a benchmark-agnostic methodology that reads only the agent's execution record and never a reference answer or an outcome score. OEB compiles the record into a unified epistemic event graph whose edges connect the propositions the agent states to the executed actions that test them; each node carries an exact excerpt that code verifies against the record. One principle governs scoring: prose can state a proposition, but only evidence returned by an executed action can support or refute it, so OEB checks what the agent writes against what it actually ran. From the graph, OEB scores four competence axes (evidence, experiment, revision, and no reward hacking), mostly as the share of opportunities for sound research that the agent took, and profiles six subjective persona traits that describe the agent's research habits. We score 119 existing runs over 12 tasks from three benchmarks: LLM post-training, chip design, and a training-speed record. Against logged results, only 16-29% of the improvements agents claim are real. On 9 of 10 tasks, the best run tries more new ideas in its second half than the worst run. The persona readings follow the model: for every trait, the model that ran explains more of its variance across runs than the task (a median of 43% against 7%).
cs.AI / 20 / 2610.02599
TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods
Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling
cs.AI
Abstract
Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials discovery. Evaluating whether a novel food tastes like its animal-based target requires expensive human sensory panels, bottlenecking the design-build-test loop. We introduce TasteBench, a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning two tasks: a food-level ranking task built on 21K+ human evaluations across 215 plant-based foods in 24 product categories, yielding 935 within-category ranking pairs, and a supporting molecular-level taste classification task over 15K flavor molecules. To enable rigorous interpretation of model performance, we characterize the ground truth: inter-rater agreement among panelists is low (Krippendorff's $α= .077$), and the split-half reliability ceiling of panel-aggregated rankings is .825, establishing the range within which ML systems on this benchmark should be assessed. We evaluate baselines across four input modalities; on the same pairs panelists rated, the best model achieves .661 pairwise accuracy, competitive with the median individual panelist (.650), and .683 across all within-category pairs. TasteBench provides the evaluation infrastructure and baselines for measuring progress on computational screening for sustainable protein discovery.
cs.AI / 21 / 2610.02608
Time Series Forecasting Benchmarks Need Scenario-Grounded Stress Testing
Yuyang Zhao, Lian Xu, Hao Xue
cs.AI
Abstract
Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly narrow: standard benchmarks reward low held-out error, while robustness studies typically reduce failure to Gaussian noise, random masking, or bounded adversarial perturbations. This obscures the real failure modes of deployed forecasting systems. Input-side anomalies are not merely noisier inputs: they often reflect structured events that alter temporal dynamics, break cross-variable dependencies, induce regime shifts, or propagate from faulty sensors to downstream decisions. These semantic, causal, and system-level failures cannot be faithfully captured by i.i.d. perturbations alone. The rise of TSF foundation models makes this evaluation gap more urgent, as unauditable pretraining corpora make held-out generalization increasingly unreliable. We therefore advocate scenario-grounded stress testing. Each test instance should include historical inputs and future targets, together with a semantic scenario, an explicit failure operator, and a measurable difficulty level. This shift makes evaluation interpretable, attributable, and deployment-relevant and friendly, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.
cs.AI / 22 / 2610.02616
VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy
cs.AI · cs.CL · cs.LG
Abstract
Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide our design. In a controlled study, optimizer self-evolution fails to improve performance without execution-based verification, but achieves the best result of that study when verification is available. Across five executors, self-evolving optimizers build their own tools for failure analysis, verification, training audits, and workflow control. Motivated by these findings, we introduce VERSE, a Verified Self-Evolving optimizer for agent harnesses. VERSE lets the optimizer test draft edits, replay failures, and perturb suspected steps before submission, while tracking fixes and regressions across rounds. Using this feedback, the optimizer revises both the executor harness and its own prompts, skills, tools, hooks, and notes, while the weights of the optimizer and executor models stay fixed. Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held-out SWE-rebench tasks and newer out-of-distribution tasks in five languages. Its best validation-selected harness reaches 42.3% and 37.7% accuracy, respectively, against 39.2% and 29.3% for the strongest baselines. Code is available at https://github.com/wzekai/VERSE.
cs.AI / 23 / 2610.02627
Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents
Feng Chen, Ritam Dutt, Atnaz Taheri, Alex Williams
cs.AI
Abstract
An email assistant should not complete less work simply because a user phrases the same request differently. Yet most benchmarks test each task with only one canonical request, leaving this form of robustness largely unmeasured. We test whether email assistants remain reliable when the requested information, available evidence, and expected outcome stay fixed, but the communication style or English variety changes. We construct validated variants along five communication-style axes and four rule-based dialect conditions, and evaluate them on three benchmarks: a retrieval-augmented generation (RAG) pipeline and two tool-using agents. Indirect requests reduce performance on all three benchmarks, while formal requests reduce performance on both agentic benchmarks. Examining the systems more closely shows that these failures have different causes. Verbose requests mainly hurt a lexical retriever by making the relevant email harder to find. By contrast, indirect and dialect variants remain harmful even when the relevant email is retrieved. In the agentic setting, indirect and formal requests mainly cause the agents to omit required actions, not to take more unsupported actions. These results show that a successful response is not enough to establish robustness: evaluations should vary how requests are expressed and separately measure whether agents complete the requested work.
cs.AI / 24 / 2610.02631
Designing the Future of User Feedback for Generative AI
Alisa Frik, Julia Bernd, Amitis Karami, Mohammad Tahaei
cs.AI
Abstract
Post-deployment feedback from users can be a cost-effective, scalable, and representative means to monitor and improve generative AI systems and features. When implemented effectively, giving such feedback can increase users' engagement with and trust in GenAI systems. Government regulations and industry guidelines call for post-deployment user engagement, but there is little guidance on designing mechanisms that are usable for consumers and provide actionable input for product teams. We conducted a multi-phase study as a collaboration between academic researchers and eBay. Our benchmark evaluation of current industry approaches identified common issues including lack of discoverability, unclear terminology, and inattention to user value. Based on these findings, we developed best-practice recommendations and designed and tested a prototype feedback-collection tool. The tool aimed to provide users with an efficient, flexible, and positive feedback-giving experience, and provide product teams with rich data on performance and potential problems in a usable format.
cs.AI / 25 / 2610.02638
Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript
Jie Jin, Ziyin Ma, Min Yin, Jinyu Chen, Haigang Song, Zhikun Pang, Xiaowen Zhang
cs.AI
Abstract
Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.
cs.AI / 26 / 2610.02654
Coherence-Driven Belief Formation and Population Dynamics of Contagion in LLM Agents
Tathagata Banerjee, Nima Moghaddas
cs.AI · cs.MA · cs.SI
Abstract
Models of social contagion usually assume how individuals adopt beliefs and derive population behavior from it. We instead empirically measure belief adoption in language model agents, quantifying the probability an agent adopts a claim given how many peers endorse it. We find this adoption kernel to be sigmoid, a characteristic of complex contagion, with a threshold that is sensitive to three sources: the claim's plausibility, the source's reliability, and the agent's disposition. These three dimensions are well approximated by a single effective dimension which we propose can be understood as the coherence of the incoming belief with the LLM agent's prior beliefs. Further, we observe a characteristic of complex contagion in the collective dynamics of belief adoption in a system of AI agents: further spread on clustered than random networks. These systems also exhibit a bifurcating cascade window, and self-sustaining hysteretic consensus which lead to consensus being far harder to remove than to establish.
cs.AI / 27 / 2610.02664
A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns
XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng
cs.AI
Abstract
Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety concern. Under benign interaction conditions, an agent may execute an action that violates a safety constraint specified many turns earlier. We term this failure mode Governance Hazard from Overlooked Safety Constraints across Turns (GHOST), which may cause irreversible damage. Our experiments reveal that GHOST events are not isolated cases: this failure mode, occurring precisely under benign interaction conditions, yields an occurrence rate of 11.5% on GPT-5.5. Furthermore, we theoretically show that if the residual conditional violation hazard along each safe prefix is bounded below by a non-summable sequence, the execution enters the hazard region almost surely. Leveraging this theoretical insight, we further propose STAR-Guard, a two-layer defense coupling historical semantic safety constraint restoration with pre-execution audit. STAR-Guard restores applicable safety constraints to reduce unsafe proposals, while its deterministic audit layer prevents residual violations from reaching the environment. Consistent with this two-layer design, we observe no GHOST events in our experiments under the GPT-5.5 setup.
cs.AI / 28 / 2610.02678
Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation
Xiang Chen, Futao Su, Kong Wang, Jiayi Chen, TanLin Li
cs.AI
Abstract
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Referenced On-Policy Distillation (SR-OPD), which reduces this cost by selecting which prompts and rollouts receive teacher supervision. When the student produces both successful and failed rollouts for the same prompt, a successful rollout can serve as a natural reference for selecting failed rollouts. SR-OPD therefore focuses on such prompts and prioritizes failed rollouts whose hidden-state trajectories show sustained divergence from a successful reference, while accounting for estimated teacher-input cost. Across three teacher-student pairs and six mathematical reasoning benchmarks, SR-OPD uses only 3.46-5.02% of the teacher-input tokens required by Vanilla OPD in the one-pass setting while maintaining comparable reasoning performance. Under a controlled setting matched to 5% of Vanilla OPD's teacher-input budget, further experiments support both key design choices: focusing supervision on prompts with both successful and failed rollouts, and using successful rollouts to guide failure selection. These results indicate that a student's own successful behavior can serve as a useful reference for allocating teacher supervision under a fixed teacher-input budget.
cs.AI / 29 / 2610.02679
DataWeave: Deploying Human-LLM Analytics for Exploratory Structured Data Analysis
Raquib Bin Yousuf, Harith Laxman, Vitaliy Shkremetko, Eunice Son, Shambhavi Verma, Brian O'Leary, Venketesh Subramony, Sylvain Nazef, Jacquelyn Elias, Ron Coddington, Chris Contakes, Michael Riley, Naren Ramakrishnan
cs.AI · cs.HC
Abstract
Data journalism, the practice of using data analysis to surface newsworthy stories, depends increasingly on the ability of reporters and investigative journalists to uncover trends, disparities, and accountability narratives. In practice, exploring large structured datasets remains slow and brittle: journalists must navigate hundreds of variables across many datasets over years, understand data coding conventions, and write non-trivial analysis code while hypotheses evolve. Although LLMs are often touted as "ask in English, get SQL/answers," real newsroom workflows expose recurring failures, e.g., schema mismatches and drift, misread domain semantics and units, and silent assumptions. We present DataWeave, a system that addresses these needs by combining conversational interaction, schema grounding, analytical planning, and executable query generation to support exploratory analysis over structured data. Rather than treating LLMs as autonomous answer engines, DataWeave frames them as interactive partners whose outputs can be inspected, corrected, and steered as hypotheses shift. We present a case study with professional journalists using our system to analyze the U.S. Department of Education's Integrated Postsecondary Education Data System (IPEDS), a high-stakes public dataset with substantial domain semantics and frequent schema updates. We also report how deployment experience and iterative refinement shaped the current DataWeave architecture and its analytical workflow. Our findings distill design principles and deployment lessons for trustworthy human-LLM collaboration in structured data analysis.
cs.AI / 30 / 2610.02704
Label-Efficient Time Series Classification at Scale: A Dual-Stream OSSE-LSTM with Counterfactual Attribution
Nguyen Ho, Bach Tung Tran, Trung Ky Nguyen, Zhenchang Xia, Bolong Zheng, Long Van Ho
cs.AI
Abstract
Time series are produced continuously at enormous scale by industrial equipment, wearables, power grids, and clinical monitors, yet annotation remains manual, expensive, and expert-dependent. The binding constraint in large-scale time series analytics is therefore not data volume but label volume, and the question facing a practitioner is concrete: how many examples per class must be labeled before a classifier becomes usable? We study this question directly, in a regime where the label space is fixed and known in advance and the decision rule must be constructed from only K labeled examples per class. We propose Dual-Stream OSSE-LSTM, an episodic metric-learning framework that pairs an Omni-Scale CNN with Squeeze-and-Excitation recalibration, for multi-scale motif extraction without per-dataset kernel tuning, with a Bidirectional LSTM for global temporal context. The two streams are independently normalized and fused into a prototype-oriented embedding. Because decisions taken from a few labels must also be explainable, we introduce Counterfactual Integrated Gradients (C-IG), which attributes the prototype margin between target and opposing classes rather than an isolated classifier logit, and reuses the resulting maps as soft masks for test-time prototype refinement without updating the encoder. On 19 univariate UCR datasets, OSSE-LSTM attains the highest average accuracy and per-dataset win count at every support size, and its accuracy remains within a 0.36-point band (96.36-96.72%) across that range. Its weakest configuration still exceeding the best result any compared baseline achieves at any K (93.99%).
cs.AI / 31 / 2610.02715
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
Qinchuan Cheng, Zhantao Gong, Pengzhan Sun, Angela Yao, Shijie Li
cs.AI
Abstract
Egocentric videos capture how people carry out everyday activities, yet testing an agent requires evaluating the consequences of actions it chooses itself. We introduce Ego2World, a benchmark that turns annotated cooking activities into executable planning environments under partial observation. Its compiler links source steps and objects to symbolic action rules, persistent world states, and explicit task conditions, so researchers can execute an agent's proposed actions and check their outcomes. World state and agent belief are maintained separately, enabling controlled studies of planning and information reuse across continuing tasks. Evaluating six planners on 105 tasks shows that accepted operations often leave task goals unmet. Execution traces and condition checks distinguish interrupted runs, partial attainment, and completed execution without goal attainment. In a separate paired Qwen-Plus study, persistent belief improves action validity by 4.15 percentage points and reduces visual-query attempts by 90.27%, with higher token use and no detected completion gain. Ego2World provides a reusable testbed for tracing how planning and memory choices affect execution, observation demand, and task attainment, connecting recorded human activity to the development and evaluation of interactive agents.
cs.AI / 32 / 2610.02741
On the Chain-of-Thought Monitorability of Looped Language Models
Han Wang, Ishwar B Balappanawar, Huan Zhang
cs.AI
Abstract
Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared transformer layers, increasing effective computational depth and enabling additional latent computation without increasing model size. However, the effect of looped architectures on CoT monitorability remains largely unexplored. In this work, we provide the first systematic evaluation of CoT monitorability in LoopLMs. We study two complementary settings: (1) varying the loop depth within the same LoopLM family to isolate the effect of additional recurrent computation, and (2) comparing LoopLMs with non-looped language models matched by parameter size, transformer-layer count, or effective depth to study whether LoopLMs are less monitorable. Across eight tasks from MonitorBench and both standard and stress-test settings, we observe task-dependent reductions in CoT monitorability under stress tests on specific Logic/Science/Engineering \texttt{Cue Answer} tasks, while other tasks exhibit weaker or qualitatively different trends. Our diagnosis suggests that these declines are not fully explained by task difficulty, verification pass rate, or generated token length; qualitative examples further suggest changes in how deeper-loop models explicitly use or attribute provided cues. Our cross-model comparison finds no evidence that LoopLMs are systematically less monitorable than non-looped language models matched on size or depth. Overall, our results suggest that deeper loop depth can reduce CoT monitorability in some tasks under stress tests, but looped transformer architecture alone does not necessarily imply lower monitorability.
cs.AI / 33 / 2610.02762
Dynamic LLM Routers are Often Misguided
Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang
cs.AI
Abstract
Dynamic LLM routers promise to cut inference costs by sending each query to the cheapest model that can answer it correctly. We analyze six commercial routers across 14 settings on a diverse benchmark spanning eight task categories, finding that none of them outperforms a router that randomly selects between two well-chosen models at matched cost. Some underperform by more than 10 percentage points. We trace this gap to four patterns prevalent across routers: difficulty blindness, length reversal, semantic matching, and roster suboptimality. We show that the first three are what the standard objective rewards: cost-accuracy Pareto efficiency on realized costs favors escalating moderately hard queries over the hardest ones, shorter queries over longer ones, and routing by a query's source over its difficulty. We also argue that the two assumptions that would justify large rosters, model granularity and model specialization, do not hold empirically. We propose an alternative evaluation methodology that does not reward these patterns, and as a proof of concept, we design a simple two-model router that avoids all four. Nevertheless, its gain over random routing is limited, because a well-chosen roster leaves little to route.
cs.AI / 34 / 2610.02793
PAPER2LLM++: Continual Self-Evolution of LLMs from Research Papers
Hongji Pu, Yilun Zhao, Wenpeng Yin
cs.AI
Abstract
Research on LLMs continually uncovers model limitations, their causes, and potential solutions. Yet these human discoveries remain largely disconnected from model evolution: an LLM does not automatically learn from new research about its own failures. We introduce PAPER2LLM++, a framework for continual self-evolution of LLMs from research papers. Rather than treating papers merely as knowledge to retrieve, PAPER2LLM++ uses the growing literature as a stream of evidence and supervision for model improvement. For each incoming paper, it extracts evidence-grounded findings, tests whether the reported limitation persists in the current model, and, when needed, converts the findings into candidate learning signals. A try-evaluate-commit procedure integrates an update only when it improves the targeted behavior without substantially forgetting prior improvements or degrading general capabilities. Across a sequential stream of research-discovered LLM failures, we show that models can progressively incorporate new findings while retaining earlier gains. PAPER2LLM++ thus takes a step toward closing the loop between human discovery and model evolution, enabling models to continually learn from research about their own limitations and improvements.
cs.AI / 35 / 2610.02796
Modeling Shared and Individual Structure for Cross-Subject Continuous Affect Regression from EEG-fNIRS
Xuan Wang, Bing Wang, Shuai Chang, Hao Yuan, Xinbo Qi, Xinyue Zhang
cs.AI
Abstract
Continuous, second-by-second valence-arousal estimation from physiological signals is typically studied in a subject-dependent setting, where the model sees labeled data from the same person it is later evaluated on. We study the harder zero-shot cross-subject variant on a synchronized EEG-fNIRS dataset: predict raw-scale ([1, 255]) valence and arousal trajectories for subjects whose labels the model never observes, given only their unlabeled EEG/fNIRS recordings while watching the same video stimuli as a disjoint set of training subjects. We decompose the affect trajectory into a structure shared across subjects who watch the same stimuli and an individual structure estimated for each test subject from a label-free EEG marker (alpha-band cross-channel synchrony), which rescales the shared trajectory around the scale midpoint. We validate the per-subject calibration mechanism on four independent axes: leave-one-subject-out correlation between the marker and each subject's true optimal gain, a functional-form comparison against non-linear alternatives, a repeated leave-4-out component ablation isolating each part of the pipeline's contribution, and a ceiling analysis bounding the remaining headroom for per-subject scaling. On held-out subjects, the model reaches an overall MAE of 25.96 / 22.80 across two evaluation batches (valence 21.94 / 19.6, arousal 29.98 / 26.0), well below EEGNet and ASAC-Net baselines reported for the same subject-independent split (raw scale score 60.6 and 55.0 respectively). We further report a systematic negative-result search across model architectures, feature representations, and prediction targets that found no signal able to improve on the single alpha-synchrony marker.
cs.AI / 36 / 2610.02800
BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration
Chence Yang, Ningxi Cheng, Arash Akbari, Qitao Tan, Qingchan Zhu, Ci Zhang, Changdi Yang, Yanzhi Wang, Wei Niu, Jinhui Wang, Jin Lu, Geng Yuan
cs.AI
Abstract
Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation, introducing non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce this overhead, yet still face trade-offs between draft quality, target quality, and storage efficiency. We propose BitNest, a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation. Instead of deriving a draft from a predefined target, BitNest first constructs a strong low-precision base and then recovers the higher-precision target through residual refinement, enabling both models to share a single physical weight representation. BitNest further extends this progressive-precision design to the KV cache for long-context inference. Across multiple 7B--8B edge-friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and delivers 1.48--1.61x end-to-end speedup over FP16 autoregressive decoding. On the LLaMA models supported by all representative self-speculative baselines, BitNest also achieves consistently competitive or higher decoding speedup.
cs.AI / 37 / 2610.02801
VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning
Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee
cs.AI · cs.LG · cs.RO
Abstract
Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.
cs.AI / 38 / 2610.02808
ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, Jun Xiao
cs.AI · cs.CL · cs.LG
Abstract
Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request support, verifier catalog, realized availability, resource accounting, online filtration, and post-trace scoring; a matched route contrast changes only the policy coordinate. ROUTEAUDIT adds three measurable objects to this contract. A contract lattice averages coordinate increments over every admissible bridge order and reports the resulting attribution together with its path sensitivity. A policy-independent response tape identifies paired sequential contrasts when adaptive policies reveal different observations. For incomplete matching, request-level bounds use whichever potential outcome remains observed and give a sharp finite-population interval. The protocol commits paid observations and ledger events before the oracle join and returns an attribution certificate for each comparison. On two held-out raw-tail caches, matched static SF+SA equals the cascade, assigning the apparent gains of 0.1797 and 0.1250 over full static to the verifier-set edge. On 1,319 held-out task requests, the learned and RLVR studies report quality 0.9522 and 0.9553 versus 0.9484 for matched static; the RLVR-static paired difference is +0.0068 with a request-paired interval $[0.0015,0.0122]$ and a training-seed-by-request hierarchical interval $[0.0006,0.0131]$. Controlled attribution recovery yields route mean absolute error 0.0011 and endpoint reconstruction error 0.0004. Factorial, bridge-order, and stochastic-provider studies evaluate the certificate interface; RLVR supplies a learned-policy stress test under the same identification contract.
cs.AI / 39 / 2610.02815
iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
cs.AI
Abstract
Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis, the stored history drifts substantially. Based on this observation, we propose iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning. iS-KV keeps a recent window exact while incrementally folding older states into bounded-rank representations. As the low-rank basis evolves, it synchronizes historical coordinates with the updated basis to maintain representation consistency. On DeepSeek-R1-Distill-Llama-8B, iS-KV achieves 82.6% accuracy at 4.06-fold persistent-KV compression, close to the original model's 83.6%. On Qwen3-8B, it achieves 89.2% accuracy at 5.64-fold compression. Under matched memory budgets, iS-KV consistently outperforms token-eviction baselines.
cs.AI / 40 / 2610.02824
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
Yuxuan Fan, Jaehong Yoon
cs.AI
Abstract
Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt's initial rubric as interpreted under each prompt's facts. We also adapt criterion weights at stage boundaries to better address observed policy errors. Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00--20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.
cs.AI / 41 / 2610.02826
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang, Ruhan Wang, Chengsong Huang, Fuxiao Liu, Haitao Mi, Jordan Boyd-Graber, LeoweiLiang
cs.AI
Abstract
Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.
cs.AI / 42 / 2610.02828
FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Daren Zha, Jun Xiao
cs.AI · cs.CL · cs.LG
Abstract
Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision-conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held-out and 59.43% OOD accuracy, compared with 64.47% and 57.03% for PB2, the strongest evaluated adaptive baseline. Three paired training seeds give gains of +2.42 and +3.19 percentage points over the contextual bandit on held-out and OOD evaluation. Under high behavior-deployment mismatch, policy-consistent risk lowers selected-decision ECE from 0.108 to 0.053; DCTC lowers it from 0.039 to 0.022 at matched acceptance; PRCC removes false-feasible admissions on an 18-action catalog ($0.197\rightarrow0.000$); and enabling all three components reduces trajectory failure from 0.181 to 0.083 in a factorial ablation.
cs.AI / 43 / 2610.02831
AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking
Wenteng Chen, Jiachen Zhu, Rong Shan, Tianyi Xu, Yuxiang Chen, Congmin Zheng, Teng Wang, Junjie Wu, Weiwen Liu, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
cs.AI
Abstract
Vision-language models (VLMs) are powerful listwise rerankers for multimodal retrieval, but high inference costs restrict them to evaluating small local candidate views. Existing multi-call strategies rely on fixed schedules, wasting expensive VLM calls on uninformative candidate pairs and easy queries. To address this, we propose Adaptive Multi-view Budgeted Elo Reranking (AMBER), an online, budgeted multi-view reranking framework that dynamically optimizes global resource allocation. AMBER treats fragmented listwise VLM outputs as local tournaments, using continuous Elo updates to maintain a lightweight global ranking state. Building on this, it allocates computation at two levels: dynamically constructing candidate views with high score ambiguity, and scheduling queries to maximize expected information gain. We show that each Elo update corresponds to a stochastic gradient ascent step on the Bradley-Terry log-likelihood, and provide a submodular information-theoretic motivation for the query-level allocation strategy. Experiments on CIRR, CIRCO, and PhotoBench demonstrate that AMBER achieves the strongest overall performance among the compared multi-call VLM reranking methods under comparable VLM-call budgets, while remaining effective in lower-budget settings. Our code is publicly available at https://github.com/wnlfc/AMBER.
cs.AI / 44 / 2610.02853
Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models
Omanshu Thapliyal
cs.AI
Abstract
Safety heads are lightweight classifiers attached to pretrained language models for flagging harmful inputs before generation. Their empirical detection performance has been studied, but their formal robustness properties remain largely unexplored. We ask when a State Space Model (SSM)-based safety head can be certified to produce the same prediction for all inputs within a bounded embedding-space perturbation. We prove that the answer turns on a single condition: the $l_\infty$ norm of the state transition matrix must satisfy $\norm{A}_\infty<1$ (the \emph{contraction condition}), which enables exact interval bound propagation (IBP) certification for linear time-invariant classifiers. When the contraction condition holds, the reachable output interval has bounded steady-state width and examples can be certified as robustly classified. When it fails, the interval grows exponentially with sequence length and certification is impossible at any practical perturbation radius. We enforce contraction with a hinge penalty and show on toxic comment data that certified fraction improves from 41\% to 59\%, with a sharp empirical phase transition at $\norm{A}_\infty=1$ matching the theory. Applying a contraction-regularized S4 head to jailbreak detection on JailbreakBench, we achieve a zero-shot transfer to AdvBench (DR=0.994) and HarmBench (DR=0.988). A logistic regression on mean-pooled Mamba-130M embeddings matches or exceeds the S4 head on every detection metric, confirming that harmful intent is already linearly separable in the embedding space. The S4 safety head's contribution is not superior discrimination but the formal certification that no probe-based approach provides.
cs.AI / 45 / 2610.02858
Harness-Aware Distillation for Small Language Model Agents
Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee
cs.AI
Abstract
Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting correctly on harness information. Standard distillation, however, imitates the teacher's full outputs and treats the harness as part of the input. We propose Harness-Aware Distillation (HAD), which focuses distillation on what the teacher adds beyond the harness. HAD complements on-policy distillation with two components: an action preference that contrasts the same teacher's actions with and without the harness information, scored after the student's own reasoning, and a validity check that drops preference pairs whose preferred action contradicts the harness records. We show that the contrast gives the student information that imitating the teacher alone cannot provide, and HAD needs no task rewards, success labels, or future information. Across multiple long-horizon agent benchmarks and models, HAD outperforms on-policy distillation baselines with the same fixed harness. Our analysis shows that HAD enters fewer unproductive loops and recovers from errors more often than the baselines, and suggests that it adaptively keeps learnable feedback in its weights while reading state information from the harness.
cs.AI / 46 / 2610.02897
Interpreting at Write Time: A Policy Ablation for Multi-Goal Agent Memory
Albert Sadowski, Jarosław A. Chudziak
cs.AI
Abstract
A long-running assistant cannot keep everything it has seen, so it summarises. Summarising is not neutral: what is kept is chosen against some notion of what the record is for, and that choice is made once, before anyone knows which of the user's standing goals will ask. Goals rarely disagree about what happened. They disagree about which parts of it were worth the space. Once the history is too long to re-read, the summary replaces the stream, and whatever it left out is gone. We ask what a memory should summarise for when it serves several standing goals at once. Three policies answer differently: summarise with no goal in view, write one summary covering every goal, or write one summary per goal and read them together. We compare them across several models and event streams, holding the read step fixed so that only the write differs. The goals do pull apart: summaries written for different goals overlap each other less than a summary overlaps a rewrite of itself. Per-goal summaries win on relevance, completeness and accuracy, and the all-goal summary loses even to the neutral one written at a fraction of its budget. Interpreting at write pays off, but only for the goal that later asks.
cs.AI / 47 / 2610.02902
LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs
Seoyeon Ye, Gayoung Kim, Jiyoung Hong, Sookyung Kim, Hyunsoo Cho
cs.AI
Abstract
Current analyses of LLMs' parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct response reflects genuine generalization or rote memorization, grounded in speculation rather than evidence. To resolve these ambiguities, we introduce LUMOS, a diagnostic framework that traces knowledge along the causal chain from training-data exposure to behavioral output, leveraging OLMo 2 with its fully transparent training corpus. By grounding analysis in verified exposure, we reveal that models internally encode rare facts with high separability (84%) yet fail to express them behaviorally (54%), though this retrieval gap narrows with scale. Furthermore, when models are asked to self-reflect on their own answers, they perform reliably on trained content (83%) but drop to random-baseline levels (49%) on unseen content. This collapse persists even under chain-of-thought prompting, which inflates confidence signals rather than improving calibration. Collectively, these findings demonstrate that incorporating the training-data axis into LLM evaluation transforms speculative diagnoses into verifiable claims, and we advocate that this axis should be a standard component of knowledge assessment in LLMs.
cs.AI / 48 / 2610.02910
Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM
Md Nurul Absar Siddiky, Liuwan Zhu, Yingfei Dong
cs.AI · cs.CR · cs.LG
Abstract
Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: router-gradient sensitivity, the sensitivity of the sequence loss to the gate weights that select an expert. Across five MoE architectures, we rank experts by each signal on 500 benign and 500 malicious prompts and measure refusal on 100 held-out malicious prompts under two budgets: equal expert counts and equal nominal malicious routing traffic (1%-5%). Under each of the two budgets, router-gradient selection reduces refusals more than activation in 24 of 25 conditions, and more than a ten-trial random mean in all 25. The largest effect is in OLMoE, where refusals fall from 34 to 9 of 100 prompts (73.53% relative) with no degraded outputs, indicating substantive compliance rather than broken generation. After matching expert counts in every layer, gradient selection still produces greater refusal reduction than activation in 23 of 25 conditions, with two ties. An exploratory cross-model analysis links larger malicious-versus-benign concentration gaps to greater peak gradient effects (rho = 0.90; exact two-sided p = 0.083, n = 5). Together, the results support gradient selection under the tested budgets.
cs.AI / 49 / 2610.02920
HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
Xiqiao Xiong, Moxin Li, Zhixin Ma, Ouxiang Li, Wenjie Wang, Fuli Feng, Xiangnan He
cs.AI
Abstract
Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief descriptions or a few attack examples in threat reports and preprints. To address this limitation, we introduce HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence through an adversarial interplay between safety-specification generation and attack-case generation. Safety specifications guide harness updates toward addressing identified safety vulnerabilities, while attack cases probe for remaining safety vulnerabilities after each update. By feeding evaluation outcomes back into both processes, HASTE enables harness evolution against emerging attacks beyond the initially observed evidence. Experimental results across multiple backbone models, attack types, and evidence forms show that HASTE consistently reduces attack success rates while preserving benign-task utility. The code is available at https://github.com/xxiqiao/HASTE.
cs.AI / 50 / 2610.02925
Positive-Unlabeled Learning for Agent Safety False Alarm Auditing
Xichen Yan, Chongyang Gao, Kezhen Chen, Guangyi Zhang, Jiaqi Wu, Lixu Wang
cs.AI
Abstract
Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtaining a reliable cutoff still requires substantial manual verification. In practice, a small set of verified-safe non-alarmed trajectories may be available while alarms remain unlabeled, naturally casting false-alarm auditing as a positive-unlabeled (PU) ranking problem. The key challenge is monitor-induced selection, since observed safe references are accepted by the monitor, while the hidden safe alarms of interest are precisely those it incorrectly flags, making the observed positives poorly representative of the positives to be recovered. To address this challenge, we propose a two-stage framework in which Trust-aware PU Supervision adapts safe references toward the alarm domain and protects plausible false alarms from excessive negative pressure, while Reliability-gated Rank Distillation consolidates consistent ordering preferences from multiple PU reference models into a single student. Consensus-guided Structural Refinement then improves the student ranking using hierarchical safe-reference support, alarm relations, and predicted reference consensus. The framework requires no alarm safety labels for fitting and leaves the underlying monitor unchanged. Across mainstream safety monitors, our method achieves a macro AUPRC of $0.6444$, outperforming eight evaluated PU baselines by 5.27--16.98 absolute percentage points; compared with PULDA, the strongest evaluated PU baseline, it recovers 33.3% more false alarms at a 5% review budget.
cs.AI / 51 / 2610.02932
When to Compile a Computer-Use Agent? Measuring Payback and Making Compilation Decisions for Token Efficiency
Yulong Ming, Jie Xu, Zihan Wu, Xiaohua Jia
cs.AI
Abstract
Compiling GUI procedures that agents execute repeatedly into programs can reduce their token costs. However, measuring payback and deciding when to compile have two challenges. First, compilation costs are uncertain because attempts can require repair and still fail to produce a usable program. Second, future reuse is unknown because tasks may stop arriving or GUI drift may stop the program from working. To address these challenges, we propose PACE (Payback-Aware Compilation from Experience), a system with a measurement protocol and an online compilation algorithm. The measurement protocol records successful and failed compilation costs, and compares agent and program execution costs on matched task inputs to estimate per-use savings and payback counts. Using these measurements, the online algorithm compares estimated future savings with compilation costs, including failed attempts, based on past task arrivals and compilation outcomes. It checks execution and compilation charges against a cumulative budget determined by observed task arrivals before allowing either action. Under stated action-cost assumptions, total cost after each arrival is at most $1+ε$ times the cost of running every task with the agent. For successful compilation attempts, estimated payback counts excluding source agent runs are 2-16 uses. In simulations using recorded task arrivals, PACE reduces token costs by 17.3% compared with ReAct, 24.9% with the AutoRPA adaptation, and 17.3% with the ToolPro adaptation on average ($ε=0.25$).
cs.AI / 52 / 2610.02945
Continual Graph Memory for Mathematical Research Agents
Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang, Jiachen Lu, Zhi Zhang, Xinjie He, Hyunsik Chae, Ethan Ji, Alexander K Taylor, Vigyan Sahai, Yiwen Kou, Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao, Wei Wang
cs.AI · cs.CL
Abstract
Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generating an enormous volume of intermediate proof results. Organizing these intermediate results throughout a long-horizon proof-search process and reusing knowledge gained from prior explorations remain major challenges. We present Ansatz, a mathematical research agent built around Continual Graph Memory, a graph-based, evolvable, cross-problem mathematical research memory system that explicitly organizes the entire proof search process and reuses information from exploration trajectories of previous problems. Specifically, we develop a unified graph memory that represents all intermediate exploration results, including facts, plans, and counterexamples, together with edges that explicitly represent the relationships among them; dependency-aware retrieval supplies precisely targeted local context; an evidence-sensitive curator updates the research frontier and distills lessons from prior attempts; and scoped recall surfaces earlier statements and negative findings for local re-proving rather than uncritical reuse. Experiments cover runs across all ten First Proof Second Batch problems, together with four component studies. Ansatz reports closure on all ten research tasks, demonstrating its ability to sustain and resume long-horizon mathematical search. Beyond these problems, Ansatz also produces solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348, and 488 without human intervention, and makes partial progress on several open problems, illustrating its strong ability to solve open mathematical research problems.
cs.AI / 53 / 2610.02979
RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction
Arya Hadizadeh Moghaddam, Mohsen Nayebi Kerdabadi, Chen Chen, Dongjie Wang, Zijun Yao
cs.AI
Abstract
Unstructured discharge notes in Electronic Health Records (EHRs) often carry signal complementary to structured medical codes, holding patient-specific evidence that standardized cohort-level codes alone cannot capture. However, this evidence in notes is frequently buried in lengthy, noisy text that is not intentionally written with any specific clinical prediction in mind. Summarization is an obvious mitigation, but generic summaries, tuned for fluency rather than the outcome, routinely omit decisive evidence while retaining plausible but uninformative detail. To this end, we propose RASPER, a Reward-Aligned Summarizer for Prediction in EHR, that optimizes note summarization directly against the downstream clinical task. RASPER employs a tunable LLM-based summarizer to extract task-relevant evidence from discharge notes and trains it via reinforcement learning from prediction feedback, using a reward derived from the downstream predictor's loss. To ground the summarizer, a longitudinal encoder converts structured codes into soft prompts that incorporate each patient's clinical context into note summarization. By rewarding the quality of the resulting multimodal prediction, RASPER encourages the summarizer to retain patient-specific evidence that complements, rather than duplicates, information captured by structured codes. RASPER consistently outperforms strong baselines on both readmission prediction and medication recommendation across MIMIC-III and MIMIC-IV.
cs.AI / 54 / 2610.02981
Safeguarding Mutual Correction in Source-Free Domain Adaptation via Cut Statistics
Seongjun Lee, Changhee Lee
cs.AI
Abstract
Source-Free Domain Adaptation (SFDA) aims to adapt a source-pretrained model to an unlabeled target domain without access to the original source domain. While early single-model approaches rely on self-refinement, they are inherently susceptible to confirmation bias and struggle to correct their own systematic errors. To overcome this limitation, recent methods introduce Vision-Language (ViL) models as external knowledge sources. However, these approaches operate in a largely unidirectional paradigm, using the ViL model primarily to supervise the source-pretrained model. This overlooks a key structural property: the two models exhibit distinct failure modes -- where one produces an incorrect prediction, the other may produce a correct one, creating a natural opportunity for mutual correction within the target domain. Yet, without ground-truth labels, identifying which model is correct on any given sample is non-trivial, and naively exchanging predictions risks propagating errors across models. To address this challenge, we propose SafeCut, a novel approach that leverages the cut statistic as a label-free measure of prediction reliability to gate cross-model supervision. Our approach dynamically controls both the direction and strength of supervision based on relative reliability, selectively amplifying true corrections while suppressing miscorrections on a per-sample basis. We further provide theoretical justification showing that this reliability-gated mechanism guarantees a net-positive correction signal. Extensive experiments across diverse SFDA benchmarks demonstrate that SafeCut achieves state-of-the-art performance, highlighting the effectiveness of safeguarding mutual correction in SFDA via cut statistics.
cs.AI / 55 / 2610.03017
Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations
David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus
cs.AI · cs.CL · cs.HC · cs.SD
Abstract
This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.
cs.AI / 56 / 2610.03020
DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
Yifei Tao, Xinyu Zhong, Henry Hengyuan Zhao, Fanyi Wang, Tengda Guo, Wentao Qiu, Ying Wang, Liujian Tang
cs.AI
Abstract
Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific agent memory implicit. Additionally, most prior works measure the model solely with final-answer QA over long interaction histories, making the assessment still incomplete and unreliable. To this end, we introduce DyadMem with the proposed new definition User-conditioned Relational Agent Memory (URAM). DyadMem jointly annotates user-side memory and URAM along the same multi-session trajectories, resulting in 6 memory categories. To summarize, it includes 3,065 episodes, 50,961 sessions, and 61,210 QA instances, with extensive session-level Capture and Update gold annotations, query-level Recall support, and two QA settings: Gold-Memory and Full-Pipeline. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong, yet Full-Pipeline QA drops sharply. Such a gap explicitly supports our fine-grained evaluation design. Additionally, several quantitative results further reveal low Capture recall, incomplete Recall, and unsafe-deletion issues arising from even the frontier LLMs. We further conduct a rigorous experiment to validate the effectiveness of our URAM and observe the positive effects for all 20 models. In summary, DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development.
cs.AI / 57 / 2610.03025
Verifiable, Articulable, and Tacit Components of Preference
Alexander Spangher, Sheldon Huang, Andreas Haupt, Noah D. Goodman, Diyi Yang, Daniel E. Ho, Sanmi Koyejo
cs.AI · cs.CL · cs.CY · cs.LG
Abstract
What makes a short story gripping; a news article newsworthy; or a math proof elegant? These constructs resist articulation or verification; their meaning is at least partially tacit. However, modern AI models are improved primarily via articulated constitutions, rubrics and verifiers (i.e. in RLAIF and RLVR); tacit components of preferences are typically understudied. We introduce a large, labeled preference dataset CreativePreferences, containing 2.8M texts labeled by 317M human preference judgments across 7 creative domains, with 42 benchmark tasks. We model these labels with executable programs, rubric banks and densely trained models (V, A and VAT, respectively). We observe robust articulability gaps, VAT-VA; and verifiability gaps, VAT-V; we estimate upper and lower bounds for each gap with a novel measurement approach that discovers articulable and verifiable metrics, identifies spurious variables and estimates the value of undiscovered metrics using capture-recapture. These gaps occur across all domains, even in domains traditionally treated as fully verifiable: correctness-centered domains (i.e. mathematics and software engineering) and claim- and novelty-centric domains (i.e. news, patents, peer review). The size of the gap varies based on domain (e.g. peer review and creative writing have the largest articulability gaps) and widens as more people take part in the judgment, consistent with Collins' collective tacit knowledge. We show two consequences: (1) on human generations, the full model more closely matches human preferences, often in disagreement with articulated criteria, and (2) in an analogy to Goodhart's law, articulating preference shifts it away from the tacit dimension. Articulability and verifiability gaps are consequential; we give recommendations on when tasks can be prompted; how learning mechanisms might improve; and when to leave judgments with humans.
cs.AI / 58 / 2610.03055
hacktrace: behavior-supervised detection of reward hacking during code generation
Hao Jiang, Xin Li, Annan Wang, Yichi Zhang, Weisi Lin
cs.AI
Abstract
A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug. Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail. We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervising shortcut behavior independently of exploit success substantially improves detection. We introduce HACKTRACE, a behavior-supervised monitor that reads the internal states the agent already computes while generating code. Reusing these states enables monitoring before a turn is complete, without additional language-model tokens or passes. Combining this evidence with static features of the final files achieves a mean per-problem AUC of 0.997 with 8 ms of monitoring overhead, improving both accuracy and latency over monitors that run the model again on an honesty question and answer. The same generation states also provide an inexpensive monitoring signal for reinforcement learning. With strong GRPO penalties, HACKTRACE reduces the cheating share of passing solutions from 82-91% to 1-5%, while retaining honest, correct solutions and maintaining high detection accuracy as the policy evolves. Our results show that both the supervision target and the source of monitoring evidence matter for turning accurate detection into a useful training signal.
cs.AI / 59 / 2610.03079
RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning
Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xiuying Chen
cs.AI
Abstract
Genuine embodied agency requires robots to turn continuous real-world experience into lasting, transferable skills. This demands continual learning that integrates new capabilities without eroding prior knowledge as tasks and environments evolve. Experience replay mitigates forgetting, but storing complete demonstrations becomes costly as tasks accumulate. World-action models offer a generative alternative, reconstructing past experience through joint predictions of actions and future observations. However, visually coherent rollouts may contain actions that cannot realize the predicted transitions, while new-task adaptation can disrupt previously learned behavior. RIFAR therefore combines reliability screening with drift-aware replay selection. It reconstructs trajectories from compact demonstration prefixes and uses a frozen inverse-dynamics model to assess action-visual consistency. Training first combines current demonstrations with the highest-quality screened trajectories. RIFAR then compares action predictions before and after this adaptation on identical historical inputs, reselecting trajectories with larger normalized drift from the same screened pool for continued training. Across three LIBERO suites and real-world experiments, RIFAR surpasses the previous state of the art in WAM-based generative replay. On LIBERO-Goal, it achieves 90.97 AUC while retaining only 320 historical time steps per task, approximately 4.9% of the steps retained using 50-demonstration replay.
cs.AI / 60 / 2610.03095
Peer Influence across Heterogeneous AI Models
Frida Nøhr Laustsen, Marie Haahr Petersen, Victoria Popa, Ariel Flint, Romualdo Pastor-Satorras, Andrea Baronchelli, Luca Maria Aiello
cs.AI · cs.CL · cs.CY · physics.soc-ph
Abstract
When two AI agents disagree, who persuades whom? As multi-agent systems increasingly combine language models of different families and sizes, the answer can determine which judgments survive interaction. Measuring persuasion as the probabilistic shift in an agent's decision after a single exchange with a dissenting peer, we test seven open-weight models across three language understanding tasks. We find that persuasion is strong: when models disagree, receivers often abandon their initial judgment after seeing a peer's answer and explanation. Surprisingly, however, neither standalone certainty nor model scale reliably predicts persuasion dynamics. Models producing almost perfectly consistent decisions in isolation can be among the most susceptible to persuasion, and small models can match larger ones as persuaders and resist their influence just as effectively. Furthermore, we show that the size of the shift depends more on the susceptibility of the listener than on the persuasiveness of the speaker. Persuasion patterns are therefore specific to each model pairing, with heterogeneity amplifying persuasion in some combinations and suppressing it in others, allowing a dissenting agent running a small model to overturn the judgments of a much larger one. These findings show that the behavior of interacting models cannot be inferred from their individual properties but must be evaluated in the combinations in which they will operate.
cs.AI / 61 / 2610.03098
Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression
Alberto Caron, Tianyu Cui, Dmytro S. Lituiev, Mangal Prakash, Artem Moskalev, Amina Mollaysa, Bo Zhai, Hirsh Nanda, Daniel M. Poole, Zhongyin Liu, Iman Farasat, Robert Davidson, Nikolay V. Manyakov, Tommaso Mansi, Scott Oloff, Rui Liao
cs.AI · cs.LG
Abstract
Codon optimization, the process of selecting synonymous codons to improve mRNA translation efficiency and protein expression, is central to therapeutic protein production and mRNA vaccines, yet it remains a hard problem. The design space is discrete and combinatorially large, precluding gradient-based methods, and existing tools rely on heuristic proxies (e.g., Codon Adaptation Index or GC-content) that poorly capture true expression. We introduce Latent-Space Codon Optimization (LSCO), which recasts this discrete problem as a continuous one by mapping sequences into the latent space of a pretrained mRNA language model, enabling efficient gradient-based search. LSCO combines four components: a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior from a protein-to-codon back-translation model, and constrained decoding for protein fidelity. On a real-world, wet-lab antibody expression dataset, LSCO outperforms simple frequency-based, as well as modern deep generative baselines in predicted expression, while retaining suitable biophysical properties.
cs.AI / 62 / 2610.03128
Trading Strategy Optimization via Textual Gradient
Chaoqun Yang, Qian Wang, Fengbin Zhu, Xinyu Lin, Bingsheng He, Roger Zimmermann, Tat-Seng Chua
cs.AI
Abstract
Quantitative trading strategy design aims to discover trading programs from historical data that remain effective in future markets, which can be viewed as a black-box program optimization problem. LLM-based textual gradients offer a promising approach by providing explicit optimization directions for iterative strategy refinement. However, directly applying textual gradients faces two challenges: (1) optimization is myopic, underutilizing experience from previous evaluations; and (2) aggregate backtest feedback overlooks temporal robustness, potentially favoring strategies that perform well only in specific market periods. To address these challenges, we propose TradeGrad, an experience-guided textual-gradient framework for robust trading strategy optimization. TradeGrad leverages accumulated optimization experience to estimate textual gradients and employs multi-scale revisions for both strategy exploration and refinement. It further introduces the Cross-Period Robust Objective (CPRO), which emphasizes performance in unfavorable historical periods to promote temporal robustness. Experiments on cross-sectional and time-series strategy design in Chinese A-share and U.S. equity markets show that TradeGrad achieves the best in-sample and out-of-sample performance across all four settings. Notably, its Chinese cross-sectional strategy achieves 27.99% annualized return, 12.19% maximum drawdown, and a Sharpe ratio of 1.63, approximately 68% higher than the CSI 300 benchmark. Further analyses validate the proposed components and show consistent improvements in both in-sample and out-of-sample performance throughout optimization. The code is available at https://github.com/transcend-0/TradeGrad.
cs.AI / 63 / 2610.03137
Keeping JEPA World Models Plannable When Little of the Frame Moves
Florian Strohm, Patrick Wagner, Jannik Schwab, Marco Huber
cs.AI
Abstract
Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.
cs.AI / 64 / 2610.03185
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu, Yue Zhang
cs.AI · cs.CL
Abstract
On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student's capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at https://github.com/HancCui/opd_hacking.
cs.AI / 65 / 2610.03198
KV$^2$: A Self-Refining KV Cache
Johannes Wesch, Danni Liu, Jan Niehues
cs.AI · cs.CL
Abstract
The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at https://anonymous.4open.science/r/KVsquared-0B97.
cs.AI / 66 / 2610.03213
Toward SLM-based agentic task-tool intent matching
Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Hervé Muyal, Marcelo Yannuzzi
cs.AI
Abstract
Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
cs.AI / 67 / 2610.03251
Learning a Fact Is Not Learning How to Retrieve It
Chaemin Jang, Jihee Kim, Dongman Lee
cs.AI
Abstract
A model trained on "The capital of X is Y" may produce "Y" after "The capital of X is" but fail after "The capital of X:". We call these different ways of eliciting the same fact request forms. To separate learning a fact from retrieving it, we train two models in two stages. In the first stage (request-form training), one model sees each fact in five forms and the other sees the same facts only as statements. In the second stage (target-fact training), both receive identical training on new facts, all as statements. Both then retrieve the new facts almost equally well from statements, but differ sharply on other request forms. Thus, a model can learn how to retrieve through a request form before it learns the facts. To understand this difference, we examine the hidden state immediately before the answer, which we call the context state. When given two different request forms for the same fact, the model trained on five forms in stage one produces more similar context states than the model trained on statements alone in that stage. Changing this state at retrieval time can enable or prevent retrieval of an already learned fact, and the same effect transfers across facts and factual relations, such as capitals and currencies. To test its role during learning, we change the context state only during target-fact training. This intervention changes later retrieval without intervention at test time. Together, these results show that later retrieval depends on earlier request-form experience and the context state during fact learning.
cs.AI / 68 / 2610.03273
EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
Geonwoo Bang, Dongho Kim, Moohong Min
cs.AI
Abstract
Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final step. Second, expert learning paths would be the natural cure for sparse-reward RL, but they do not exist in educational data, because student logs record what learners did, not what they should have done. We address both obstacles by importing a recipe from simulator-based demonstration learning in robotics: the knowledge tracing simulator is used both to synthesize per-learner expert demonstrations through evolutionary search and to train a deployment-free policy that distills these demonstrations into a feed-forward learner. Our framework, EVOL, instantiates this pipeline with an asymmetric actor-critic where the actor commits to deployment-realistic blind planning while the critic exploits the privileged simulator state during training. Across three datasets (ASSIST15, Junyi, and EdNet; 39-189 concepts) and path lengths L = 5, 10, and 20, EVOL surpasses 8 baselines spanning heuristic, sequential, RL, graph-enhanced RL, and LLM-enhanced methods. We further compare three imitation strategies (BC, AWR, and DAPG) and show that final performance is governed by the quality of evolutionary experts rather than by the particular imitation objective.
cs.AI / 69 / 2610.03296
JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs
Haoran Zhang, Dongjun Kim, Seohyeon Cha, Kevin S Chan, Ananthram Swami, Gustavo De Veciana, Haris Vikalo
cs.AI · cs.LG
Abstract
Complex reasoning queries can be decomposed into directed acyclic task graphs and distributed across heterogeneous LLMs, reducing latency through parallelism and enabling smaller models to solve complex tasks. In practice, however, the suitability of an LLM for a given subtask may be a priori unknown, and execution alone does not reveal output correctness. We propose JOVE, an online framework that jointly assigns executor LLMs and selects intermediate outputs for paid verification. Verification runs asynchronously and is used to improve future allocations, so the system must balance spending on execution now against learning for later. We study how to optimize this trade-off under a long-term budget and a per-query latency constraint, with stochastic, initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs. Online learning updates task-dependent estimates of LLM quality based on verification feedback, while an information-gain bonus incorporates the value of learning into allocation decisions. Under a natural set of assumptions, we establish sublinear quality-learning regret for JOVE. Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times.
cs.AI / 70 / 2610.03312
Optimal Planning in a Dynamic World
Devin Wild Thomas, Solomon Eyal Shimony, Wheeler Ruml, Erez Karpas, Shahaf S. Shperberg, Andrew Coles
cs.AI
Abstract
Background: We address the problem of planning when the set of feasible states or actions changes over time. For example, in the problem of path planning among moving obstacles (sometimes known as SIPP), the feasibility of being at a particular location can change as the obstacles move. Or, the action of boarding a particular train is feasible only while it is stopped at the station. This dynamism means that the optimal plan and its duration can change depending on when execution begins. In practice, execution start time is often unknown until planning has completed or another agent gives the go-ahead. However, most prior planning work either ignores dynamism or assumes a known start time. This makes it straightforward to assess state and action feasibility but is impractical for some applications. Objectives: In this paper, we relax the assumption of a known start time. We define the setting of {\em any-start-time planning} and provide algorithms for it. Methods: We present a data structure called a compound arrival time function (cATF) that compactly encodes the optimal plan as a function of start time. We provide general-purpose planning algorithms, based on heuristic graph search, that assemble cATFs by propagating functions along edges instead of scalar costs. Results: We prove that the size of a cATF is at most linear in the problem size. An experimental evaluation of an implementation for the specific problem of SIPP shows that, on difficult problems, agents that rely on replanning often fail, while any-start-time algorithms using cATFs can quickly look up the optimal plan once the execution start time is known. Conclusions: By enabling efficient representations and reasoning for time-dependent plans, this work provides a foundation for planning in dynamic worlds.
cs.AI / 71 / 2610.03315
Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents
Linh-An Phan, MingXue Wang, Guangyu Wu, Feng Pan, Zhaoyu Pang, Yanbin Zhang
cs.AI
Abstract
Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equally useful for diagnosis. We present \textit{LiteTrajEval}, a lightweight architecture for budget-bounded trajectory evaluation. LiteTrajEval derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports. Evaluated on public Magentic-One-style and $τ$-bench-style trajectory datasets, LiteTrajEval improves failure-localization alignment with human annotations by roughly 20--35 percentage points on Magentic-One and up to 23 percentage points on $τ$-retail compared with AgentRx, while reducing cost by about 6$\times$ and evaluation time by more than 8$\times$. This solution has also been deployed in our enterprise agentic platform.
cs.AI / 72 / 2610.03363
Geometry Meets Physics: Data-Efficient Pre-Training for Unstructured Neural PDE Solvers
Luis Medrano-Navarro, Giacomo Baldan, Qiang Liu, Benjamin Holzschuh, Jan Hagnberger, Mathias Niepert, Nils Thuerey
cs.AI
Abstract
Neural surrogate models for Partial Differential Equations (PDEs) on unstructured 3D geometries are often limited by poor generalization and the high cost of generating large-scale training datasets. Consequently, pre-training on massive datasets of related PDE dynamics has emerged as a critical alternative to enhance the robustness and scalability of these models. However, this strategy is neither compute- nor data-efficient, as it relies on massive pre-computed data that is very costly to generate. In this work, we introduce a disk-data-free pre-training framework tailored to both steady-state and transient regimes. For steady-state problems, we propose a geometry-driven strategy that leverages intrinsic shape descriptors to learn representations of complex 3D domains. For transient problems, we introduce a physics-driven approach based on online generation of synthetic PDE data, enabling scalable pre-training without reliance on expensive datasets. Across multiple experiments, our approach achieves faster convergence, greater data efficiency, and higher accuracy during fine-tuning, particularly under realistic low-data regimes. This methodology provides a practical pathway toward data-efficient neural emulators for large-scale simulations.
cs.AI / 73 / 2610.03367
Multilingual GSM-Symbolic: What determines capability transfer across languages?
Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni, Isaac Chung, Sofie Bruun, Ayush Sunil Munot, Max Müller-Eberstein, Adnan El-Assadi, Elisa Bassignana, Gianluca Barmina, Hafsteinn Einarsson, Iben Nyholm Debess, Linda Freienthal, Lukas Galke Poech, Mike Zhang, Nicolas Legrand, Vladimir Salnikov, Yevhen Kostiuk, Zafar Hussain, Sagandeep Kaur, Agnes Toftgård, Marie Mattson, Kristoffer Nielbo
cs.AI · cs.CL
Abstract
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size ($β= 1.77$), language resource level ($β= 0.77$), reasoning ($β= 0.67$) and typological distance ($β= -0.25$). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages ($β= -0.27$ and $β= -0.20$, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.
cs.AI / 74 / 2610.03387
Benchmarking Candidate Coverage in Typed Decision Models
Jiawen Lu, Tongtong Wu
cs.AI · cs.CL
Abstract
Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initial evaluation of Laya and Jev across AG News, DBpedia, Emotion, and TREC. The models receive identical frozen texts and requests: 300 calibration and 589 test texts yield 23,932 predictions per model. Present/absent pairs match ordinary candidate count, and name variants preserve descriptions, members, and order. Native rejection behavior differs sharply: at five TREC candidates with natural names, Laya detects 97.2% of missing-answer cases but falsely rejects 69.7% of present controls; Jev's rates are 24.8% and 0.0%. Calibration-only none-score thresholds change these rates to 33.9%/3.7% and 45.0%/1.8%, respectively. On DBpedia, Jev's high coverage-score AUROC supports a stronger operating point, whereas both models have weak complete-set accuracy on Emotion. Competence-conditioned analysis, probability-precision sensitivity, and interface audits show why classification, score ranking, and rejection policies need separate measurement. This initial benchmark is descriptive and limited to reference-label omission; it does not establish natural out-of-scope generalization, causal mechanisms, or a new rejection method.
cs.AI / 75 / 2610.03425
Becoming Suspicious Across Borders: Algorithmic Extraterritoriality and AI-Driven Financial Surveillance
Georgios Pavlidis, Savvas Chatzichristofis, Eleni Gavriil
cs.AI
Abstract
Suspicion is an important, yet elusive concept in anti-money laundering and counter-terrorist financing (AML/CFT), which allows for intervention below the threshold of proof. In its traditional form, suspicion can be understood as a situated legal judgement by human actors within identifiable jurisdictions. It is argued that this understanding is no longer adequate. As artificial intelligence (AI) becomes an integral part of financial surveillance, suspicion is increasingly produced through data-driven processes. This transformation is epistemic, but also spatial. Since AI-driven financial surveillance operates through transnational data infrastructures, regulatory reach is less a matter of where conduct occurs than a question of whether such conduct becomes visible within data systems. This article develops the concept of algorithmic extraterritoriality, understood as a form of regulatory power mediated by data infrastructures rather than formal assertions of jurisdiction. Moreover, since individuals are increasingly constituted as datafied subjects of suspicion, they are rendered governable through dispersed and opaque processes of evaluation. This constitutes a challenge for accountability and contestability because suspicion becomes more difficult to locate, explain or contest.
cs.AI / 76 / 2610.03458
A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control
Zhe Zhou, Tianhua Tao
cs.AI · cs.CL
Abstract
Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low monitor readout does not identify whether such an intervention controls behavior. In a code-generation environment whose dominant exploit is available at the start of the reasoning trace, we train policies against three monitors that pass the same offline gate: an in-domain activation probe and two penalties conditioned on how early the policy commits to its own final answer. The probe score is at its numerical floor from the first recorded training step, and the trained-score median is zero for every prefix-trained run at the endpoint. These readouts estimate different quantities, and we do not compare their scales; within each monitor family, however, low values do not establish behavioral control. Within one fixed configuration, prefix-trained runs with the same zero-median trained score range, by seed alone, from a mixed regime with a low hacking share to near-pure reward hacking. All probe runs reach the hacking regime, but their floor-level readout reflects a mismatch between the position where the probe was validated and the position where it was read during training, not a second instance of this ambiguity. Text-level analysis identifies a prefix failure mode: generic planning and filler shells postpone the exploit past the cut without eliminating it from the final output. Low measured commitment therefore does not distinguish a low hacking share from delayed commitment to the exploit. Offline discrimination and low monitor-aligned readouts are insufficient evidence of behavioral control; an out-of-band behavioral check is required. We characterize the endpoint readout, not its evolution. Code is available at https://github.com/zhezhou1106/spoof-cost.
cs.AI / 77 / 2610.03519
Reasoning Models Are Accurate but Unsound on Identification
Arman Behnam, Binghui Wang
cs.AI · cs.CR
Abstract
A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model's training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at https://anonymous.4open.science/r/certid-D718.
cs.AI / 78 / 2610.03524
From Benchmarks to Production: A Text-to-SQL System for Complex Financial Data
Arijit Sehanobish, Bruno Gomes Coelho, Guillaume Michel, Sophia Zhi, Valerie Faucon-Morin, Kristen Howell
cs.AI · cs.DB · cs.LG
Abstract
General-purpose Text-to-SQL systems achieve strong performance on academic benchmarks like Spider and BIRD, where schemas are relatively shallow and column values are often human readable. In production financial databases, where concepts are stored as opaque integer keys rather than human-readable strings, these methods fall below 50%, as even simple queries require multiple joins and filter predicates reference opaque IDs. We present Financial LINking Text-to-SQL (FLINT), a domain-specialized Text-to-SQL system that closes this gap through three key components: (1) a lookup agent that dynamically resolves natural-language concepts to question-specific reference table constraints, (2) embedding-based retrieval of structurally similar query templates from a compact, expert-authored bank, and (3) schema linking that prunes a large table schema to the relevant subset by traversing foreign-key chains, rather than relying on name similarity alone. We evaluate on two datasets totaling 359 questions over production financial schemas. FLINT outperforms various state-of-the-art baselines using the same LLM. The system is deployed in production as part of a financial data retrieval service.
cs.AI / 79 / 2610.03548
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
Wenlong Zhang, Zhengbo Jiao, Chenxu Zhang, Lekang Jiang, SiYuan Ma, Qituan Zhang, Guo Chen, Linfeng Zhang
cs.AI
Abstract
Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver accuracy decreases from 100.0% to 54.8% over fourteen evolution rounds. Ablations show that combining both update schedules produces harder tasks than fixed-harness recursion or either schedule alone. The resulting data improves downstream SFT and GRPO performance. In particular, a 27B student fine-tuned on 10K synthesized mathematics examples achieves 62.5% mean-16 accuracy on APEX, competitive with selected frontier-model references. These results support adapting the synthesis harness alongside the tasks to generate increasingly challenging data with downstream training value.
cs.AI / 80 / 2610.03564
Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows
Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun
cs.AI
Abstract
Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation suite of 2,603 point in time episodes across 12 subtasks in portfolio construction, risk management, and fundamental analysis, with hidden regenerable ground truth and task specific deterministic verifiers. Executing 17,820 episodes across 9 models and 3 resource conditions, the paired analysis across 8 models shows that curated skill packages raise mean scores by +16.2 points (0.366 to 0.528), whereas skills generated within a single episode add only +0.5 points while consuming more tokens and turns. We then decompose the curated premium by granting human authored procedural documents and executable domain tools separately: documents alone add +5.6 points, tools alone add +19.5 points, and their combination is subadditive. The premium is strongly workflow dependent: executable tools dominate numerically intensive workflows, documentation matters more when procedural or output schema guidance is the bottleneck, and interpretive tasks benefit from both. The effects are sign stable across 10 scoring variants and cluster bootstrap analyses, and an independently implemented second harness reproduces the directional pattern while showing that effect magnitudes depend on how tools and data are exposed. Overall, a measured "skill premium" is a property of the full model, resource, and harness system rather than of the underlying model alone.
cs.AI / 81 / 2610.03570
Learning to Assess Heartbeat Observability for mmWave Heart-Rate Sensing
Yuxuan Hu, Shilin Shan, Jianfei Yang, Feng Xu
cs.AI
Abstract
Contactless heart-rate sensing with millimeter-wave (mmWave) radar requires assessing whether individual measurements support reliable estimation. We study learning to assess heartbeat observability, defined as the readability of the heartbeat component in an acquired phase spectrum, for selective heart-rate estimation. Coherent superposition of scatterer returns can suppress this component even under similar macroscopic observation geometry, motivating assessment directly from acquired measurements. To obtain training supervision across different observability conditions, we develop a controllable multi-scatterer frequency-modulated continuous-wave (FMCW) simulator. Agreement between the dominant heartbeat-band peak and the known heart rate provides an automatic observability label for each simulated measurement. We propose HEAR (Heartbeat Estimation with Assessed Reliability), a compact dual-task Transformer that jointly predicts an observability score and heart rate. Its input combines spectral magnitudes with frequencies relative to the respiration fundamental, providing context for respiratory harmonics. Trained solely on simulated observations, HEAR transfers zero-shot to two public real-world datasets collected at 60 and 120 GHz from 134 subjects. The same learned score supports selective prediction with both HEAR's own heart-rate head and multiple existing estimators. On the 120 GHz dataset, score-based selection reduces the HR head's mean absolute error from 17.9 BPM at full coverage to 1.6 BPM at 50% coverage. The complete pipeline achieves an end-to-end processing latency of 50.8 ms on an edge device. Project page: https://yuxuanhu9.github.io/HEAR/.
cs.AI / 82 / 2610.03574
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina
cs.AI · cs.LG
Abstract
We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.
cs.AI / 83 / 2610.03618
Low-Cost Video--Time Priors as a Strong Baseline for EEG--fNIRS Emotion Regression on Familiar Videos
Minghao Kong, Jiurun Chen, Ying Gao, Xiangbin Meng, Rongjie Wang
cs.AI · cs.CV · cs.MM
Abstract
Continuous emotion regression estimates moment-to-moment valence and arousal while a viewer watches a video. In familiar-video deployment, responses fron training participant-specific estimate, and prior-dominating fixed fusion tests whether physiology adds residual correction. In five-fold subject-held-out evaluation on 24was within 0.05 and 0.32 MAE of fusion in the internal and external evaluations, respectively. Source-explicit ablations showed that video identity and within-video tine accounted for most of the reduction, while EG-FNIRS gains were smaller and varied across participants and videos. These results identify the video-time prior as a strong, low-cost baseline and position EEG-fNIRS as an optional residual signal for familiar-video emotion regression.
cs.AI / 84 / 2610.03631
NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents
Lijie Ding, Changwoo Do
cs.AI · physics.ins-det
Abstract
Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder's partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.
cs.AI / 85 / 2610.03634
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng
cs.AI
Abstract
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.
cs.AI / 86 / 2610.03651
MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search
Sean Culatana, Shang-En Huang, Kang Li
cs.AI · cs.IR
Abstract
Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and {4, 8, 16}-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.
cs.AI / 87 / 2610.03693
Transcriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies
Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. Esteva, Rohit Bhargava, Marcin Braun, Kamila Paździerz, Jakub Czerwiński, Hanna Romańska-Knight, Albert Grinshpun, Bareket Daniel, Michele Buchinger, Frederick Howard, Piotr Wysocki, Brie Chun, Freya Schnabel, Rich Caruana, Jan Witowski, Krzysztof J. Geras
cs.AI
Abstract
Scarcity of labeled data limits development of deep learning biomarkers in oncology. We develop a two-stage AI model predicting pathological complete response (pCR) to neoadjuvant therapy in breast cancer. The first stage learns the transcriptome from histopathology using 8,742 patients across 32 cancer types, corroborated by pathologist review and spatial agreement with measured expression. This simplifies the second stage to predicting pCR from inferred expression and clinical variables. Developed using 1,080 patients (five cohorts) and evaluated in 1,412 patients (nine cohorts), the model achieves a pooled AUROC of 0.79 (95% CI, 0.73-0.85), discriminating responders within molecular subtypes. It outperforms histopathological biomarkers, remaining stable across intratumoral sampling and with minimal biopsy tissue. Ablations show transcriptome-wide inference improves discrimination over clinical variables alone or one-stage pathology models, and robustness by avoiding genomic assays' gene selection constraints. These results indicate that biologically informed compression may generalize to data-sparse applications in precision oncology.
cs.AI / 88 / 2610.02320
DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar
cs.CV · cs.AI · cs.LG
Abstract
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: https://saidgurbuz.github.io/deskforge/
cs.AI / 89 / 2610.02375
EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision
Ruiyang Hao, Zhi Qin Tan, Yulan He, Owen Addison, Yunpeng Li
cs.CV · cs.AI
Abstract
Dento-maxillofacial cone-beam CT (CBCT) reports may contain dozens of tooth-specific, anatomical, and spatial findings from a single 3D scan. Learning to generate such reports from limited clinical data is challenging because routine reports may not exhaustively document image findings, and a non-mention may reflect either absence or non-reporting. We present EviDent-CBCT, an evidence-bottlenecked framework designed for this incomplete supervision. An anatomy-aware network maps each CBCT scan to a discrete record of tooth-level, global, and tooth-IAC evidence. A dental-logic consistency projection reconciles incompatible evidence before a deterministic renderer and an image-blind local language model generate the report using only this record. For tooth-level evidence, reliability-aware training uses eligible non-mentions as reduced-weight negatives, while unreported global and tooth-IAC labels remain unknown. A metal-sensitive input channel preserves intensity cues from dental materials. Across three validation runs, EviDent-CBCT achieves $0.666\pm0.006$ merged evidence set-F1 and $0.402\pm0.003$ RadFact-Lite-Dental logical-F1, versus $0.371\pm0.018$ for the strongest controlled direct baseline. In the ODIN 2026 challenge, it ranked second in automated evaluation and third in blinded clinical Arena comparison on the hidden test set. These results support the discrete evidence record as an effective and auditable interface for CBCT report generation.
cs.AI / 90 / 2610.02718
Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis
Peilin Yang, Xiaoyu Liu, Jian Sun, Qinghua Tao
cs.CV · cs.AI · cs.LG
Abstract
Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.
cs.AI / 91 / 2610.02967
Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh
cs.CV · cs.AI
Abstract
Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (https://arena.ai/), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.
cs.AI / 92 / 2610.03015
OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection
Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang
cs.CV · cs.AI · cs.RO
Abstract
Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes geometry and visual information differently, making object-relevant cues hard to model, localize, and preserve. We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. To resolve geometric mismatch, the ERP-Ray Geometry Adapter (ERGA-Ray) models spherical viewing rays and periodic spatial structure. To localize evidence in scene-wide context, the Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence and converts it into a structured geometric action. To recover local cues lost under fixed token budgets, the Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution for heading estimation. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. With target-specific geometry adaptation, VARC retains 95--98% of the same-configuration mAP, indicating reusable object-level 3D reasoning across sensing configurations. The source code will be made publicly available at https://github.com/FeiT-FeiTeng/OmniAct3D.
cs.AI / 93 / 2610.03084
NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models
Omar Elfatairy, Maria A. Bravo, Jessica Bader, Zeynep Akata
cs.CV · cs.AI · cs.LG
Abstract
Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.
cs.AI / 94 / 2610.03099
Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning
Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng
cs.CV · cs.AI
Abstract
E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. Cross- video reasoning requires models to locate fine-grained evidence among many redundant frames and integrate visual details, speech, and on-screen text. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. To address the sparse credit assignment of reinforcement learning, we develop an offline trajectory rectification mechanism that identifies reasoning errors and missing multimodal evidence in RL-generated trajectories. The corrected trajectories provide supervised fine-tuning signals that reduce biases learned during RL. This mechanism supports a rectified bootstrapping pipeline in which initial RL exposes reasoning bottlenecks, supervised fine-tuning corrects them, and a final RL stage further improves the policy. AdSeek achieves 74.30 percent accuracy on the AdsCVR test split, outperforming its Qwen3-VL-8B-Instruct backbone by 27.90 percentage points. It also generalizes to the open- domain CrossVid benchmark, demonstrating effective active evidence gathering.
cs.AI / 95 / 2610.03123
Foresight: planning future perception in streaming VLMs without retraining
Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari, Samip Ghimire, Saroj Poudel, Binod Bhattarai, Danda Pani Paudel
cs.CV · cs.AI
Abstract
Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.
cs.AI / 96 / 2610.03202
Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation
Divya Jyoti Bajpai, Arun Verma, Manjesh Kumar Hanawal
cs.CV · cs.AI
Abstract
Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose COFLOW, an inference-time method that adaptively selects the step counts each generation based on the prompt features. Our context-aware COFLOW is trained online with an unsupervised reward that balances inference efficiency and generation fidelity. Our method is plug-and-play, requiring no retraining of the underlying generative model. It generalizes to image and video generation, achieving over 2.5x speedup while preserving perceptual and semantic quality. We further provide a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.
cs.AI / 97 / 2610.03403
ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation
Zhihao Zhan, Le Tao, Yifei Tian, Xin Liu, Jie Yuan
cs.CV · cs.AI
Abstract
Forest point cloud segmentation is fundamental for fine-grained 3D forest scene understanding, yet remains challenging due to irregular tree structures, severe occlusions, density variations, and ambiguous instance boundaries. Recent query-based forest segmentation methods have shown promise for unified semantic and instance prediction, but they still insufficiently exploit forest-specific spatial structure and account for boundary uncertainty. In this paper, we propose ForestQuery, a boundary-aware and spatially anchored query learning framework for unified forest point cloud segmentation. ForestQuery enhances instance and semantic query learning through two complementary designs. Specifically, boundary uncertainty is explicitly modeled to guide reliable instance query construction and modulate query optimization through adaptive loss reweighting. Meanwhile, spatially anchored semantic query enhancement (SA-SQE) introduces learnable 3D anchors encoding forest vertical stratification priors to enrich semantic queries with explicit spatial references. We evaluate ForestQuery on multiple public forest point cloud benchmarks and a self-collected annotated real-world dataset. Extensive experiments demonstrate consistent improvements in both individual-tree segmentation and semantic segmentation across diverse forest scenes. Code and data are publicly available at https://zhan994.github.io/ForestQuery
cs.AI / 98 / 2610.03445
Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama
cs.CV · cs.AI
Abstract
A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image's representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p<0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder's prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.
cs.AI / 99 / 2610.03467
Preserving Anatomical Continuity: Three-Stage Pipeline for Colon Segmentation in 3D Abdominal CT Scans
Deshan Kalupahana, Sonit Singh, Praveen Ravindran, Arcot Sowmya
cs.CV · cs.AI
Abstract
Accurate colon segmentation from CT images is essential for colorectal disease analysis, yet deep learning based methods often produce disconnected predictions due to complex anatomy. This study introduces a three-stage, topology-preserving segmentation pipeline to address this issue. The first stage performs initial deep learning-based segmentation, followed by centreline bridging to reconnect disjoint regions and a reconstruction stage to refine continuity. Evaluations on TotalSegmentator and RAOS datasets using overlap, distance and topology-based metrics demonstrate improved structural consistency while maintaining segmentation accuracy. The proposed method enhances topological integrity, enabling more reliable colon segmentation for clinical and research applications.
cs.AI / 100 / 2610.03510
Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation
Ziyi Wang, Junchi Yao, Heqian Qiu, Wenbo Shi, Chengjiu Wang, Jinyang He, Binkai Hong, Hongliang Li
cs.CV · cs.AI
Abstract
Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual content. To address these problems, we present Weave Forcing, a training-free framework for compositional memory reuse in interactive long video generation. First, we use an LLM for semantic slot routing to decompose user prompts into character and background descriptions and explicitly select suitable historical references for each component. To isolate the required content, masked memory weaving uses contrasting attention maps conditioned on semantic slots to construct refined semantic masks, selectively exposing relevant tokens from compressed historical KV memories to guide the generation of the current shot. We further introduce coverage adaptive RoPE to adjust temporal offsets and memory retention according to no, partial, or full reference coverage, addressing visual artifacts observed when incomplete historical references are positioned close to the current generation. Extensive experiments demonstrate that Weave Forcing improves cross-shot subject and background consistency while maintaining competitive visual quality and text alignment.
cs.AI / 101 / 2610.03636
LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei, Ben Mildenhall, Georgia Gkioxari, Justin Johnson, Gowthami Somepalli
cs.CV · cs.AI
Abstract
Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: https://ziqi-ma.github.io/logo-website/
cs.AI / 102 / 2610.03649
On-Board Anomaly Detection for Efficient Marine Environmental Monitoring
Thomas Goudemant, Clotilde Szywala, Benjamin Francesconi, Michelle Aubrun, Yves Bobichon, Marjorie Bellizzi, Adrien Girard
cs.CV · cs.AI · cs.LG
Abstract
Marine ecosystems are impacted by various threats such as oil spills, algal blooms, and sediment floods, which disrupt habitats, wildlife, and human activities. Advances in satellite imagery and Artificial Intelligence (AI) have enhanced our capabilities for early detection and mitigation of such hazards. In this paper, we propose a marine event detection pipeline for Earth observation satellites equipped with multi- or hyperspectral sensors. Our approach includes a self-supervised neural network encoder that compresses satellite images into a reduced latent space, enabling efficient onboard processing. A machine learning anomaly detection model identifies deviations from normal sea patterns to detect environmental anomalies. We compare its performance against traditional algorithms such as Isolation Forest, One-Class Support Vector Machine and Local Outlier Factors. Our lightweight, resource-efficient pipeline is optimized for deployment on satellites with limited computational resources, ranging from embedded CPUs to AI hardware accelerators. By prioritizing the transmission of critical information, our solution enhances system responsiveness and optimizes satellite communication bandwidth. Demonstrated through current integration across multiple missions, including European Space Agency's (ESA) Phisat-2 mission and Microsoft/Thales Alenia Space IMAGIN-e mission, our pipeline aims to improve marine environmental monitoring by providing timely alerts and efficient data reduction.
cs.AI / 103 / 2610.03715
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu
cs.CV · cs.AI · cs.GR
Abstract
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench
cs.AI / 104 / 2610.03717
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg
cs.CV · cs.AI · cs.RO
Abstract
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. https://snap-nvs.github.io
cs.AI / 105 / 2610.02370
Network-in-the-Loop at Scale: GPU-Batched 5G Simulation for Massively Parallel Robot Learning
Zifan Zhang, Mingzhe Han, Kannan Athreya, Yuchen Liu
cs.NI · cs.AI · cs.DC · cs.RO
Abstract
Massively parallel GPU simulators train multi-robot policies in thousands of environments, and many fleets use private Fifth-Generation (5G) networks, where each robot's delay depends on its teammates' traffic. Network-in-the-loop training places a simulated 5G network inside this loop. However, GPU robot simulators reduce the network to an independent delay per message, while packet-level simulators run one scenario per CPU process and cannot keep pace with thousands of parallel environments. To bridge this gap, we present Isaac-Net, a GPU-batched 5G New Radio (NR) module that advances the uplink of thousands of environments in lockstep with Isaac Lab physics. Isaac-Net simulates every slot, the 0.5~ms interval in which the base station decides which robots transmit, for all environments at once. Extensive experiments confirm that its NR engine reproduces the median delay of ns-3 5G-LENA across loads, with a median delay 5--10\% low on an unseen carrier and 9\% high at 32 robots per environment in closed loop. The engine also reproduces the Age of Information (AoI), the age of each robot's newest delivered report, while an independent delay per message leaves the AoI tail about three times too light. In a configuration validated against 5G-LENA, Isaac-Net keeps the network in the loop for about one million robots on one GPU at 83\% of the Isaac Lab rate without the network, measured under a random policy. Isaac-Net is open source at https://github.com/ZzZTripleZzZ/isaac-net
cs.AI / 106 / 2610.02376
Coco: An Agentic Copilot for the Hardware--Software Co-Design Lifecycle
Samuel Kushnir, Kavya Sreedhar, Yeshwanth Reddy Pogula, Amir Yazdanbakhsh, Narges Shahidi, Ming Liu, Varun Gohil, Ravi Iyer, Parthasarathy Ranganathan, Christina Delimitrou, Suvinay Subramanian
cs.PL · cs.AI
Abstract
Co-designing ML models and the accelerators that run them is an unusual reasoning task: architects must draw confident, high-stakes conclusions about systems that do not yet exist, and the pace of both model evolution and hardware cadence means the analysis burden grows every quarter. The evidence behind each decision--hundreds of gigabytes of fresh simulation sweeps over novel design points--is by construction absent from any LLM's pretraining corpus, and there is no external literature to retrieve; naive "chat-with-your-data" approaches hallucinate exactly where correctness matters most. We present Coco (Copilot for Codesign), an agentic platform deployed with TPU architects that accelerates the co-design lifecycle of setting up experiments, sweeping simulators, and deriving insights. Coco is built as four layers: (i) a datastore that automatically registers every simulation sweep into a normalized relational schema, so agents ground every number in a SQL query rather than scraping heterogeneous files; (ii) a library of tools with typed APIs that agents compose without human orchestration; (iii) agents that encode recurring analysis workflows--most notably iso-execution analysis, which compares systems at matched execution configurations, including swept-but-dominated points off the Pareto frontier; and (iv) a platform UX whose navigation state doubles as agent context. We report early deployment experience toward a reduction in time-to-simulation and time-to-insight, and argue that co-design is a distinct agentic domain: its data must be retrieved rather than memorized, its workflows are recurring but context-dependent, and expert adoption hinges on UX that balances IDE-style control with interactive exploration.
cs.AI / 107 / 2610.03234
WAMpy: Efficient Synthesis of Prolog Programs in Python
Dominik Magiera, Lukas Röhrig, Frank Jäkel
cs.PL · cs.AI
Abstract
We present WAMpy, a Python framework optimized for synthesizing Prolog programs. Unlike general-purpose Prolog systems, WAMpy targets workloads that repeatedly generate and evaluate small candidate programs. WAMpy compiles Prolog clauses into NumPy array-based WAM instructions and supports partial recompilation of hypotheses against fixed background knowledge. Performance-critical routines are accelerated using Numba just-in-time (JIT) compilation. In a benchmark of repeated compilation-and-evaluation workloads, WAMpy improves end-to-end performance compared with SWI-Prolog accessed from Python using Janus.
cs.AI / 108 / 2610.02832
FastOPD: On-Policy Distillation for Lightweight VLA Deployment
Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye
cs.RO · cs.AI · cs.CV · cs.LG
Abstract
Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of $π_{0.5}$ with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.
cs.AI / 109 / 2610.03476
MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation
Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi
cs.RO · cs.AI
Abstract
Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms $π_{0.5}$-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.
cs.AI / 110 / 2610.03498
Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi
cs.RO · cs.AI
Abstract
Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-10. Conditional intervention improves success rate under intermittent attacks, whereas continuously applying the same intervention substantially degrades policy performance. These results show that attack-related internal representations can provide useful targets for VLA adversarial defense and that controlling when to intervene is important for limiting disruption to nominal policy behavior.
cs.AI / 111 / 2610.03710
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Kanazawa
cs.RO · cs.AI
Abstract
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)
cs.AI / 112 / 2610.03656
Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles
Junyoung Koh, Hao-Wen Dong
cs.SD · cs.AI
Abstract
Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.
cs.AI / 113 / 2610.03321
Information Limits of Low-Rank Approximation Certification
Kang Liu, Bohao Qu
math.OC · cs.AI · cs.IT
Abstract
Low-rank approximation can require additional matrix--vector products to verify that its error meets a prescribed tolerance. We characterize this certification cost for both relative matrix error and mean-square output error. For a single approximation matrix candidate, we determine the exact dimension-uniform minimax query constant as the allowed failure probability vanishes. Our main result concerns reusing validation responses as the approximation space expands. For a candidate family constructed independently of validation, one batch supports an entire nested path without increasing the query budget with the number of checks. Across \(W\) paths, a concentration bound exploiting shared residual energy yields a \(\sqrt{\log(W+1)}\) dependence. A matching lower bound establishes its optimality for fixed interior error targets and sufficiently small separation gaps. Finally, we compare two uniformly valid certificates on the same dispersed-spectrum family. Optimizing the validation budget within each rule family yields costs of orders \(N^{1/3}\) and \(N^{2/3}\) for validation and construction beyond the true target. Code is available at https://anonymous.4open.science/r/Low-rank-approximation-1275/
cs.AI / 114 / 2610.03106
S2S-JEPA: Predicting the Predictable at Subseasonal-to-Seasonal Timescales
Chenyu Dong, Gianmarco Mengaldo
physics.ao-ph · cs.AI · cs.LG
Abstract
The subseasonal-to-seasonal (S2S) timescale, roughly from two weeks to two months ahead, is a critical forecast window for sectors such as agriculture, energy, and water management. Yet, it is widely known as the `predictability desert'. Recent AI weather models excel up to two weeks ahead but deteriorate beyond, largely because they are trained to predict fine-scale details that are neither predictable nor essential at S2S timescales. We argue that a more physically grounded objective is to forecast only the slowly varying components that remain predictable. Computer vision reached the same conclusion with the Joint-Embedding Predictive Architecture (JEPA), which predicts in latent space, discarding unpredictable details. In this work, we introduce S2S-JEPA, which brings the JEPA paradigm to S2S forecasting. It is tailored to this task through design elements from state-of-the-art AI weather models. S2S-JEPA achieves comparable skill to the gold-standard ECMWF physics-based ensemble and surpasses it on multiple metrics at weeks 5 to 6.
cs.AI / 115 / 2610.02651
Equivariant Flow Matching for Electron Density Prediction
Chenxing Liang, Chengdong Wang, Yuchao Lin, Xiaofeng Qian, Shuiwang Ji
physics.chem-ph · cs.AI
Abstract
Machine learning surrogates for density functional theory (DFT) have been increasingly used to reduce the cost of first-principles calculations. In this arena, predicting real-space electron densities offers a scalable and transferable initialization for self-consistent field (SCF) procedures. However, current methods face a clear dilemma. That is, grid-based architectures incur a high computational cost, while basis-set methods fail to capture the structural correlations inherent in the coefficient space. Here, we develop OrbFlow, an $\mathrm{SE}(3)$-equivariant generative model that predicts Gaussian-type orbital (GTO) coefficients via flow matching. OrbFlow retains the efficiency of a compact atom-centered basis while replacing pointwise regression with a learned probability path over the full coefficient space. It is trained through a two-phase trajectory curriculum that mitigates discretization drift during numerical integration. OrbFlow achieves state-of-the-art accuracy on QM9, reducing density error by 13.6% relative to the previous best model, and reduces error by 51% to 63% on every molecule of the MD benchmark relative to the strongest prior method sharing its basis. The predicted density also cuts SCF iterations by up to 68% with zero-shot transfer to unseen exchange-correlation functionals and recovers dipole and quadrupole moments to within a few percent of DFT references without any SCF calculation.
cs.AI / 116 / 2610.02515
IGNITE Tokamak World Model Architecture
Peter Steiner, Azarakhsh Jalalvand, Nathaniel Chen, Kouroche Bouchiat, Ricardo Shousha, SangKyeun Kim, Egemen Kolemen
physics.plasm-ph · cs.AI · cs.LG
Abstract
We introduce IGNITE, a generative world foundation model for fusion plasma behavior simulation trained in a self-supervised manner from over a decade of unlabeled experimental data at the DIII-D National Fusion Facility. The core of IGNITE is a dynamics model that can simulate DIII-D discharges from a given set of actuator trajectories. These trajectories can be supplied or generated on-the-fly from a textual prompt or from desired experimental outcomes. The model architecture consists of several spatio-temporal tokenizers that embed the different input modalities, including time-series like spatio-temporal measurement data, image sequences, and high-resolution spectrograms, each of which collected at vastly different time scales. The backbone is composed of an auto-regressive dynamics model that has the capacity to predict entire DIII-D discharges given initial latent plasma states and actuator trajectories over a theoretical infinite horizon. IGNITE paves the way towards efficient AI-driven experimental planning and world modeling for nuclear fusion.
cs.AI / 117 / 2610.03160
Multimodal reasoning for broadly neutralizing antibody discovery from label-free human B cell repertoires across virus families
Hantao Lou, Jianqing Zheng, Can Yue, Meihan Zhang, Yuanchao Bao, Yu Chen, Mengting Huang, Yupeng Yang, Qianyu Pan, Nana Fu, Yansong Shi, Hongli Li, Yangyang Chai, Ruyi Chen, Wansheng Li, Zhu Liang, Rongmei Yao, Yuanhan Mo, Lei Wang, Chunmei Wang, Yun Quan, Qiong Zhang, Xiangxi Wang, Xuetao Cao
q-bio.QM · cs.AI · cs.CE · q-bio.CB
Abstract
Discovering broadly neutralizing antibodies (bnAbs) from human natural immune repertoires remains a fundamental challenge in immunology, hindered by: the extreme rarity of bnAb, incomplete understanding of their cellular origins across pathogens, and the inability of existing computational tools to generalize across emerging viral threats. Here we present ImmuneAgent, a closed-loop AI system that integrates multimodal reasoning with continual meta-learning and wet-lab feedback to overcome these barriers. Applied to screen the natural BCR repertoires from vaccinated or infected cohorts, the system achieves a ~55% neutralization antibody discovery rate (60 of 110 cloned candidates) and a ~11% bnAb yield (12 of 110), substantially outperforming a state-of-the-art sequence-based neutralization predictor or cofolding models evaluated at the same cloning budget. Five ImmuneAgent-discovered antibodies conferred 100% in vivo protection against lethal influenza challenge, comparable to the clinical-stage therapeutic MEDI8852. The system recovered the cellular and structural determinants of bnAb activity and identified FCRL5+CD27+ atypical memory B cells as a conserved bnAb reservoir and hydrophobic interface enrichment as a cross-viral structural signature, which generalized to unseen antigens, discovering human metapneumovirus (hMPV) cross-neutralizing and human papillomavirus (HPV)-neutralizing antibodies without antigen-specific sorting. These results validate that ImmuneAgent is a generalizable framework for rapid therapeutic antibody discovery against emerging viral threats.
cs.AI / 118 / 2610.03598
When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game
Siu Tung Wong, Carlo Campajola
q-fin.TR · cs.AI
Abstract
Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
cs.AI / 119 / 2610.02437
Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks
Haodong Liang, Yanhao Jin, Krishnakumar Balasubramanian, Lifeng Lai
stat.ML · cs.AI · cs.LG
Abstract
Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class express the same semantic answer in different styles. The tasks share an underlying semantic rule but differ in their prompt distributions and teachers' stylistic preferences. Using a tractable linear-softmax policy, we derive an exact decomposition of the updates into semantic and style components. We show that, at a common policy and prompt, SFT and RFT have parallel semantic updates but differ in their style dynamics. Starting from a policy with no within-class style preference, RFT with exact policy gradients preserves this symmetry, whereas SFT with a nonuniform teacher develops off-axis style drift along a nonzero task mean under population updates. We use this drift to establish a separation under explicit conditions: for population updates from a common perfectly fitted checkpoint, SFT forgetting admits a strictly positive lower bound over a finite training interval, while RFT retains zero semantic error. Simulations over task sequences support these theoretical predictions.
cs.AI / 120 / 2610.03483
AREX: Affine-Residual Exponential Integrator for Few-Step Sampling in Flow Matching
Shizheng Lin, Soon Hoe Lim, N. Benjamin Erichson
stat.ML · cs.AI · cs.LG
Abstract
We introduce AREX, a training-free sampler for pretrained flow matching models that uses the target mean and covariance to capture an analytically tractable part of the sampling dynamics. We show that the velocity field of the moment-matched Gaussian target is the $L^2$-optimal affine approximation to the marginal velocity field. This motivates decomposition of the learned dynamics into an affine component over the whole sampling path, determined by the first two target moments, and a neural residual term. AREX keeps the affine component and integrates it using an explicit matrix-valued propagator. In turn, we only require to integrate over the residual term. This differs from scalar exponential integrators, which analytically handle only isotropic linear dynamics. Across image and text-to-image generation tasks, AREX consistently improves sample fidelity in the few-step sampling regime without retraining the underlying model.
机器学习 (cs.LG)
180
cs.LG / 1 / 2610.03389
From Patching to Pruning Visual Computation in Vision Language Models
Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Shaoyi Huang, Pu Zhao, Yanzhi Wang
cs.CV · cs.LG
Abstract
Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.
cs.LG / 2 / 2610.03147
TSGuard: A Real-Time Framework for Detecting and Imputing Missing Data in Streaming Time Series
Imane Hocine, Asma Abboura, Soror Sahri, Abhijith Senthilkumar, Yacine Hakimi, Grégoire Danoy
cs.DB · cs.HC · cs.IR · cs.LG
Abstract
Streaming sensor applications routinely suffer from delayed or missing observations caused by faults, communication losses, or environmental interference. Although recent imputation methods exploit temporal and spatial dependencies effectively, most either assume offline access to future observations or prioritize throughput without enforcing domain plausibility. We present TSGuard, a real-time demonstration system for monitoring, validating, and imputing missing values in streaming time series. TSGuard combines a lightweight graph-aware temporal imputation model with constraint-aware validation, fallback estimation, and operator-facing explanations. Rather than treating imputation as an isolated prediction task, TSGuard integrates it into a broader data-quality loop: detect problematic observations, impute missing values, validate estimated against physical and spatial constraints, and either retain the original value as a plausible anomaly or replace it when it violates domain constraints. Using environmental sensing as a motivating setting, the demo enables users to inspect delayed sensors, compare imputers, define constraints, and validate flagged values in real time. The combination of lightweight online spatiotemporal imputation, domain-aware validation, and explicit retain-or-replace decisions is our central contribution, while interactive explanations make these decisions inspectable and actionable. for operators.
cs.LG / 3 / 2610.03621
Normal-Form Correlation in Markov Games
Ioannis Anagnostides, Constantinos Daskalakis, Gabriele Farina, Noah Golowich, Tuomas Sandholm, Brian Hu Zhang
cs.GT · cs.CC · cs.DS · cs.LG
Abstract
There has been a surge of recent work on correlated equilibrium concepts in Markov games. However, existing results focus on concepts weaker than normal-form correlated equilibria (NFCEs), leaving open the more challenging question of computing such equilibria, which goes back to the seminal work of Papadimitriou and Roughgarden (JACM'08). Here, we establish the first efficient algorithm for NFCEs in finite-horizon Markov games with a fixed number of players $n$. In particular, with $S$ states, horizon $H$, and at most $A$ actions per player, it computes an $ε$-NFCE in time $S(AH/ε)^{O(n)}$. This is the first algorithm polynomial in $1/ε$ and the description of the game for NFCEs in an interesting class of problems beyond the normal-form setting. Moreover, under the usual assumption that recommendations are independent across states, we show PPAD-completeness---that is, computational equivalence to Nash equilibria---either in many-player games or when the precision is exponentially small. The key idea behind our approach is to run backward induction on a sequence of auxiliary stage games, but with the twist that in each step we compute a constant-expectation correlated equilibrium. This is a natural refinement of correlated equilibrium in which the conditional expected payoff from obeying is independent of the recommendation. In fact, our reduction goes both ways, establishing an equivalence between constant-expectation CEs and NFCEs in Markov games. For a fixed number of players, we observe that a constant-expectation CE can be computed approximately by combining linear programming with suitable discretization. In contrast, it is PPAD-hard in i) polymatrix (many-player) games at constant precision, and ii) two-player games at exponentially small precision. The latter result follows from an unexpected connection to rank-2 two-player games.
cs.LG / 4 / 2610.02711
Characterizing the Performance Gap in Human Activity Recognition for Older Adults
Hossein Khayami, Sungjin Hwang, Eshed Ohn-Bar, David E. Conroy, Amanda Lazar, Eun Kyoung Choe, Hernisa Kacorri
cs.HC · cs.LG
Abstract
Human activity recognition (HAR) from wrist-worn accelerometers is increasingly used for health and behavioral tracking. Yet, most wearable HAR models are developed and evaluated on datasets dominated by younger adults, leaving it unclear whether benchmark progress generalizes across age groups. In this work, we leverage MyMove, our carefully annotated, free-living older-adult HAR dataset (mean age 71), to evaluate deep-learning architectures and training regimes under both leave-one-subject-out and cross-dataset transfer. We find that improvements on younger-adult benchmarks fail to transfer equally to data collected from older adults, resulting in a persistent and often widening performance gap. However, richer representations, particularly frozen self-supervised features pretrained on the age-diverse UK Biobank dataset, substantially improve performance on data from older adults and consistently narrow the performance gap, at modest cost to younger-adult performance, though disparities remain. These findings suggest that benchmark gains and architectural scaling alone provide an incomplete picture of progress in wearable HAR, and broader advances may require representations that better capture population diversity, alongside personalized adaptation to individual movement patterns and routines.
cs.LG / 5 / 2610.02749
Learning Query Encoders Can Be Hard Even When Vector Retrieval Is Geometrically Easy
Anders Wikum, Nina Mishra, Amin Saberi, Tal Wagner
cs.IR · cs.LG
Abstract
Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top-$k$ answer sets of $n$ documents. We study a different notion of geometric capacity--the maximum recall achievable for a frozen document index--and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support. Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline $k/n$. Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval.
cs.LG / 6 / 2610.02334
Joint Movement and Compression Ratio Design for Mobile Embodied AI Networks (MEAN)
Yahao Ding, Jiaxiang Wang, Zhouxiang Zhao, Zhaohui Yang, Mingzhe Chen, Mohammad Shikh-Bahaei
cs.IT · cs.LG
Abstract
Mobile embodied AI networks (MEAN) enable embodied agents to perceive, reason, communicate, and act in wireless environments. In such networks, agent mobility can improve channel conditions, while semantic compression can reduce transmission payloads. However, movement consumes energy, and stronger compression incurs additional computational cost. This paper studies joint movement, semantic compression, and transmit power design for an uplink MEAN system. We formulate a max-min energy efficiency (EE) problem by jointly optimizing transmit power, movement distance, and semantic compression ratio under controllable power constraints. The problem is non-convex due to the coupled signal-to-interference-plus-noise ratio (SINR), mobility-dependent channel gains, and fractional EE objective. To solve it, we propose an alternating optimization (AO)-Dinkelbach algorithm, where the fractional objective is handled by the Dinkelbach transformation, transmit power is updated via successive convex approximation (SCA), and movement distance is updated by coordinate-wise grid search. Simulation results show that the proposed scheme outperforms no-mobility and no-compression baselines, demonstrating the benefit of jointly exploiting mobility control, semantic compression, and power allocation in MEAN.
cs.LG / 7 / 2610.02344
Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures
José Lucas De Melo Costa, Seong Woo Ahn, Fabrice Popineau, Arpad Rimmel, Bich-Liên Doan
cs.LG
Abstract
Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ($γ$) and a decay effect ($σ$). Under approximate spectral decoupling, a per-mode stability ratio $μ_i = γ_i / σ_i$ factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting $μ$. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available at https://github.com/jose-melo/drive-vs-decay.
cs.LG / 8 / 2610.02345
Mitigating Convergence Collapse in Fixed-Target Anomaly Detectors via Kernel-Anchored Locality Regularization
José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel, Bich-Liên Doan
cs.LG
Abstract
A family of tabular anomaly detectors trains a neural map toward a fixed target under squared-error loss and scores anomalies by the test-time residual; contraction matching, one-step rectified flow, and reconstruction autoencoders all fit this template. We characterize a convergence collapse: better optimization makes the detector worse. At convergence, the learned map tracks the target even off-distribution, so the residual signal vanishes on anomalies as well as on normal data. These detectors therefore rely on implicit non-convergence (early stopping, capacity caps) to retain signal. We argue this is structural: effective anomaly detection requires a locality constraint that blocks unconstrained extrapolation. Classical detectors (kNN, KDE, isolation forests, LOF) enforce locality explicitly; fixed-target neural detectors do not. We formalize the connection by showing that the kernel-regression analog of a fixed-target detector is a finite-bandwidth Nadaraya-Watson smoother, which we call Kernel Contraction Matching (KCM). KCM is closed-form, training-free, and CPU-efficient, yet matches established neural baselines on ADBench. Building on this bridge, we introduce the Kernel-Anchored Regularizer (KAR), which penalizes deviation of the neural prediction from a kernel-weighted average of training targets. Across collapse-prone ADBench datasets and three backbones, KAR mitigates collapse and improves AUROC under prolonged training.
cs.LG / 9 / 2610.02347
From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data
Mohamed Bouadi, Nassim Bouarour, Shivam Dubey, Aditya Tanna, Vinay Kumar Sankarapu
cs.LG
Abstract
Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002$\pm$0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution
cs.LG / 10 / 2610.02359
Lexicographic Multi-Objective On-Policy Distillation
Doseok Jang, Jon Ander Campos, Youran Qi
cs.LG · cs.AI · cs.CL
Abstract
Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.
cs.LG / 11 / 2610.02363
ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time
Pranay Kothari
cs.LG · cs.DB
Abstract
Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, duplicated, out-of-order and retried records) and requires its final state to equal a batch recomputation of the complete log. Because the oracle recomputes rather than classifies, a wrong table and a crash are distinct verdicts: a crash is visible to monitoring a team already runs, and a wrong table is not. On 40 tasks we built, our reimplementation of single-execution grading certifies 86-100% of the pipelines eleven models produce; re-executing the same artifacts finds 7.0-79.2% of the certified ones silently wrong. The gap is not produced by the repair loop: within the same model and task, pipelines repaired against the snapshot test fail replay about as often as those that passed it first time. In every model, idempotency hazards fail more often than ordering hazards. Separating a wrong answer from a crash also changes how interventions read: a hazard warning cuts one model's silent failure from 48.2% to 10.5% while raising its crash rate from 9.0% to 37.0%, so all-in failure moves only from 51.0% to 44.0%. All eleven arms were independently re-run, and rates moved by at most 5.9 points.
cs.LG / 12 / 2610.02366
Co-design Gym: A Unified Benchmark for Embodiment-Policy Co-optimization
Aviraj Newatia, Yordan Tsvetkov, Leonard Pleiss, Andrew Spielberg, Rika Antonova
cs.LG · cs.RO
Abstract
Finding an optimal behaviour policy within a given environment is a widely studied problem in domains as diverse as games, robotics, energy infrastructure, communication networks, and multi-agent systems. Numerous benchmarks have been developed to support such research, but the vast majority assume that the agent's embodiment (design) is fixed, focusing instead on policy learning alone. Lifting this assumption gives rise to a broader class of problems in which optimizing embodiment and policy separately is highly suboptimal. An agent's embodiment strongly shapes which control policies can be discovered, while the optimal embodiment is in turn defined by the policies it admits. To help the research community study this class of problems explicitly and systematically, we introduce Co-Design Gym - a suite of benchmark environments for jointly optimizing embodiment and policy. Our environments span domains such as robotic manipulation and locomotion, multi-robot cooperation, deformable and soft dynamics, video games, electricity grids, wireless networks, F1 racing, multi-agent warehouses, and optimal control, offering 20 environment families (domains), with over 85 distinct co-design presets in total. We further contribute a systematic evaluation of representative co-design algorithms, characterizing the current state of the art. Together, these contributions lay the groundwork for cumulative, comparable progress in co-design.
cs.LG / 13 / 2610.02381
Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li
cs.LG
Abstract
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.
cs.LG / 14 / 2610.02383
The Surprising Effectiveness of Shared Memory in Looped Transformers
Giovanni Monea, Keshav Ramji, Yousef El-Kurdi, Luis A. Lastras, Yoav Artzi, Nathan Godey, Ramón Fernandez Astudillo
cs.LG · cs.AI
Abstract
Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters. Each recursion, however, writes its own key-value cache, so memory still grows with compute. Inference-time techniques can shrink this cache at a cost in quality. We pretrain looped language models to share memory: only the first recursion writes a cache, and later recursions read it while keeping a short window of their own. Surprisingly, we find that sharing memory does not cost quality and instead improves it. At 150M-1B parameters, our Looped Prediction Transformer (LPT) and its hybrid variant set a new quality-memory frontier for looped models: with five recursions, the hybrid lowers validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer while using 76-79% less context memory. Through an extensive analysis, we investigate why memory sharing helps. Shared and local memory develop different representations, and later recursions attend mostly to the shared memory, which also acts as a gradient highway to the first recursion.
cs.LG / 15 / 2610.02397
Validated Data Onboarding for AI Demand Forecasting on U.S. Building Meter Data: Design, Controlled Evaluation, and a Corrected Negative Result
Yixuan Liang
cs.LG
Abstract
Electric utilities and grid operators increasingly rely on machine-learning models to forecast next-day demand, and those models learn from meter data that is routinely defective: readings go missing, sensors freeze, buildings read zero for hours, and units change by a factor of 100. This report presents a data-onboarding pipeline that detects and repairs such defects before a model is trained, using only information available at forecast time, and a controlled experiment that measures whether the pipeline protects a 24-hour-ahead forecast. On hourly electricity data for twelve U.S. buildings from the public Building Data Genome 2 dataset (210,528 rows, 2016-2017), seeded, hash-logged defects touching 0.10% of the training period raised the error of a gradient-boosting forecaster by 86%; after detection and past-only repair the error returned to the clean-data level (mean absolute scaled error 0.760 clean, 1.415 corrupted, 0.729 repaired) while 93% of training targets were retained. At a defect prevalence calibrated to published field studies (1.6% of training rows) the unprotected forecaster's error reached 4.4 times that of a seasonal-naive rule, and the repaired forecaster again matched the clean baseline. The same pattern held for ridge regression and a random forest and across horizons of 1 to 24 hours. A first version of the pipeline over-cleaned natural data and made forecasts 25% worse; that result is retained, its cause is traced in the published artifacts, and the per-building calibration that corrects it is documented as a dated amendment. Every number is reproducible from pinned public inputs with SHA-256 verification, 84 automated tests and continuous integration.
cs.LG / 16 / 2610.02399
VisAudit: Evaluating Multimodal Agents for Visual Diagnosis and Repair
Shicheng Liu, Adam Kahirov, Qi Zhang, Zhimin Hu, Song Wang, Junhong Lin, Julian Shun, Yada Zhu
cs.LG
Abstract
Multimodal agents are increasingly used for data visualization tasks but remain limited in autonomous review. Unlike humans, they may fail to recognize when a visualization is incorrect, determine what to change, repair it without disrupting correct content, and verify whether the intervention succeeded. Existing benchmarks largely evaluate predefined individual capabilities such as chart generation, instruction-guided editing, or defect detection, and therefore do not capture this gap in autonomous review. We introduce VisAudit, a benchmark for evaluating visualization diagnosis, repair, and verification. Given a rendered chart and configurable auxiliary evidence, including its source data table, intended text summary, and visualization code, an agent iteratively diagnoses potential defects, modifies and executes visualization code, inspects execution and visual feedback, and determines when no further intervention is needed. VisAudit defines three tracks spanning diagnosed repair, autonomous repair, and open-world verification, and contains 1,900 flawed instances across 21 chart types and 10 flaw categories, together with 300 initially correct charts. We construct the benchmark through controlled perturbations of validated source visualizations, with systematic verification and human-aligned quality control to ensure that injected defects are well-defined and recoverable from the available evidence. Experiments with leading multimodal models reveal a substantial gap from reliable autonomous review: the strongest evaluated model fully recovers only $47.4\%$ of flawed charts in the autonomous-repair setting.
cs.LG / 17 / 2610.02404
Trained Agentic Context Management
Bryce Sandlund
cs.LG · cs.CL
Abstract
We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a diverse synthetic dataset using this harness. With only 8,000 tokens of context, our small model is as strong as GPT-5.4 with 1M tokens of context on the OOLONG-synth benchmark when document length exceeds 40K tokens.
cs.LG / 18 / 2610.02410
Efficient Neural Field Learning via Adaptive Coverage and Focused Sampling
Guang Zhao, Xihaier Luo, Huan-Hsin Tseng, Seungjun Lee, Shinjae Yoo, Yihui Ren, Wei Xu
cs.LG · cs.AI
Abstract
Implicit neural representations (INRs) provide a flexible framework for modeling high-dimensional continuous fields, but their training is often inefficient due to uniform subsampling that ignores spatial heterogeneity. Existing adaptive sampling methods partially address this issue by prioritizing high-error samples, but typically operate at the point level, often leading to redundant sampling in localized regions and insufficient coverage of the domain. We propose ACES (Adaptive Coverage-aware Efficient Sampling), a structured sampling framework that improves training efficiency by decoupling coverage and importance. ACES constructs adaptive spatial partitions to ensure domain coverage and reduce redundancy, and applies region-level importance weighting to prioritize informative regions during training. We provide a theoretical analysis showing that adaptive partitioning reduces gradient variance by increasing within-region homogeneity, and that controlled bias in region-level weighting may improve optimization efficiency relative to standard unbiased estimators. Experiments on scientific field learning tasks demonstrate that ACES achieves faster convergence and lower error than uniform and pointwise adaptive sampling baselines, with the largest gains in fields with highly localized complexity.
cs.LG / 19 / 2610.02413
Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
Elad David, Max Fomin
cs.LG
Abstract
LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to append a short classification instruction after the user's turn and read the probe at that point, to sharpen it: the instruction asks the model to represent the incoming request as a class, concentrating the signal the probe must separate, at negligible serving cost. But does the wording of that suffix matter, and does its benefit hold in the wild, on attack types the probe never saw in training, the regime a deployed monitor faces? We test this with a controlled ladder of post-user suffixes under strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B). On a single-position probe, a classification suffix consistently improves out-of-distribution detection over no suffix (up to ~4 AUC points); yet which suffix matters: prompting the model to classify the input, even into content-free labels, reliably wins; an off-topic or merely-attentive suffix helps little. The gain comes from the classification format, not the named criterion: a content-free suffix matches the real malicious/benign one, with the criterion adding precision only at strict thresholds. This is not an artifact of the single-position read: the benefit carries to the multi-position pooling probes used in production (attention, multi-max, MLP), though the best-performing suffix there is readout-dependent. Served through a KV-cache fork, it is a cheap drop-in for any activation-probe monitor, though not an automatic win: which suffix helps, and by how much, depends on the model and the readout.
cs.LG / 20 / 2610.02417
The AI Theorist reveals excitonic structure in $α$-RuCl$_3$
Hongjian Zhou, Xianfan Nie, Sean Wu, Tarun Patel, Jinge Wu, Andrew Liu, Adam Wei Tsen, David A. Clifton
cs.LG · cond-mat.mtrl-sci
Abstract
Advances in experimental instrumentation and automation generate increasingly rich datasets, but turning experimental observations into microscopic understanding remains a bottleneck in scientific discovery. To accelerate this process, we introduce AI Theorist, a system of artificial intelligence (AI) agents for autonomous discovery of physical models through hypothesis generation, first-principles calculations and evidence-driven refinement. We apply the framework to $α$-RuCl$_3$, a leading candidate material for realizing a Kitaev quantum spin liquid, to investigate its electronic structure through optical spectra. AI Theorist develops a new interpretation of the optical and photocurrent observations, identifying distinct excitonic states with contrasting optical selection rules and real-space distributions. To our knowledge, this is the first demonstration of an AI system autonomously developing a physical model to explain previously unpublished experimental observations in a quantum material, utilizing first-principles electronic-structure and many-body calculations. Our results establish a route to autonomous theoretical discovery in materials science, in which AI agents use first-principles calculations to turn experimental observations into physical models and testable predictions.
cs.LG / 21 / 2610.02427
Geometry-Aware Time Reparameterization for Flow-Map Distillation
Félix Dedek, Makoto Yamada
cs.LG · cs.AI
Abstract
Flow-map distillation enables one- and few-step generation by learning finite-time transitions of a pretrained generative ODE. We investigate whether changing the teacher's time parameterization can make these transitions easier to learn. Motivated by the hypothesis that trajectory segments with large normal acceleration are harder to distill, we propose a geometry-aware time reparameterization that allocates more student time to these regions while preserving the teacher's geometric paths and terminal distribution. We derive a shared clock that equalizes a population normal-acceleration statistic under suitable assumptions, and construct a practical approximation from robust, regularized estimates across teacher trajectories. We incorporate this clock into Lagrangian flow-map distillation, using the transformed time coordinate to condition the student. The clock is estimated once before distillation and requires neither teacher retraining nor additional student parameters or inference-time network evaluations. Experiments on synthetic data, CIFAR-10, and CelebA-64 show improved sample quality over identity-time distillation at matched inference budgets, including improvements in one- and two-step image generation. The gains in one-step generation, where no intermediate sampling times can be adjusted, highlight the benefits of time reparameterization during distillation.
cs.LG / 22 / 2610.02431
CRISP: A Framework for Clause-Reconstructed Interpretable NeuroSymbolic Propositions
Alex Chan, Shafi Muhtasim Chowdhury, Ekin Can Erkuş, Ole-Christoffer Granmo, Alex Yakovlev, Rishad Shafik
cs.LG
Abstract
Deep neural networks achieve high accuracy through layered numerical transformations, yet their decisions remain difficult to audit because decision evidence is encoded in hidden activations rather than explicit rules. This paper introduces CRISP, a framework that reconstructs the last-layer activation vector (LLAV) of binary neural teachers as Tsetlin Machine (TM) clauses. CRISP sign-binarizes the teacher's penultimate pre-logit activations, and assigns one Individual TM (ITM) to each LLAV neuron. Each reconstructed hidden bit is represented by propositional clauses over Booleanized input features, which gives a direct symbolic trace from named input thresholds to a named teacher neuron. CRISP is evaluated on MNIST, KMNIST, FashionMNIST (FMNIST), SVHN, and CIFAR10 using a BinaryConnect convolutional neural network (BCCNN) teacher and a fully binary neural network (BNN) teacher, with an additional study on binary thresholding, thermometer encoding, and quartile binning at multiple bit depths. The results show that LLAV sign-binarization does not reduce teacher-head accuracy in the tested BNN setting, while ITM reconstruction error is the main limiting factor. Quartile one-bit Booleanization gives the strongest reconstruction fidelity on SVHN at 87.52% test fidelity and is competitive on CIFAR10, and the reconstructed LLAV preserves 78.41% teacher-head accuracy on FMNIST. Pooled clause-evidence visualizations show that the learned ITM literals concentrate on the object region in centered benchmarks. CRISP therefore provides a clause-level route for inspecting the final hidden representation of binary neural teachers.
cs.LG / 23 / 2610.02439
A Generative Model of Complex Networks Using Graphons and Neural Inverse Operators
Wooseong Choi, Italo'Ivo Lima Dias Pinto, Chen Sun, Gaurav Gupta, Dong Song, Paul Bogdan
cs.LG
Abstract
Generative graph models are central to understanding and simulating complex networks. However, existing approaches have complementary strengths and limitations. Mechanistic models offer interpretability but rely on instance-specific estimation methods. Deep generative models, on the other hand, offer amortized inference at the cost of interpretability and are largely limited to graph sizes seen during training. Scientific applications motivate a framework that retains the strengths of both paradigms. We bridge them by formulating both the generative model and parameter recovery in function space. A multifractal step graphon extends standard step graphons with a recursive construction that compactly parameterizes complex networks. This formulation admits a neural inverse operator to recover its parameters, enabling inference on unseen graph sizes. We evaluate our model, trained only on synthetic multifractal step graphon realizations, against both paradigms. Against a graph foundation model pretrained on empirical networks, our method achieves the best average performance on three of four metrics in a zero-shot graph-generation benchmark, indicating that the model transfers to real-world graphs. We also apply our method to single-observation networks, a regime largely inaccessible to deep models that require training corpora, where it performs comparably to an instance-specific method that optimizes on each graph. In a multi-subject EEG case study, the inferred parameters track a reversible change in brain state more sensitively than traditional network statistics. Together, these results indicate that mechanistic interpretability and amortized inference can be effectively unified in a generative graph model to enhance our understanding of complex networks.
cs.LG / 24 / 2610.02440
Bandits via Additive Quantized Representations
Ami Tavory, Noam Touitou, Tal Sarig, Frank Cheng, Ido Guy
cs.LG
Abstract
Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles and neural methods capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across multiple levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines using up to 1000 times less memory.
cs.LG / 25 / 2610.02461
A Composable AI-Accelerated Iterative Solver for 3D-IC Thermal Modeling
Yixing Li, Jiahang Zhou, Zhiyu Zeng, Xin Ai
cs.LG
Abstract
Accurate thermal analysis of heterogeneous 2.5D/3D-IC packages is essential yet computationally prohibitive. A single full-package FEM simulation can take hours, while AI-based surrogates treat the entire stack as a monolithic prediction target and must be retrained whenever the die count or topology changes. To address this limitation, this work proposes Domain-Decomposed AI-Accelerated Iterative Solver for Thermal Analysis (DAIST), a composable thermal solver that decomposes the global package simulation into block-level subdomain problems, replaces subdomain solvers with neural operators, and couples them through iterative exchanges of interfacial temperature and heat flux. This local-to-global architecture eliminates the topology lock-in of monolithic models: block-level neural operators can be directly reused in unseen package assemblies without retraining. The iterative coupling strategy further provides a controllable accuracy-runtime tradeoff, where the iteration budget can be adjusted to trade accuracy for runtime. Evaluated on a multi-chiplet system and an advanced packaging system, DAIST achieves up to $178\times$ speedup over traditional FEM solvers with mean temperature errors of 0.068% and 0.323%, respectively, while demonstrating cross-topology reuse of block-level models across structurally distinct package assemblies.
cs.LG / 26 / 2610.02462
Capability Scaling-Down Laws for LLM Compression
Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong
cs.LG · cs.CL
Abstract
LLM compression reduces inference costs and memory requirements, but selecting a method and configuration remains largely empirical because comparable resource reductions can produce different capability losses. We systematically investigate capability scaling-down laws for LLM compression across pruning, quantization, and distillation. Our framework measures capability loss in mathematics, code generation, and question answering, and relates these measurements to model size, training stage, compression settings, data availability, and training exposure. We develop simple predictive relations and evaluate their accuracy, measurement efficiency, and generalization to unseen configurations and model states. Sharing the density response across pruning levels halves the configuration measurements needed to fit a pruning predictor: on new Pythia states, on pre-registered OLMo-2 test states and under Wanda pruning, the compact relation matches a regression fitted with all measurements on math and code to within 0.020 nats per token, with coefficients refitted for each setting. Controlled distillation experiments show that the cost of heavy data reuse recurs across question-answering distributions, while the net benefit depends on the evaluation distribution. We further evaluate the decision value of these predictions by comparing numerical selection with configuration medians and fixed method priorities. Independent evaluations across two model families show that selection captures most of the available cross-method benefit for question answering within the tested candidate sets, where a fixed method priority attains the same regret, with smaller opportunities for mathematics and code. These results clarify the predictive scope of capability scaling-down laws and their use in compression method selection. Our code is publicly available at: https://github.com/LabRAI/scaling_down_law.
cs.LG / 27 / 2610.02487
Threshold-Aware Conformal Routing
Shiwei Tan, Huzefa Rangwala, Danielle C. Maddix
cs.LG · cs.CE
Abstract
High-fidelity simulations are essential to scientific and engineering design, but can be expensive to run repeatedly. Learned surrogates offer a faster alternative, yet their higher errors may alter downstream decisions. This accuracy-speed tradeoff creates a need to determine whether a surrogate can be used or the full simulator remains necessary. We study decisions determined by whether a scalar quantity of interest lies above or below a fixed threshold. For each input, we use the surrogate when its conformal interval lies entirely on one side of the threshold and route the input to simulation when the interval intersects it. Standard conformal prediction constructs intervals without reference to the downstream decision threshold: even a narrow interval near the threshold can cross it and trigger simulation, whereas a wider interval farther away can remain entirely on one side and require no simulation. We introduce Threshold-Aware Conformal Routing (TACR), which learns an input-dependent scale using a threshold-aware objective that concentrates interval tightness near the decision boundary. Exact split-conformal calibration on held-out data preserves distribution-free marginal coverage, which also upper-bounds the probability of an incorrect threshold decision that is not routed. Across various scientific and engineering datasets, TACR reduces simulator deferrals by 14-75% relative to standard conformal prediction at the same coverage target. Against a variant without threshold-local weighting but with similar predictor accuracy, TACR further reduces deferrals by 10-24% on four datasets. These results show that optimizing interval allocation for routing can reduce simulator calls without weakening the standard conformal guarantee.
cs.LG / 28 / 2610.02488
Harnessing LLMs as Agents: What Does It Cost?
Zelin Zhao, Xinyu Guo, Jingyuan Zhang, Yuxuan Zhang, Yongxin Chen
cs.LG
Abstract
Language-model agents increasingly rely on harnesses that manage bounded context, persistent memory, tools, verification, and repeated execution, yet existing notions of model capability do not quantify the computational resources these mechanisms consume. We introduce the Language Model Agent Machine (LAM), a resource-bounded abstraction that fixes the underlying semantic model while explicitly charging harness-level resources. We establish four classes of results. Communication: LAM execution is instancewise equivalent to red--blue pebbling under simultaneous call--transfer budgets, transferring classical I/O lower bounds to context--memory traffic. Access: memory interfaces induce asymptotic separations, including a $Θ(n)$ gap between random and non-speculative sequential access on pointer chasing. Recomputation: bit-reversal DAGs require $Θ(n^2/(C+S)+n)$ model calls with context capacity $C$ and persistent-memory capacity $S$, quantifying when stored intermediate state avoids repeated semantic computation. Reliability: we derive tight stage-local sampling bounds, exact imperfect-verification costs, and a Young--Daly-type checkpoint law with a closed-form optimal verification interval. Controlled and held-out experiments on GPT-6 Astra test communication and reliability predictions, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks. Together, these results provide a resource theory for the computational cost of language-model agent harnesses.
cs.LG / 29 / 2610.02497
LiteEMG-FM: An Efficient and Deployable Foundation Model for Robust EMG Sensing
Tianhao Wu, Xu Wu, Amirmohammad Radmehr, Jiawei Yu, Yi Wu, Phuc Nguyen, Jian Liu
cs.LG
Abstract
Electromyography (EMG) signals vary substantially across individuals, body regions, recording sessions, and sensing hardware, limiting the generalization of models for assistive devices and human-computer interaction. Existing time-series foundation models are also computationally expensive for real-time wearable deployment and often fail to capture EMG-specific time-frequency characteristics. We present LiteEMG-FM, an efficient hybrid CNN-Transformer foundation model for practical EMG sensing. Pretrained on 16 diverse upper- and lower-limb EMG datasets, LiteEMG-FM learns representations that generalize across users and datasets. For resource-constrained deployment, we implement a hierarchical wake-up architecture in which a lightweight, always-on 1D-CNN filters rest and non-target activity and activates LiteEMG-FM only for valid gestures. We evaluate full inference offloading, split inference, and full on-device processing, characterizing their trade-offs in latency, power consumption, and memory footprint. Across diverse evaluation settings, LiteEMG-FM outperforms state-of-the-art time-series foundation models and supervised baselines, particularly under zero-calibration cross-participant and data-scarce conditions. These results demonstrate that LiteEMG-FM is an effective, efficient, and deployable foundation model for EMG applications.
cs.LG / 30 / 2610.02505
Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning
Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil
cs.LG · cs.AI · cs.RO
Abstract
Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a simplified simulator. Most existing methods directly optimize biased objectives based on LF data. In contrast, the recently introduced multi-fidelity policy gradient (MFPG) framework uses LF data solely to construct a control variate that reduces variance and improves HF data efficiency without biasing the policy gradient estimator. However, published work on MFPG is limited to REINFORCE on small-scale simulation tasks. We develop MFPG for modern actor-critic learning in GPU-parallel simulation and on a physical robot. Our analysis and experiments show that naive extensions to proximal policy optimization (PPO) can lose cross-fidelity correlation or inflate variance. Our MFPG-PPO addresses these failures by redesigning the sampling, advantage estimation, and control variate construction to preserve cross-fidelity correlation, and by monitoring estimator uncertainty to prevent variance inflation. We also introduce a budget-aware MFPG-PPO to divide a fixed sampling budget among high- and low-fidelity data sources. Across simulated robot locomotion tasks of varying LF-to-HF transfer difficulty and HF data budgets, MFPG-PPO improves upon PPO trained on HF data alone in nearly all settings, and consistently matches the performance of PPO trained with 16x more HF data on the hardest task at the smallest HF budgets. In contrast, most baselines that use LF data perform well only where direct LF-to-HF transfer succeeds. MFPG-PPO enables stable learning on a physical Franka arm using only 4 real-robot episodes per update and no human demonstrations.
cs.LG / 31 / 2610.02511
Post-Training Quantization of Autoregressive Weather Models
Ananyo Bhattacharya, Swastik Bhattacharya, Christiane Jablonowski
cs.LG · astro-ph.EP
Abstract
Advancements in high-resolution numerical weather prediction (NWP) and data assimilation (DA) have shaped the developments in deep learning (DL) architectures emulating atmospheric dynamics. Emulators for weather forecasting exhibit forecast quality comparable to physics based models at forecast horizon scaling from few days to subseasonal time scales. The emulators are driven by hardware-accelerated matrix multiplication in autoregressive inferences, significantly reducing the computation time and resources required for NWP. Optimization of the matrix multiplication processes in GPU architectures provides opportunities to scale towards high-resolution domain, and offers implementation of out of the box solutions. Post-training quantization (PTQ) has been demonstrated across multiple DL architectures to accelerate and increase the number of computations in unit time while consuming less power, enabling applications on edge hardware. In this study, we investigate the effect of PTQ on pre-trained AI emulators for global-scale weather forecasting. We implement PTQ algorithms in Deep Learning Weather Prediction (DLWP) and FourCastNet (FCN) models as a proof of concept for geophysical fluid dynamics applications. We systematically investigate the effect of PTQ on emulator inferences over short-range forecast horizons. Evaluation of PTQ configurations using simulated quantization hints at qualitatively meaningful forecasts over short-time horizons. These results provide a first benchmark of PTQ for autoregressive weather emulators and a basis for quantization-based optimization of DL models for dynamical systems.
cs.LG / 32 / 2610.02516
Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers
Haifeng Wu, Srinivasan Manoharan, Jian Wan, Fangbo Tu, Junhua Zhao, Xin Chen
cs.LG · cs.AI
Abstract
Zero-shot classifiers are useful for routing user requests to specialized LLM tasks, but scoring every request against a large candidate set is expensive: a zero-shot NLI classifier must evaluate one premise-hypothesis pair per label, so cost scales linearly with taxonomy size. We study a student-guided teacher distillation pipeline for a fixed taxonomy of 60 LLM task categories: a compact ModernBERT classifier predicts the full category distribution in one forward pass and retrieves a small top-k candidate set, and a larger DeBERTa-v3 zero-shot NLI classifier reranks only those candidates rather than all 60 labels; the resulting teacher labels iteratively improve the student, which produces sharper candidates for the next round. Unlike generic embedding retrieval or clustering-derived shortlists used in extreme multi-label classification, our candidate generator is trained end-to-end on the target taxonomy and is the same model serving production traffic, distinguishing it from LLM-routing work that routes between candidate models, and from concurrent System-1 encoder-classifier proposals (e.g. TypeSafe AI's Jev and the open-source Laya project) whose training methodology is undocumented or RL-based. Our best student checkpoint reaches 77.5% teacher agreement on a 200-example evaluation set, and preliminary coverage measurements show Coverage@16 of 91-100%, suggesting top-k sets retain most of the teacher's decision-relevant information. We further show truncated top-k teacher scores should not be treated as full 60-class soft targets for KL distillation: zeroing untruncated classes destroys the dark knowledge soft-label distillation depends on, introducing systematic bias rather than a harmless sparse approximation. A complete evaluation, including coverage at multiple k on a held-out set, an embedding-retrieval baseline, and a larger human-reviewed test set, remains in progress.
cs.LG / 33 / 2610.02517
Learning the Latent Structure: A Feature-Centric Approach to Graph Data Augmentation
Yu Song, Zhigang Hua, Yan Xie, Bingheng Li, Jingzhe Liu, Bo Long, Jiliang Tang, Hui Liu
cs.LG
Abstract
Graph-structured data plays a pivotal role in modeling complex relationships. However, real-world graphs are often incomplete due to data collection and observational constraints, severely limiting the effectiveness of modern graph learning pipelines. While existing Graph Data Augmentation (GDA) methods attempt to refine graph structures for improved downstream performance, they are typically label-dependent, computationally expensive, and inherently transductive, limiting their applicability in practical scenarios. In this work, we present a novel feature-centric graph data augmentation framework that bypasses explicit structure modeling by operating directly in the embedding space. Through a self-supervised inverse masking process, our method captures latent ties between observed and complete graphs, enabling recovery of unobserved structural signals through refined node representations. To enhance robustness under noisy and sparse supervision, we introduce a message regularizer and a bootstrap strategy for effective training and generalization. Evaluated on ten graph datasets spanning multiple domains, our approach, SelfAug, consistently outperforms state-of-the-art methods in both accuracy and efficiency across inductive and cold-start settings, highlighting its potential as a scalable and generalizable solution for real-world graph learning scenarios.
cs.LG / 34 / 2610.02520
Instance-Dependent Regret for CMDPs with Step-Wise Constraints
Qian Zuo, Francesco Emanuele Stradi, Leyang Xue, Sattar Vakili
cs.LG · cs.AI
Abstract
We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints. In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning. Exploiting this structure, however, requires learning which actions are safe while controlling constraint violations. We propose Safe Variance-Adaptive Exploration (SVAE), an efficient algorithm that learns candidate safe subgraphs and performs variance-adaptive optimistic planning within them. With high probability, SVAE achieves cumulative regret of order $\widetilde{\mathcal{O}}(\sqrt{SAH\min\{\mathbb{V}_Σ,K\mathrm{Var}^{\star}\}}+S\sqrt{AH^3\min\{K,\mathcal{C}\}}+S^2AH^2)$ over $K$ episodes, where $H$ is the horizon of a single episode, while $S$ and $A$ are the numbers of states and actions, respectively. Here, $\mathrm{Var}^{\star}$ is the maximum return variance among safe policies, $\mathbb{V}_Σ$ is the variance accumulated before the first unsafe action is encountered, and $\mathcal{C}$ captures the statistical complexity of eliminating actions incorrectly considered potentially safe. SVAE additionally attains $\widetilde{\mathcal{O}}(H\sqrt{SAK}+S^2AH^2)$ step-wise constraint violation and a gap-dependent violation bound that is polylogarithmic in $K$. Finally, we establish a lower bound showing that dependence on these instance-specific quantities is unavoidable.
cs.LG / 35 / 2610.02524
BaCP: Backbone Contrastive Pruning for Preserving Representations in Extremely Sparse Neural Networks
Mohammad Haroon Khawaja, Muhammad Haseeb, Mohammad Fatim Shoaib, Muhammad Tahir
cs.LG
Abstract
Unstructured pruning at extreme sparsity often suffers from representational collapse, causing sharp drops in accuracy. To address this, we study Backbone Contrastive Pruning (BaCP), which regularizes the sparse network's embedding space by aligning it with pretrained, fine-tuned, and historical snapshot models. Building on the contrastive decomposition of the CAP framework (Xu et al., 2022), we provide a rigorous matched-budget characterization of this approach across multiple pruning criteria. Evaluated across 90 settings, BaCP improves accuracy substantially in extreme sparsity regimes where standard pruning fails, and is close to baseline where representations remain intact.
cs.LG / 36 / 2610.02528
Autoregressive Differentiable Method for Integer Programming
Ouns El Harzli, Yudong Cao
cs.LG · math.OC
Abstract
We introduce an autoregressive differentiable method to solve 0-1 integer programs. We fix an arbitrary order of the binary variables and we train a transformer to predict the next bit while remaining in the feasible set. Our method is first trained on feasible incumbents provided by any solver, thus allowing us to initialize the transformer in the feasible set. Our procedure then implements a Lagrangian penalty to penalize infeasible solutions, and the transformer is further trained to explore the feasible set using Gumbel-softmax activations on the relaxed objective. We have tested our method on non-convex instances of quadratic knapsack problem and demonstrated consistent improvement upon state-of-the-art open-source solvers for dense problems up to 10,000 binary variables. In particular, we empirically demonstrate a phenomenon akin to a tunneling effect where the effective change of variables from binary variable to the continuous weights of the transformer that the method implements enables crossing barriers in the relaxed objective landscape.
cs.LG / 37 / 2610.02545
Reward Inflation: A Healthy Stimulus for Reinforcement Learning
Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang
cs.LG
Abstract
Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.
cs.LG / 38 / 2610.02554
Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
Dongsu Lee, Haoran Xu, Amy Zhang
cs.LG · cs.RO
Abstract
Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.
cs.LG / 39 / 2610.02559
Neuron merging via inverse-activation regression for post-training compression of sigmoid neural networks
Ao Kuniya, Jun Ohkubo
cs.LG
Abstract
As neural networks continue to grow in scale, model compression is becoming increasingly important for efficient inference under limited computational resources. Structured pruning methods remove neurons or channels that are estimated to be less important, but the removed units may still contain useful information. From the viewpoint of coarse-graining a trained network, it is valuable to ask which information should be retained when multiple neuronal degrees of freedom are consolidated. In this paper, we discuss cluster-based merging methods for compression of trained neural networks. In addition to a data-free contribution-weighted averaging method, we propose neuron-merging methods in which neuron responses are mapped back to the pre-activation space via the inverse activation function, and the weights and biases of each representative neuron are estimated using the least-squares method. We also examine both a data-assisted strategy with actual training inputs and a data-free strategy using randomly generated inputs. The comparisons provide empirical evidence, in the tested sigmoid networks, that weight information is particularly useful for clustering whereas activation information is useful for representative-neuron reconstruction in the merging process.
cs.LG / 40 / 2610.02563
OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang
cs.LG · cs.AI
Abstract
We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks. Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks. We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at https://github.com/Roblox/open-game-eval.
cs.LG / 41 / 2610.02574
DISSOLVR: An Interpretable and Fast Framework for Aqueous and Organic Solubility Prediction
Vansh Ramani, Har Ashish Arora, Dhairya Kuchhal, Sayan Ranu, Tarak Karmakar
cs.LG · physics.chem-ph
Abstract
High-fidelity solubility prediction is fundamental to pharmaceutical development and environmental partitioning, where accurate modeling must couple molecular structure with thermodynamic behavior across diverse chemical environments. However, recent advancements have been dominated by deep learning architectures that often sacrifice physical interpretability for predictive power. We challenge this trend by showing that state-of-the-art performance does not require such non-transparent architectures. To address this, we introduce DISSOLVR, a transparent framework for molecular solubility prediction. In addition, we perform a comprehensive literature review and a benchmarking study against various methods. We show that DISSOLVR approaches the aleatoric limit of experimental uncertainty and achieves OOD generalization through structural invariance, derived by mapping molecules to physically-grounded descriptors. Then, we present an LLM-assisted post-hoc explanation pipeline that bridges the gap between symbolic model artifacts and chemically grounded narratives. Finally, a comparative benchmark of a survey involving 22 expert chemists reveals that expert evaluators provide deep insights.
cs.LG / 42 / 2610.02584
Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute
Vu Quang Hoang, Nghia Hieu Nguyen
cs.LG
Abstract
Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with $K=2$ performing best among the configurations tested.
cs.LG / 43 / 2610.02594
How Causality Bridges the Semantic Gap
Shuhao Zhang, Xuran Zhou, Han Guo, Pengtao Xie, Yujia Zheng
cs.LG · cs.AI · cs.CL
Abstract
Numerical measurements capture how a system behaves, but often leave the meanings of its variables unspecified. Some variables are measured but never labeled, and others are never measured at all. Existing methods assign semantics to such variables by consulting general human knowledge, but this inherits its biases where that knowledge exists and offers nothing where it does not. We bridge this gap between measurements and their meanings with causal structure instead, reading a variable's semantics from how it acts on other variables. We formalize this as structure-constrained semantic alignment, in which the embedding of each unnamed variable is solved under the dependence relations implied by the causal graph, with the embeddings of a few known names as anchors. Accordingly, we build CausalBridge, a framework that discovers the causal graph from the measurements, latent variables included, solves for the embeddings under those relations, and expresses them as names through a language model. The causal structure reflects the mechanism that generated the measurements and is recovered from the measurements alone, which may make it the one source of information free of bias from human knowledge. We evaluate CausalBridge on five questionnaires and three robotics scenarios, with 20 to 90% of the variable names masked. It recovers the semantics of observed and latent variables more accurately than existing methods that rely on association, and its lead widens as less of the system is documented. The graph it discovers names variables as accurately as the documented one, and a new system is named in minutes and at a fraction of the cost of sampling methods. Once the semantic gap is bridged faithfully, machines can understand the world and take actions causally.
cs.LG / 44 / 2610.02598
Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights
JuneHyung Kim, Sankeerth Durvasula, Nandita Vijaykumar
cs.LG · cs.PF
Abstract
Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these transfers by skipping weights associated with zero or near-zero activations. However, as more activation contributions are omitted, model quality eventually degrades rapidly, indicating that weights associated with small-magnitude activations collectively influence model quality sharply. In this work, we improve the trade-off between model quality and decoding performance when exploiting activation sparsity. Our key idea is to replace the binary choice of whether or not to read a weight with three options: fully retain it, approximate it using a compressed weight representation, or omit it entirely. SpAx skips weights associated with activations closest to zero, reads approximate weights for smaller-magnitude activations, and reads original weights for the largest-magnitude activations. Smaller-magnitude activations attenuate the errors introduced by approximate weights, while compressed weight representations require fewer bytes to be transferred. With weights offloaded to CPU memory, SpAx speeds up decoding by 3.86X on average (up to 5.57X) with 16-bit weights and 2.06X (up to 2.74X) with 4-bit weights, at a WikiText-2 perplexity increase of at most 10%. With weights offloaded to flash storage, the speedups are 3.31X on average (up to 4.81X) and 1.54X (up to 2.03X).
cs.LG / 45 / 2610.02610
Seer: Maximum Likelihood Regression for Learning-Speed Curves
Carl Myers Kadie
cs.LG
Abstract
The research presented here focuses on modeling machine-learning performance. The thesis introduces Seer, a system that generates empirical observations of classification-learning performance and then uses those observations to create statistical models. The models can be used to predict the number of training examples needed to achieve a desired level and the maximum accuracy possible given an unlimited number of training examples. Seer advances the state of the art with 1) models that embody the best constraints for classification learning and most useful parameters, 2) algorithms that efficiently find maximum-likelihood models, and 3) a demonstration on real-world data from three domains of a practicable application of such modeling. The first part of the thesis gives an overview of the requirements for a good maximum-likelihood model of classification-learning performance. Next, reasonable design choices for such models are explored. Selection among such models is a task of nonlinear programming, but by exploiting appropriate problem constraints, the task is reduced to a nonlinear regression task that can be solved with an efficient iterative algorithm. The latter part of the thesis describes almost 100 experiments in the domains of soybean disease, heart disease, and audiological problems. The tests show that Seer is excellent at characterizing learning-performance and that it seems to be as good as possible at predicting learning performance. Finally, recommendations for choosing a regression model for a particular situation are made and directions for further research are identified.
cs.LG / 46 / 2610.02611
Scale-Recursive Rectified Flows for Few-Step Precipitation Ensembles
Shunya Nagashima, Takumi Bannai
cs.LG · cs.CV
Abstract
Fine-resolution precipitation estimates support flood risk assessment and water management, but coarse satellite products cannot resolve rainfall within each grid cell. Generative models address this ambiguity by producing ensembles of plausible high-resolution rainfall fields. Among these models, rectified flows generate samples by iteratively transforming random noise into rainfall fields. Reducing the number of sampling steps accelerates generation but can make ensemble members too similar, understating uncertainty. We propose a scale-recursive rectified flow that generates broad patterns before local details and guides sampling-step allocation by comparing ensemble variability with prediction error across spatial scales. Validation scores and rainfall power spectra constrain the allocation to avoid excessive amplification. In satellite-to-radar downscaling over the contiguous United States, our analysis identified broad rainfall patterns as the main source of insufficient ensemble variability under reduced sampling budgets. Allocating more steps to the coarse flow improved probabilistic accuracy and rain detection across training seeds at fixed architecture and computational cost. The proposed model also achieved better probabilistic accuracy with shorter sampling time than a nonrecursive flow using more steps.
cs.LG / 47 / 2610.02615
Quantifying the Value of Constructive Induction, Knowledge, and Noise Filtering on Inductive Learning
Carl M. Kadie
cs.LG
Abstract
Learning research, as one of its central goals, tries to measure, model, and understand how learning-problem properties affect average-case learning performance. For example, we would like to quantify the value of constructive induction, noise filtering, and background knowledge. This paper describes the effective dimension, a new learning measure that helps link problem properties to learning performance. Like the Vapnik-Chervonenkis (VC) dimension, the effective dimension is often in a simple linear relation with problem properties. Unlike the VC dimension, the effective dimension can be estimated empirically and makes average-case predictions. It is therefore more widely applicable to machine and human learning research. The measure is demonstrated on several learning systems including Backpropagation. Finally, the measure is used to precisely predict the benefit of using FRINGE, a feature construction system. The benefit is found to decrease as the complexity of the target concept increases.
cs.LG / 48 / 2610.02659
Distributed Learning with Selective State Space Models: Architecture-Aware Convergence Analysis
Adam Piaseczny, Md Kamran Chowdhury Shisher, Shiqiang Wang, Christopher G. Brinton
cs.LG · cs.AI · math.OC
Abstract
Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated learning methods are largely architecture-agnostic, and do not account for the stability, selectivity, and state-space parameterization that characterize modern selective SSMs. To address this, we derive architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs, and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization. We then numerically validate the single-layer bounds on sequences generated by a teacher SSM, using a learner that follows the analyzed recurrence. We use this analysis to formulate expectations about the effects of local training and client heterogeneity, and examine these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains. These experiments illustrate how SSM-specific bounds can provide a basis for interpreting the behavior of practical federated learning algorithms.
cs.LG / 49 / 2610.02661
AIGS: Adaptive Incremental Gating System for Online Representation Learning in Non-Stationary Data Streams
SiRui He, Kai Liang Lew, Chui Zi Ong, Chean Khim Toa
cs.LG
Abstract
Real-time data streams in Web of Things (WoT) and edge computing environments often evolve through latent regime changes. For online representation learning under strict computational constraints, the central problem is resolving the stability-plasticity dilemma: keeping useful historical knowledge while rapidly reacting to concept drift. Existing methods employ fixed update schedules or rolling windows. However, they suffer from parameter ossification during sudden shifts and waste computational resources when the stream remains stable. This paper proposes the Adaptive Incremental Gating System (AIGS), a lightweight closed-loop state-aware adaptation framework. AIGS introduces the Shock Ratio, an endogenous residual feedback mechanism that normalizes current reconstruction error against recent variation. This signal drives a Continuous Plasticity Controller that smoothly interpolates between learning plasticity and memory retention. By treating representation learning as a closed-loop control mechanism, AIGS avoids catastrophic forgetting and maintains a strictly linear $\mathcal{O}\left(k\cdot d\right)$ per-step complexity suitable for latency-sensitive edge devices. Experiments on real-world smart city dynamic streams-spanning traffic networks, meteorological systems, and industrial infrastructure-demonstrate distinct domain-dependent advantages. On Electricity Transformer Temperature datasets, AIGS achieves preventative early-warning lead times of 8.31 (ETTm1) and 9.88 (ETTm2) steps under gradual degradation. On Performance Measurement System traffic datasets, it shows significantly faster post-shift recovery after abrupt mutations. On the highly noisy Weather dataset, it improves anomaly recall while resisting stochastic noise overfitting. These findings establish AIGS as a practical, plug-and-play adapter for resource-constrained edge monitoring systems.
cs.LG / 50 / 2610.02662
Mind the Refinement Gap: When Safe High-Level Robot Plans Produce Unsafe Executions
Stabak Das, Priyesh Ranjan, Xiangfang Li, Lijun Qian
cs.LG · cs.RO
Abstract
Language-enabled robot systems increasingly combine semantic-graph planning with temporal-logic safety monitors. We investigate a trace-completeness assumption in these systems: whether the high-level action sequence checked by a monitor represents the navigation and implicit action effects induced during execution. We audit this assumption in RoboGuard by comparing its verdict on a surface plan with its verdict on a graph-refined trace under the same Linear Temporal Logic (LTL) specification. Our evaluation comprises 28 controlled cases spanning five action-abstraction families and 14 end-to-end cases in which SPINE [1] generates plans from natural-language instructions while RoboGuard generates scene-grounded safety specifications. In the controlled evaluation, all 12 targeted abstraction cases exhibit the predicted surface-versus-refined discrepancy while all 16 controls behave as expected, motivating graph-based trace refinement as a lightweight mitigation and a diagnostic tool for physical-AI safety monitors.
cs.LG / 51 / 2610.02670
LEAP: Learning Efficient Action Proposals For LLM Agents
Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang
cs.LG · cs.AI · cs.CL
Abstract
LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals for target model to verify. Large drafters match the target more often but take longer to propose, while small off-the-shelf models are fast but rarely make the same decision as the target. We ask a more general question: what determines the end-to-end speedup of action speculation? To answer it, we develop a latency framework for the speculative round. The framework compares what a round gains with what it costs. The gain depends on how well the drafter predicts the target and on how many steps the task can take before it ends. The cost comes from drafting, from waiting for target verification and from executing tools. Guided by the framework, we introduce LEAP (Learning Efficient Action Proposals) which keeps the drafter small and makes it accurate by training it on the target actions sequences. With a small 0.6B model, LEAP agrees with the target on most decisions and makes agents up to 60% faster in end-to-end wall clock time, with no systematic change in task success. Across various datasets, target models and draft models, the framework accounts for most of the measured speedups. We also show the draft model can be online trained with no prior trace collection and match the performance of offline training, making LEAP practical to deploy in the real world.
cs.LG / 52 / 2610.02683
RAOA: Alternating-Operator Neural Computation with Programmable Radio Propagation
Toshiaki Koike-Akino
cs.LG · cs.ET · cs.LO
Abstract
Can programmable radio propagation serve as computational depth rather than only as a communication channel or one-shot analog transform? We introduce the Radio Alternating Operator Ansatz (RAOA), a recurrent computing architecture that alternates an energy-derived problem update with a mixing update over a persistent latent state. Recomputing the problem field after each mix makes repeated passes compositional even when the same learned controls are reused across depth. We evaluate this idea through exact discrete optimization, constrained programmable-propagation simulation, and pretrained-model adaptation. On discrete objectives, repeated execution can improve solution quality without increasing the learned-control count, and the same formulation handles higher-order interactions directly. A passive phase-only free-space model further shows that the required operators can be approximated by programmable propagation while retaining useful downstream behavior despite realization error. When inserted as a zero-initialized residual adapter, RAOA adapts pretrained language models with WikiText performance close to a matched shallow MLP across three model families, while reasoning-task transfer remains model-dependent. Together, these results connect alternating-operator computation, programmable radio propagation, and neural adaptation within one recurrent framework. The RF realization evidence is simulation-based rather than a hardware demonstration.
cs.LG / 53 / 2610.02700
Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
Rui Li, Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, Qi Liu
cs.LG · cs.CL
Abstract
On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.
cs.LG / 54 / 2610.02701
Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers
Kaushik Pendiyala, Haris Zia, Trevin Lee, Timothy Legge, Alejandro J. De Leon, Zihan Zhao, Aaron Wang, Abhijith Gandrakota, Jennifer Ngadiuba, Richard Cavanaugh, Javier Duarte
cs.LG · hep-ex
Abstract
Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly associated with particle identity and kinematics in some configurations, but this structure does not increase monotonically with classification performance. These results highlight the need to distinguish stored parameter capacity, active computation, routing capacity, and routing organization when evaluating sparse expert models for jet classification. Code and experiment configurations are available at https://github.com/kpendiyala/MPT.
cs.LG / 55 / 2610.02705
MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads
Linkai Ma, Xinyu Luo, Mengbo Wang, Ananth Grama, Petros Drineas, Brian Bullins
cs.LG · cs.AI
Abstract
The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (language model head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the language model head $\mathbf{L} \in \mathbb{R}^{V \times d}$, we motivate the use of the $2\to\infty$ operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table $\mathbf{E} \in \mathbb{R}^{d \times V}$, we draw on the $1 \to 2$ operator norm, based on the one-hot input geometry identified by Bernstein & Newhouse (2025). The identity $\lVert\mathbf{L}\rVert_{2\to\infty}=\lVert\mathbf{L}^\top\rVert_{1\to2}$ then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for $\mathbf{E}$ and row normalization for $\mathbf{L}$. Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by $\sim$46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.
cs.LG / 56 / 2610.02716
Differential Privacy of Gradient Descent on Perturbed Objectives
Austin Watkins, Raman Arora
cs.LG · cs.CR · stat.ML
Abstract
Objective perturbation adds a random linear term to a regularized empirical risk and releases the exact perturbed minimizer. We study the finite computation obtained by releasing the $N$-th iterate of deterministic gradient descent on $w\mapsto F(w;S)+\langle z,w\rangle$, where $z\sim\mathcal N(0,σ^2I_d)$ is drawn once before optimization. For strongly convex and smooth objectives with Lipschitz Hessian, we prove an explicit condition under which the map $z\mapsto w_N$ is a $C^1$-diffeomorphism on the bounded domains used in the privacy argument, with a quantitative lower bound on the smallest singular value of its Jacobian. This permits a direct change-of-variables analysis of the finite iterate. For generalized linear models, the resulting privacy-profile bound has no explicit ambient-dimension factor once the iteration condition holds, and its finite-iteration correction decreases geometrically. By letting the free truncation parameter grow slowly with $N$, we recover the corresponding exact-minimizer certificate in the limit. We also bound the expected excess empirical risk by $dσ^2/(2μ)$ plus a geometrically decreasing optimization term, and transfer the result to population risk without an additional multiplicative condition-number factor in the leading statistical terms.
cs.LG / 57 / 2610.02730
Bellman Error Minimization Via Linear Programming Normalization
Haining Yu
cs.LG
Abstract
This paper proposes a new functional approximation approach to reduce Bellman error in high-dimensional dynamic programming and Reinforcement Learning problems. Using a classic dynamic programming problem (network capacity control in revenue management) as the motivational example, the paper illustrates that deep neural networks and linear programming approximation algorithms can be combined to derive approximate solutions to dynamic programming problems. Simulation results show the proposed approximation algorithms achieves competitive performance when compared with benchmark.
cs.LG / 58 / 2610.02738
Inner Momentum for Differentially Private Muon
Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai
cs.LG · cs.CR
Abstract
Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.
cs.LG / 59 / 2610.02740
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu
cs.LG · cs.AI
Abstract
Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.
cs.LG / 60 / 2610.02763
A Two-Stage Cascade for Near-Real-Time Forest Anomaly Detection from Sentinel-1 SAR Time Series
Pann Thinzar Seint, Subas Chhatkuli, Bryan Atwood
cs.LG
Abstract
Tropical forest monitoring is essential for global climate stability and biodiversity preservation. To address the urgent need for rapid, reliable detection of forest loss which is essential for timely intervention against illegal logging, supply chain transparency, land-use governance and carbon market standards, we introduce a two-stage statistics-encoder cascade for near-real-time anomaly detection using Sentinel-1 time series. Our system is designed to overcome two fundamental challenges in remote sensing: the cloud-cover limitations that restrict optical monitoring and seasonal backscatter variation that causes SAR systems to mistake natural moisture changes for forest loss. The architecture integrates two distinct analytical engines to ensure high-fidelity detection: (1) an adaptive, robust-statistics z-score test on co-registered Sentinel-1 VH backscatter, same-season historical baseline and (2) a learned confirmation gate based on the latent-space structural similarity (SSIM) of a convolutional autoencoder trained on stable-forest patches. A candidate disturbance is confirmed as an alert only when both stages agree, and is assigned a confidence score and a Low/Medium/High risk tier from its repeat-occurrence history. The system produces per-alert auditable confidence scores and area-in-hectares estimates directly compatible with Monitoring, Reporting and Verification (MRV) workflows, sustainable forestry management, operational field checks and environmental risk assessments. Beyond its primary application, the model's flexibility allows for critical environmental applications ranging from selective logging to large-scale agricultural encroachment mapping, flood mapping and so on.
cs.LG / 61 / 2610.02764
A Controlled Audit of Personal AI Memory for Rating Prediction
Shivam Gupta
cs.LG
Abstract
In structured rating prediction, does a personal AI use historical item-rating associations, or mainly the user's rating tendencies? We audit this distinction by permuting historical ratings within each user while preserving the exact rating distribution, item support, and metadata. We combine this control with full history, native memory extraction, and matched numerical readers in a publicly frozen evaluation of 400 held-out user profiles and 6,160 target ratings across Coat and MovieLens. On Coat, the tested Qwen-written Mem0 pipeline increases user-macro mean absolute error relative to full history by 0.084 for Qwen and 0.149 for Phi; both family-adjusted bootstrap intervals exclude zero. Correct historical assignments help both readers on Coat, but the corresponding MovieLens effects are smaller and inconclusive after adjustment. A history-only ridge reader outperforms Qwen in both domains and Phi on MovieLens, while the Coat Phi comparison is unresolved. All 2,800 reader calls, including 150 invalid outputs, are retained under a fixed fallback rule. A separate implementation verifies inputs, metrics, and all ten primary contrasts. The contribution is a reproducible diagnostic study showing why extraction, association use, output reliability, and reader choice require separate evaluation.
cs.LG / 62 / 2610.02766
Exact Memory-Time Optimization for Prefix-Cached Language Model Serving
Shivam Gupta
cs.LG
Abstract
Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available. We introduce Prefix-Certificate Retention (PCR), an exact finite-trace formulation for static, grouped, reset-on-access timeouts. Usable-prefix rewards become nodes whose prerequisites are timeout thresholds and preceding hit certificates. The resulting maximum-weight closure reduces to one minimum cut, with graph size linear in the number of block lookups and timeout choices. A breakpoint theorem extends the construction to all nonnegative timeouts without discretization error. We also derive a linear-time-in-grid-size dynamic program for ordered timeouts and bounds that certify the cost of this restriction. Exhaustive small-instance checks and chronological replay of 39,632 public Mooncake requests validate the formulation. On the fixed grid, ordered timeouts attain the unrestricted training optimum in 118 of 120 trace-grouping-price cases. Heterogeneous retention improves several held-out memory-time tradeoffs, but finer training optimization does not uniformly improve transfer. The contribution is a tractable optimization model and an auditable benchmark for retention policies; the experiments measure usable prefix blocks and storage time, not GPU latency.
cs.LG / 63 / 2610.02768
No-Free-Graph: Learning When Multimodal Data Should Be Graphified
Zekai Chen, Kai Hu, YuXin Zeng, Xunkai Li, Xun Wu, Yinlin Zhu, Zhengyu Wu, Xu Wang, Rong-Hua Li
cs.LG
Abstract
Multimodal graph learning has recently emerged as an effective paradigm for in corporating inter-entity relationships into multimodal representations. Existing studies have made substantial progress on how to construct and optimize graphs, but rarely consider a more fundamental question: whether additional relational structures should be introduced for a given dataset and task. Through empirical studies across diverse datasets, tasks, and graph constructors, we reveal that graphification is not consistently beneficial: introducing relational structures can provide substantial improvements in some cases, while offering limited or even negative gains. This observation motivates a new perspective that graph construction should be treated as a selective decision based on its expected utility rather than a default preprocessing step. To address this issue, we propose MAG-SCOUT, a pre-construction graph assessment framework that estimates whether introducing graph structures is beneficial before generating the complete topology. MAG-SCOUT collects limited relational evidence, analyzes its potential taskspecific contribution, and estimates the expected utility of graphification together with construction cost to make a build-or-skip decision. Extensive experiments across six multimodal datasets, three downstream tasks, and diverse graph constructors demonstrate that MAG-SCOUT effectively identifies when graph structures should be introduced, saving 33.1% of task-macro graph work while retaining 96.7% of held-out positive-gain mass under the pre-registered floor.
cs.LG / 64 / 2610.02771
Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback
Khang Luong, Dinh Thai Son, Hoang Ta, Hung The Tran, Tuan Quang Dam
cs.LG · cs.AI · stat.ML
Abstract
We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localization, where direct empirical mean estimation is no longer available and clipping becomes unavoidable. We first formulate a time-uniform 1-bit mean-estimation primitive based on randomized threshold queries and a clipped tail-integral identity. We then embed this primitive into candidate-challenger best-arm identification algorithms. A fixed-clipping algorithm gives a simple anytime $(ε,δ)$-PAC guarantee, while a phased adaptive-clipping algorithm matches the clipping level to the current resolution and yields a gap-adaptive sample complexity. We also prove a $K$-arm worst-case information-theoretic lower bound showing that the logarithmic penalty caused by finite-variance 1-bit feedback is intrinsic. This bound matches the leading dependence of the phased algorithm up to lower-order $\log\log$ factors.
cs.LG / 65 / 2610.02781
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak, Yun He, Richard Yuanzhe Pang
cs.LG · cs.AI · cs.CL
Abstract
Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.
cs.LG / 66 / 2610.02795
Efficient Memory Crystallization for Graph Learning under Non-Stationary Distribution Shifts
Yue Hou, Ruomei Liu, Yingke Su, Junran Wu, Ke Xu
cs.LG
Abstract
Deep graph learning models deployed in real-world systems often need to cope with non-stationary environments, where the underlying graph distribution drifts continually over time. Prevailing solutions rely on training auxiliary generative modules to synthesize memory graphs for cross-domain adaptation, which incurs substantial computational overhead and scales poorly under prolonged distribution shifts. We argue that a more economical path exists: rather than generating memory, one can crystallize it. To this end, we propose Efficient Memory Crystallization (EMC), a training-free test-time framework that distills each incoming graph domain into a compact, semantically faithful memory through a closed-form solution to a memory-oriented distribution-matching objective, thereby eliminating redundant domain information under continual covariate shifts. To preserve both generalizability and adaptability as the model traverses a long sequence of target domains, EMC further models inter-domain dependencies through state-evolving memories and admits a theoretically grounded, tighter generalization error bound than direct adaptation. Extensive experiments demonstrate the superior performance of EMC over state-of-the-art baselines on graphs under non-stationary distribution shifts, while reducing average runtime by 87.4% and GPU memory consumption by 92.4% relative to the recent competitor, making continual graph adaptation practical at scale.
cs.LG / 67 / 2610.02798
Muon Learns Facts Better: Understanding the Role of Spectral Orthogonalization
Xuheng Li, Qiwei Di, Yuan Cao, Quanquan Gu
cs.LG · math.OC · stat.ML
Abstract
The Muon optimizer applies spectral orthogonalization to matrix-valued updates and has shown strong performance in large-scale neural network training, yet the mechanisms of this transformation in feature learning remain poorly understood. In this work, we investigate this question through a tractable factual-recall model, where a fact maps each subject-relation pair to an answer, and a linear transformer learns the subject- and relation-dependent information required to recover this mapping. The transformer is optimized with gradient flow (GF), spectral GF, or Sign GF, which are continuous-time limits of gradient descent, Muon, and Adam, respectively. Prior studies (Nichani et al., 2025) have shown that when the number of subjects exceeds the number of relations, GF learns relation-dependent information before subject-dependent information, producing a feature-separation phase during training. We characterize this separation with the learning times when the subject- and relation-dependent components of the prediction reach a target accuracy. With $S$ subjects and $R$ relations, GF has a learning-time ratio of $\widetildeΘ(\sqrt{S/R})$, whereas Spectral GF reduces this ratio to $\widetildeΘ(1)$. In addition, for fixed $S$ and $R$, the subject- and relation-dependent errors decay as $1/(T\log T)$ in training time $T$ under GF, but as $\exp(-\mathrm{poly}(T))$ under spectral GF. Finally, we show that GF and spectral GF are equivariant under orthogonal transformations of the token embeddings, whereas Sign GF is not: Different orthonormal embeddings can potentially produce no feature separation, a large feature-separation phase, or even a reversed learning order. These results provide a mechanistic view of how spectral orthogonalization can fundamentally reshape feature-learning dynamics.
cs.LG / 68 / 2610.02816
Gated Slot Attention-2: Two-Sided Associative Memory Correction in Linear Attention
Ruijie Li, Shengnan Ding, Weimin Zhang, Derick Tang, Zhanpeng Zeng, Qinsong Zeng, Ming Chen, Jiaxi Hu, Yuxuan Liang
cs.LG
Abstract
Linear attention models have emerged as efficient alternatives to standard attention, but effectively managing their fixed-size recurrent memory remains challenging. To improve memory, recent work has explored two distinct directions: delta-rule variants for precise correction of values associated with keys, and slot-based architectures such as Gated Slot Attention for modeling key and value memories in two stages. We observe that these directions are complementary--the delta rule provides effective memory correction, while the two-stage structure provides a natural way to operate on both sides of an association. Building on this insight, we introduce a new Gated Oja Rule for key-side correction and extend it with decoupled erase and write control to obtain Gated Oja Rule-2. We then introduce Gated Slot Attention-2 (GSA2), which combines Gated Oja Rule-2 for key-side correction with Gated Delta Rule-2 for value-side correction through shared latent slots. We further derive a hardware-efficient chunkwise algorithm for parallel training. Experiments demonstrate that GSA2 consistently improves over strong linear-attention baselines across benchmarks while retaining linear-time sequence modeling and constant-memory recurrent decoding.
cs.LG / 69 / 2610.02822
Adaptive Spectral-Koopman Dynamics Modeling for Temporal Domain Generalization
Tengxue Zhang, Yu Ke, Yang Shu, Chenchen Sun, Yisheng An, Chenjuan Guo, Bin Yang
cs.LG · cs.AI
Abstract
Temporal Domain Generalization (TDG) has emerged to address real-world streaming data with distribution shifts over time. However, existing methods are either prone to overfitting to domain-specific noise in the data space or become overly complex and less interpretable in the parameter space. To bridge these gaps, we propose \textbf{AdaSpecK}, a spectral-Koopman framework with adaptive context extraction for TDG. To mitigate noise fitting to irregularly sampled domains, we introduce spectral-regularized Koopman dynamics modeling, which applies spectral-aware filtering in the latent space to extract denoised low-frequency trajectories and learn a Koopman operator to model the system dynamics in a linearized space. To model complex historical environments under non-stationarity, we design a context-informed heterogeneous pattern extraction mechanism. Specifically, we employ a target-conditioned attention module to attend to distinct past windows, producing a dynamic, target-specific historical summary. By constructing an environmental signature from the current evolutionary pattern, our model adaptively perceives which aspects of the past context are most informative for future prediction via a learned router. Extensive experiments on eight diverse classification and regression benchmarks demonstrate that AdaSpecK achieves state-of-the-art performance. The code and datasets are available at \href{}{https://anonymous.4open.science/r/Ada-Spec-K}.
cs.LG / 70 / 2610.02839
To Explore The Strange New World Beyond Data Distribution: System Behavior, Causality Tax, and Non-causal Base Model
Xianzhi Zeng, Jiangneng Li, Gao Cong
cs.LG · cs.CL
Abstract
We show that the causality of language models (LMs) may not be necessary nor optimal. This is the case when system behavior (denoted as $S$) is incorporated as a first-principle Bayesian feature. Here, $S$ refers to extra dominant factors beyond the data space, and they involve coupled effects. Despite being the de facto foundation of modern architecture, recent studies indicate persistent mismatches and contradictions with causality. These issues largely stem from system behavior rather than the data distribution. We therefore propose the SBD framework, which incorporates $S$ as an irreducible component of the evidence lower bound (ELBO). SBD theoretically reveals a counter-intuitive Causality Tax phenomenon, where causality emerges as a suboptimal approximation with an additional structural error, due to the obliviousness to $S$. To address the challenge of latent variable analysis, we validate the SBD-predicted impact of $S$ via implicit measurements, theoretical-bound-guided controls, and Neural Tangent Kernel (NTK) evaluations. In particular, we construct Green Shell (GSH) to show the possibility of reducing Causality Tax. GSH is a non-causal variational family, and it replaces the sequential dependency chain of $S$ components with a divide-and-conquer partition. NTK spectra in the lazy-training regime confirm that GSH always achieves significantly tighter error bounds than causality, with $7dB+$ improvement in signal-to-noise ratio. In the relatively later stage of lazy-training, GSH further leads to superior generalization (up to $20\%$ richer multi-scale fitting capabilities). Taken together, SBD establishes system behavior as a complementary theoretical abstraction besides causality and distribution fitting, opening new research avenues such as designing and optimizing LM base models.
cs.LG / 71 / 2610.02846
Understanding Enrichment in Reinforcement Learning
Jinwoo Kim, Shraddha Barke
cs.LG
Abstract
When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in order to avoid the high variance of correction. It thus remains unclear what exactly is gained or lost in RLVR by correcting enriched rollouts. We show, mathematically, that omitted or truncated correction implicitly reweights the defined reward, and we decompose the resulting gradient error into scale, rotation, and variance to explain their distinct effects on learning. To make correction practical, we develop a novel sequential Monte Carlo (SMC) weight correction mechanism that, under stability and particle-order assumptions, tempers the exponential compounding of standard correction variance over the length of a sample to an additive accumulation. We then apply our analysis of enrichment to interpreting the results of a fine-tune of Qwen3-1.7B on a sparse band of OpenMathReasoning, establishing a concrete mechanism of how enrichment, both corrected and uncorrected, helps avoid collapse in sparse domains. Our main contribution is to understand, in general, how enrichment and correction can affect training, as opposed to claiming that either mode of operation is superior to unenriched RL.
cs.LG / 72 / 2610.02860
Counterfactual Action Evaluation, Observation Bottlenecks, and Representation Geometry in Joint-Embedding Predictive World Models
Arjun Subramanian
cs.LG
Abstract
Low latent prediction error does not establish that a world model distinguishes the consequences of its actions. We introduce an evaluation protocol that traces the same intervention through simulator state, raster observations, target embeddings, and predictor outputs. Exact simulator-state forks in a controlled deformable-physics testbed reveal distinct bottlenecks. Changed commands alter particle motion, yet 41.5% of one-step raster pairs are identical. Observation loss is not the whole explanation: among 579 high-visibility counterfactuals, median predictor-to-target response is 0.0051 and 0.0217 across two seeds, falling to 0.0027 and 0.0116 after variance normalization. An isotropic state perturbation matched to the target counterfactual embedding shift produces 190x and 53x larger predictor changes on the same visible pairs, isolating action-path under-use rather than a dead or globally shrunk predictor. MSE-only training gives 8.36x lower 10-step latent error in matched seeds, but in spectrally concentrated spaces; one VICReg target encoder is also strongly concentrated, so neither error nor rank alone certifies physical state. Finally, stiffness remains near chance even from full-resolution rasters and mechanical state while privileged material parameters decode perfectly, indicating weak identifiability under this excitation rather than encoder discard. These results motivate auditing physical effect, observation visibility, representation geometry, and action dependence separately.
cs.LG / 73 / 2610.02864
NeuroLens: Learning Latent Embeddings of Neural Semantics from Chronic Recordings
Hanrui Lyu, Baiyuan Chen, Tianshu Tan, Matthew R. Whiteway, Maxwell D. Melin, Ji Xia, Linyang He, Bradly C. Stadie, Anne Churchland, Liam Paninski, Yizi Zhang
cs.LG · q-bio.NC
Abstract
Understanding how neural activity represents higher-order cognition and how these representations evolve over time has long been a central pursuit in neuroscience. However, current analytical tools cannot easily distinguish representational plasticity from recording instability in chronic neural recordings. Here, we introduce NeuroLens (Latent Embeddings of Neural Semantics), a self-supervised model based on the Joint-Embedding Predictive Architecture (JEPA) framework that learns denoised, semantically informative latents from chronic neural recordings. An adaptive encoder maps changing neural populations into a common latent space, while a temporal predictor learns structure that supports prediction of future latent states. By predicting in latent space, NeuroLens captures temporally predictable structure and reduces sensitivity to transient, recording-specific variability. Across chronic intracortical data in mice and humans, the learned representations improve decoding of decision-making and semantic task variables. Multi-day pretraining enables generalization to future sessions, rapid few-shot adaptation to unseen neural populations, and more stable decoding over time than state-of-the-art baselines. Together, these results establish NeuroLens as a new paradigm for studying how neural representations change during learning and over long timescales.
cs.LG / 74 / 2610.02866
DIVINE: Simple Cross-Market Stock Pretraining via Diverse Indicator Reconstruction
Kuan-Yu Chen, Shu-Cheng Zheng, Yu-Chen Den, Wei-Cheng Liao, Tien-Hao Chang
cs.LG · cs.CE
Abstract
Financial time-series pretraining typically learns from masked observations, contrastive relations, or future outcomes---yet existing objectives struggle to simultaneously avoid future-supervision uncertainty and maintain return-prediction alignment. We propose DIVINE (DIVerse INdicator rEconstruction), a simple cross-market pretraining framework that reconstructs technical indicators from raw OHLCV history. Computed from observed price-volume history, technical indicators provide consistently defined supervision across markets while summarizing diverse market dynamics with established relevance to return prediction. Pretrained jointly on six-equity market datasets, DIVINE reconstructs 77 targets derived from 16 standard indicators and transfers only the learned encoder to downstream stock ranking. Across all six markets, DIVINE achieves the strongest average portfolio performance with a lightweight 0.05M-parameter encoder, outperforming pretraining baselines and matching or exceeding substantially larger financial foundation models, while remaining robust and data-efficient. Systematic analyses show that indicator diversity and market diversity provide complementary gains in transfer. Together, these results suggest that supervision design and cross-market diversity---rather than model scale---are the key drivers of strong, transferable financial representations.
cs.LG / 75 / 2610.02868
Distributionally Robust Survival Models under Subpopulation Shift and Outlier Contamination
Seonghwi Kim, Sung Ho Jo, Minwoo Chae
cs.LG · cs.AI
Abstract
Learning robust survival models under distribution shift is an important but challenging problem in many applications. In heterogeneous populations, a model that performs well on average may still perform poorly on certain subpopulations, and this issue becomes even more severe when the training data are contaminated by outliers. In this paper, we propose a novel distributionally robust framework for survival analysis that jointly addresses latent subpopulation shift and outlier contamination. The proposed method combines an outer minimization that selects a refined nominal distribution by reducing the influence of contaminated samples and an inner maximization that focuses on the most challenging subpopulation. This formulation directly accommodates non-decomposable survival losses while preserving interactions across samples, including the risk-set structure of the Cox negative partial log-likelihood. We develop an alternating gradient-based algorithm with outer updates derived from the KKT conditions of the inner maximization. Experiments on simulated data and two survival benchmarks demonstrate that the proposed method remains robust when subpopulation shift and outlier contamination occur simultaneously. It stabilizes training in contaminated settings and substantially improves worst-group performance across both linear and nonlinear survival models, while maintaining competitive and sometimes superior overall performance.
cs.LG / 76 / 2610.02872
Peer Effects in Signed Networks: Separating Influence Through Positive and Negative Ties
Xiaojing Du, Jiuyong Li, Lin Liu, Debo Cheng, Jixue Liu, Thuc Duy Le
cs.LG
Abstract
Evaluating network interventions requires understanding how treatment affects people through their social relationships. Counting treated neighbors without distinguishing supportive and antagonistic ties can conceal opposing influences. We define effects through positive and negative ties, their interaction, and a sign-composition effect of reallocating treatment between the two types at a fixed total, and give their identification formulas. Under sign-blind assignment, we show how ignoring signs mixes the effects of the two tie types. We propose SiDE (Signed-exposure Doubly robust Estimator), which combines sign-specific outcome models with exposure probabilities induced by individual treatment assignment. We establish double robustness of its score and assess approximate intervals that account for overlapping neighborhoods. Semi-synthetic experiments on six real signed networks demonstrate accurate effect estimation and examine the limits of interval coverage. An exploratory reanalysis of published school-experiment data yields a positive estimate of the peer effect through spend-time ties on wristband wearing, but the intervals for all four effects include zero after adjustment for multiple comparisons. This framework can inform network intervention design by showing when influences through the two tie types reinforce or offset one another.
cs.LG / 77 / 2610.02881
Toward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach
Xunkai Li, Chenxi Wan, Yinlin Zhu, Wang Luo, Hongchao Qin, Rong-Hua Li, Guoren Wang
cs.LG
Abstract
Multimodal graph foundation models (MGFMs) seek to learn generalizable representations from large-scale graphs with heterogeneous node modalities. However, real-world Multimodal-Attributed Graphs (MAGs) often contain incomplete node attributes, limiting the scale and diversity of available pretraining corpora. Besides, existing MGFMs primarily incorporate graph topology as structural context, overlooking its role in guiding multimodal binding and shaping a unified representation space. To address these challenges, we propose GraphBind, a topology-driven approach that uses graph topology to bind rich modality information into a unified shared space. GraphBind is motivated by the stability of graph topology, which provides structural references and complementary semantic information for multimodal binding. Concretely, GraphBind uses topology to organize self semantics and reliable neighborhood semantics into a global shared space that integrates structure and semantics, and adapts this space to discriminative and generative tasks through lightweight interfaces. Extensive experiments against 11 representative baselines demonstrate that GraphBind achieves leading performance on both discriminative and generative tasks, with relative improvements of up to 28.1% over the strongest baseline.
cs.LG / 78 / 2610.02882
DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs
Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim
cs.LG · cs.AI
Abstract
Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a 1.5$\times$ end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3$\times$ relative to weight-only baselines.
cs.LG / 79 / 2610.02894
When Can We Trust the Matching Principle? Robust Deployment Geometry Under Finite-Sample and Model Uncertainty
Vishal Rajput
cs.LG · cs.AI
Abstract
Match only geometry you can identify; otherwise spread the penalty. We quantify that decision by the trust ratio tau = epsilon / gamma (estimation uncertainty over spectral separation). Under the linear-quadratic Matching response, oracle-relative drift between estimated and oracle projector matching scales as tau^2 for probes in the chosen top-r deployment subspace -- O(tau^2) in the Davis-Kahan separation region tau < 1/2, with practical usefulness depending on constants. Confidence-Calibrated Matching (CCM) turns tau into a policy -- directional when tau is small, progressively isotropic when not -- with thresholds from calibration, not from the theorem (match sits in the separation region; soft is mostly heuristic). Experiments show both regimes, including UCI HAR embeddings where always-match is worse than abstain on every cell.
cs.LG / 80 / 2610.02907
Do ResNets Route? Sparse Interaction Experts in Residual Networks
Liang Yan, Siying Chen, Kaijie Chen, Bo Li, Jinghao Zhang, Mu Miao
cs.LG
Abstract
Residual networks execute every block for every input, yet their functional contributions need not be input independent. We formulate a trained ResNet as a set function over binary residual-branch masks and apply Möbius inversion to decompose its output exactly into individual residual corrections and higher-order interactions. For smooth residual stacks, we show that each fixed $k$-way interaction scales as $\mathcal{O}(λ^k)$ under residual scaling. Exhaustive analysis of ImageNet-pretrained ResNet-18 and ResNet-34 reveals that interaction mass peaks at orders five and ten, respectively, rather than at low orders. Reducing the residual scale shifts both spectra toward lower orders, but also changes model predictions. The interaction coefficients are concentrated in magnitude but not hard sparse, and prediction-preserving sparsity weakens with depth. Crucially, the dominant interactions vary across inputs and predicted classes around a shared global core, while their overall complexity changes little with sample difficulty. These results show that dense ResNets implement an implicit form of soft routing: every block is executed, but different inputs rely on different residual interaction experts. Routing can therefore emerge at the level of functional contribution without an explicit router or sparse execution.
cs.LG / 81 / 2610.02909
Constraint-Aware Training
Jinwoo Kim
cs.LG
Abstract
When generating programs with language models, constrained decoding can apply program analyses to exclude tokens that violate syntax, scope, or typing rules. However, there is a duplication: standard training already teaches the model to suppress the tokens rejected by these analyses. This duplication leads to the question: if we will perform some analysis to filter a set tokens out during inference anyways, can we avoid teaching the model the said analysis altogether during training, and does this externalization lead to more efficient models? This paper defines a general constraint-aware objective satisfying this externalization desideratum and formalizes the benefits of externalization into three concrete theorems about model size and data efficiency. We show, through a controlled synthetic experiment, that the theorems survive training dynamics: constraint-aware training yields lower prediction loss at a matched parameter count and data compared to ordinary cross-entropy training, motivating training objectives that incorporate the analyses used during generation.
cs.LG / 82 / 2610.02911
Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models
Taiheng Pan
cs.LG · cs.CL
Abstract
Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differed from their description. In each case the logged quantity looked consistent with a working setup, while the quantity that defines the comparison went unchecked. With the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin. We contribute the signature of each detail and its effect on the comparison, reference results for TIS and uncorrected GRPO under sampler lag, and the PTH checklist.
cs.LG / 83 / 2610.02951
Dynamic Expert Pruning for Multi-Agent Systems
Jabin Koo, Soheil Abbasloo, Sungjae Lee, Jungseul Ok
cs.LG · cs.AI · cs.MA
Abstract
Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning reduces this footprint, yet existing methods are static --- a single mask, calibrated offline, is applied to the model for every subsequent request. This assumption can fail when the workload is heterogeneous, most prominently in multi-agent systems, where one backbone serves many tasks and roles at once: our analysis shows that different tasks and roles recruit different experts, while static methods assign one fixed subset to all of them. We therefore propose Dynamic Expert Pruning (DEP), which rests on a finding we establish here: an agent's system and task prompts are by themselves sufficient to identify the experts that agent and its task require, since that text already describes what the agent will do. A lightweight predictor, trained once on workflow transcripts, turns those prompts into a specialized per-request mask in a single forward pass, with no per-configuration calibration. Across diverse tasks and roles, model scales, and MoE architectures, DEP achieves better overall accuracy than static pruning and merging baselines, and generalizes to workflows unseen in training without retraining. Its margin over those baselines is largest when few experts are retained, suggesting that the role specialization inherent to multi-agent systems permits sparser serving than static pruning allows.
cs.LG / 84 / 2610.02953
SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention
Zihan Teng, Jiayu Zhao, Wentao Ren, Minhao Fan, Tianrui Ma, Song Chen, Weichen Liu
cs.LG
Abstract
Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compression strategies, token-wise methods reduce cached states but risk information loss through eviction or condensation, while feature-wise methods reduce per-token KV dimensions but can require full-dimensional reconstruction to apply positional embedding, limiting decoding speedups. We introduce SlimKV, a question-agnostic joint token-feature KV-cache compression method. SlimKV uses low-rank-aware training to compress long contexts into beacon memory states with latent KV representations, together with layer-adaptive rank allocation. We further uncover a positional asymmetry: removing key-side RoPE affects beacon and raw tokens differently, with much smaller degradation for beacon tokens. Exploiting this asymmetry, SlimKV trains beacon KV projections under a K-RoPE-free constraint and enables latent-space attention during decoding, mitigating reconstruction latency. On LongBench, SlimKV outperforms baselines at 16x/32x compression and remains leading at 4x/8x, where it retains over 96% of the uncompressed model's score. Needle-in-a-Haystack confirms robustness across evidence positions, and efficiency evaluation shows up to 7.34x attention speedup and 3.38x end-to-end decoding speedup over the uncompressed model at 128K length.
cs.LG / 85 / 2610.02957
Understanding Trajectory Heterogeneity in Federated World Model Learning
Yipan Wei, Zhaokun Yan, Ziming Hong, Jiaqi Wu, Lixu Wang
cs.LG · cs.CL
Abstract
World models learn state evolution from trajectories, making access to temporal context a central training requirement. Federated learning can use distributed records, while ownership boundaries within a trajectory restrict the examples each client can construct. Our study benchmarks this cross-time setting through hourly action-conditioned clinical prediction on eight MIMIC-IV disease cohorts, comprising 40.87 million transition memberships. We specify severity-based client ownership, patient-separated construction, local history and future-window rules, and paired rollout evaluation from one to 32 hours. A matrix of ten federated algorithms covers 32 disease--partition configurations under five rounds of ten-percent participation. Three findings emerge from existing results and training logs. First, client ownership and participation jointly restrict long-window coverage: only 7.55\%--21.36\% of pooled-available 32-step windows have a locally complete anchor visited during training, averaged across diseases. Second, finer severity partitions accompany higher FedAvg error in 15 of 16 paired comparisons, while algorithm gains are small and horizon-dependent: FedProx reduces mean error by 0.56\%, with no consistent improvement at 32 steps. Third, algorithm labels conceal distinct update behavior, including inactive extrapolation and orders-of-magnitude differences in update scale. Cached-update performance also varies strongly across trajectory partitions under the same benchmark protocol. These results establish temporal access, participation coverage, optimization behavior, and horizon-resolved prediction as complementary dimensions for evaluating federated clinical world models.
cs.LG / 86 / 2610.02960
Dirac-Interconnected Neural Elements: Discovering Modularity in Physical Systems Without Reduction
Reiho Li, Razmik Arman Khosrovian, Takaharu Yaguchi, Hiroaki Yoshimura, Takashi Matsubara
cs.LG
Abstract
Deep learning has shown remarkable success in the data-driven modeling of dynamical systems. Much of its success is attributed not to the flexibility of neural networks but to inductive biases based on physical prior knowledge, such as energy conservation and symplecticity. However, existing methods do not fully exploit the fact that real-world physical systems are interconnections of components. Some methods require the interconnection to be known a priori, while others assume the system to be reducible to an ordinary differential equation (ODE) and learn only the reduced ODE, discarding the algebraic constraints imposed by the interconnection. Here, we propose Dirac-interconnected neural elements (DINEs), a neural network model that represents a physical system as a differential-algebraic equation (DAE), whose algebraic constraints are given by a Dirac structure in kernel representation. With DINEs, we simultaneously identify from data the interconnection among the components as a Dirac structure and learn the characteristics of the components as neural networks. This allows us to keep the learned subsystems in unreduced form and isolate or compose them to make a new system without retraining. Moreover, DINEs can handle partially observable systems. Experimental results demonstrate these capabilities on physical systems beyond the reach of existing methods.
cs.LG / 87 / 2610.02990
Differentiable Koopman Operator for Contrastive Learning on Dynamic Graphs
Md Abrar Jahin, Taufikur Rahman Fuad, Md Rizwan Parvez
cs.LG
Abstract
Real-world interaction networks are inherently dynamic: edges form and dissolve as node behavior shifts over time. Most snapshot-based contrastive methods encode temporal dependencies implicitly in encoder weights, without an explicit model of how node representations evolve, making them brittle under distribution shifts. We propose KAIROS (Koopman-Aligned Invariant Representations for Open Dynamic Systems), a self-supervised framework that embeds a differentiable Koopman operator within a dynamic graph contrastive learning loop to linearize temporal evolution in the learned embedding space. A dual-view encoder pairs raw node features with a graph-diffused structural view and is optimized with multi-granularity contrastive objectives across temporal windows. For anomaly detection, KAIROS uses the Koopman prediction residual together with temporal inconsistency and local neighborhood deviation to separate irregular behavior from predictable graph evolution. Evaluated on nine dynamic graph benchmarks, KAIROS achieves state-of-the-art anomaly detection results on all nine datasets, with gains of up to 23.15 ROC-AUC points over prior work, while remaining competitive for unsupervised node classification. These results show that explicit dynamics modeling provides a scalable and effective inductive bias for temporal graph representation learning.
cs.LG / 88 / 2610.02994
Sentry: Learning to Recover from LLM Agent Failures at Test Time
Changxiu Ji, Amy Lu, Qizheng Zhang, Kunle Olukotun
cs.LG · cs.AI · cs.CL
Abstract
LLM agents often fail mid-task due to invalid tool calls, repeated actions, or poorly grounded reasoning, and learning from these failures is a path to reliability. We find that how failure knowledge reaches the agent matters as much as what it contains. Failure lessons are conditional: kept in the agent's context, they misfire when their failure is absent, and removing them from an evolving playbook improves performance. Runtime interventions, in contrast, act only when a failure occurs but do not learn from their repairs. We argue that failure knowledge is conditional knowledge and should be conditionally exposed, and instantiate this principle in Sentry, a failure-management layer that runs alongside the agent. When Sentry detects a failure, it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did; the full playbook never enters the agent's context. Across multiple agentic benchmarks, Sentry outperforms the strongest runtime-intervention baseline on every benchmark, by 37\% on average, and the strongest context-evolution baseline by 39\% on the two benchmarks where both are evaluated; combining Sentry with context evolution yields further gains. Learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance even when relevant lessons remain available on demand.
cs.LG / 89 / 2610.03000
Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability
Ambarish Moharil
cs.LG · cs.AI · cs.NE · physics.data-an
Abstract
Intrinsic explainability remains a challenging problem, particularly in contexts where multilayer perceptrons (MLPs) require dynamic re-training within an optimization environment. This paper investigates how MLPs and their training dynamics can be represented and studied in non-Euclidean spaces; our representation features the Poincaré model of hyperbolic geometry. We aim to capture the geometric evolution of their weighted topology and self-organization over time. Instead of restricting the analysis to single checkpoints---as per established measure-based explainability methods---we construct temporal \textit{parameter graphs}, i.e., snapshots over time $T$ steps of the optimization/training process for MLPs. This reflects the view that neural networks encode information not only in their weights but also in the trajectory traced during training. Drawing on the idea that many complex networks admit embeddings in hidden metric spaces where distances correspond to connection likelihood, we present a geometric and temporal graph-based metalearning framework for obtaining dynamic hyperbolic representations of the underlying neural parameter graphs. Our model embeds temporal parameter graphs in the Poincaré model ball, and learns from them while maintaining equivariance to within-snapshot neuron permutations and invariance to permutations of past snapshots. In doing so, the approach preserves functional equivalence over time and recovers the latent evolving geometry of the network. Experiments on regression and classification tasks with trained MLPs show strong meta-network performance, accompanied by hyperbolic temporal representations. This reveals how the network structure emerges over time under specific training environments, thus providing insights into the network's self-organization.
cs.LG / 90 / 2610.03001
Neural Data Needs Semantic Tokenization: Behavioral Events as Boundaries of Session-Transferable Tokens
Sangyoon Bae, Jiook Cha
cs.LG
Abstract
Extracellular electrophysiology records a different set of neurons in every session. Neural foundation models embed each neuron and each session into their tokens, so every new session is an input they have never seen, and they fail to generalize to it. A tokenizer for new sessions needs a unit that every session shares and that carries behavior. Population activity offers such a unit. It evolves on a low-dimensional manifold that persists across neuronal turnover and across animals once sessions are aligned. This manifold changes regime at task events such as stimulus onset and movement onset. Within each regime the population occupies a state, the part of the manifold it spans between two events, and each state carries its own behavioral meaning. We propose Tokenization with States (TWS), which segments each trial, one repetition of the task, at these events and converts every state into tokens of population geometry, with no neuron or session embedding. On held-out International Brain Laboratory (IBL) sessions, TWS decodes movement even from regime boundaries that carry no information about the target, while a per-neuron foundation model pretrained on those sessions decodes at chance. Frozen after training on mice alone, TWS transfers to macaques and Utah arrays with only a linear probe. On reach direction, a target that its boundaries do not define, it achieves a Matthews correlation of 0.23, where the event time alone achieves $0.01$. For cross-session generalization, the token matters more than the model on top.
cs.LG / 91 / 2610.03007
AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning
Han Yu, Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Hejian Sang, Han Shi, Menglin Zhou, Xuanzhao Dong, Minzhou Huang, Rui Cai, Hao Wang, Alborz Geramifard
cs.LG · cs.AI
Abstract
Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.
cs.LG / 92 / 2610.03019
Signal Simplification Is Not Predictive Simplification: Diagnosing Residual Neural Forecasting in Short-Horizon Volatility
Bingqi Lian, Linfeng Cheng, Mei Lu, Jerry Wu
cs.LG
Abstract
Hybrid statistical-neural pipelines often assume that a successful statistical first stage leaves a cleaner and more learnable residual target. We examine that assumption in short-horizon volatility forecasting through a signal-forecast-system diagnostic framework. Across five liquid U.S. assets, a volatility-aligned HAR-style model outperforms AR, MA, and ARIMA. Within expanding training windows, the pre-standardization fitted residual process used to construct residual-LSTM sequences has about 82% lower variance than the corresponding target and near-zero lag-1 autocorrelation; independently, rolling pseudo-out-of-sample HAR errors show about 74% variance reduction and similarly weak lag-1 dependence. Residual-only LSTM augmentation nevertheless raises mean squared error from 0.3049 to 0.3594 on average, with deterioration on every asset. Pure LSTM records the lowest selected pseudo-out-of-sample MSE, 0.2649, while the residual hybrid requires substantially more end-to-end runtime without improving accuracy. We describe this pattern as forecaster-preconditioner asymmetry: first-stage forecasting success and statistical residual simplification need not translate into useful downstream neural preconditioning.
cs.LG / 93 / 2610.03027
Tailoring the Quantization Space for 1-Bit KV Cache Compression
Minsoo Cheong, Donghyun Son, Sungjoo Yoo
cs.LG · cs.AI · cs.CL
Abstract
The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to $14\times$ larger batch sizes and achieves $1.87\times$ higher peak throughput compared to the BF16 baseline.
cs.LG / 94 / 2610.03035
Balancing Multimodal Learning via Functional Progress
Zhongjing Gu, Fengqiang Wan, Yiming Cui, Yufa Feng, Yang Yang
cs.LG
Abstract
Multimodal learning often suffers from modality imbalance, where the joint optimization process is dominated by a single modality. Existing methods typically estimate modality imbalance from score disparities derived from prediction uncertainty or optimization statistics. However, due to distinct prediction uncertainty and learning dynamics across modalities, direct comparison of such scores may misinterpret intrinsic modality differences as progress gaps, leading to biased imbalance estimation. In this paper, we propose Function-Space Guided Multimodal Optimization (FGMO), which leverages a function-space progress signal to assess modality-wise optimization progress and coordinate optimization across modalities to alleviate modality imbalance. Specifically, we introduce Functional Progress Estimation (FPE) to measure each modality's update-induced function-space response and calibrate it against a loss-aligned unimodal reference, producing a comparable progress signal. Based on this signal, Functional Response Control (FRC) redistributes modality-level function-space budgets and realizes the target responses through tensor-wise learning-rate adjustment. Theoretical analysis establishes a one-step target-contraction property of FRC under bounded controller-state mismatch, and extensive experiments demonstrate the effectiveness of FGMO across multiple multimodal benchmarks.
cs.LG / 95 / 2610.03054
RIPPLE in Still Water: Zero-Shot Clustering in Federated Learning with Wavelet Scattering Transform
Alessandro Licciardi
cs.LG
Abstract
Clustered Federated Learning (FL) partitions a client population into groups of similar local distributions and trains one specialized model per cluster, mitigating client drift that degrades single-model methods under non-IID data. Prior methods discover cluster structure inside the training loop through gradient similarity, loss evaluation, or EM-style updates, thus increasing communication overhead, exposing gradients to inversion attacks, and providing no mechanism to assign clients absent from training. We propose RIPPLE, a clustered FL framework in which cluster assignment is computed entirely offline from a spectral characterization of each client's local data: a variance-weighted principal-component prototype embedded via the Wavelet Scattering Transform and decoded by a Gaussian Mixture VAE trained server-side on synthetic client populations before federation begins. Per-round communication cost matches FedAvg exactly, and a client absent from training obtains a personalized model from a single forward pass, without gradient computation, model evaluation, or extra communication round. We prove that the gap between RIPPLE's surrogate clustered objective and the oracle is bounded by a computable quantity decaying with client sample size and independent of federation duration; per-cluster convergence matches the minimax-optimal rate for non-convex smooth objectives. Across five benchmarks spanning controlled and realistic heterogeneity, RIPPLE consistently outperforms all baselines, with margins growing on the most realistic partitions.
cs.LG / 96 / 2610.03057
When Does Synthetic Relational Data Teach Models to Use Relations? Tracing Predictive Structure from Pretraining Data to Model Behavior
Shivam Dubey, Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Aditya Tanna, Vinay Kumar Sankarapu
cs.LG
Abstract
Relational foundation models are increasingly pretrained on synthetic databases, yet downstream benchmarks reveal little about why one synthetic corpus produces a better model than another. In particular, strong performance may arise from realistic row-level statistics without the model ever learning to use relational structure. We study this as a data-attribution problem: which property of synthetic pretraining data induces relational computation? Using four Relational Transformer checkpoints trained with the same architecture, initialization, objective, and compute budget on corpora produced by four relational data generators, we trace a measurable property of the data to learned computation and downstream behavior. We hypothesize that relational mechanisms emerge when cross-table information is predictively necessary for the masked-cell pretraining objective. RelDiff exhibits by far the largest predictive gain from foreign-key-linked parents, and its corresponding model is uniquely sensitive to foreign-key interventions on unseen databases. This dependence survives a random-initialization control, grows monotonically with the fraction of corrupted links, and localizes to a serial cross-table pathway. Finally, disrupting the same mechanism during downstream inference removes RelDiff's advantage on relational tasks while leaving structure-insensitive models nearly unchanged. These results connect a property of synthetic training data to a learned mechanism and, through intervention, to downstream behavior.
cs.LG / 97 / 2610.03065
Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings
Niklas Emonds, Georgia Koppe
cs.LG
Abstract
Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.
cs.LG / 98 / 2610.03082
Smart Sensing for Safer Bridges: From Sensor Signals to AI-Driven Anomaly Detection
Rahul Jaiswal, Joakim Hellum, Halvor Heiberg
cs.LG
Abstract
Bridges contribute significantly to transportation connectivity and urban development. Therefore, reliable bridge monitoring is crucial for protecting public safety and detecting anomalous behavior in bridge sensor data that may provide early indications of abnormal structural conditions. This paper investigates anomaly detection in real-world bridge sensor data using two different complementary approaches, namely signal processing and the data-driven machine learning model Isolation Forest. The real-time bridge sensor data is collected from an iBridge sensor device installed on a bridge in Norway. The methods are evaluated using anomaly counts, anomaly detection time, processing rate, anomaly rates, visualization, and temporal agreement. Moreover, a controlled anomaly-injection analysis is performed to evaluate the sensitivity of each method. Numerical results demonstrate distinct detection characteristics and computational requirements, highlighting the potential of machine learning, particularly the data-driven Isolation Forest, alongside signal processing for identifying anomalies in bridge sensor measurements.
cs.LG / 99 / 2610.03085
Light Entropic Optimal Transport on Riemannian Manifolds
Xavier Aramayo-Carrasco, Petr Mokrov, Alexander Korotin
cs.LG
Abstract
Entropic Optimal Transport (EOT) has become a practical framework for learning stochastic couplings between complex distributions, with applications in generative modeling and domain adaptation. However, most EOT solvers are designed for Euclidean spaces, while manifold extensions remain limited and often rely on costly iterative methods, simulated dynamics, or generic neural models that do not fully exploit the underlying geometry. We introduce ManifoldLightOT, a light approach for learning kernel-induced EOT couplings directly on common manifolds. Using the kernel form of the EOT solution, we construct geometry-specific Gibbs kernels together with compatible potential parameterizations for spheres, tori, $\mathrm{SO}(3)$, and $\mathrm{SE}(3)$. These choices yield closed-form normalization and directly sampleable conditional distributions. Our formulation naturally extends to products of manifolds, making it applicable to more complex geometries. The parameters of the potentials are optimized directly from samples using Monte Carlo estimates of the learning objective. Through synthetic and real-world experiments, we show that ManifoldLightOT often outperforms existing manifold OT methods while retaining direct sampling.
cs.LG / 100 / 2610.03087
Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines
Maximilian Böther, Josh Wills, Ties Robroek, Sonnet Xu, Paul Burstein, Daniel Zayas, Cody Blakeney, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Haakon Mongstad, Luke Merrick, Pratyush Maini, Ari Morcos, Matthew Leavitt, Ana Klimovic, Bogdan Gaza
cs.LG · cs.AI · cs.DB
Abstract
Deterministic data loading is important for foundation model development: model researchers need confidence that differences they observe across costly ablations are caused by the parameter they changed rather than non-determinism in the training data sequence. The data loader must provide elastic determinism, i.e., a deterministic sequence of global training data batches despite changes to the GPU topology across runs (e.g., due to GPU scarcity), frequent checkpoint-resume cycles, and different data processing execution backends. Achieving this is difficult because modern foundation model data pipelines tokenize, pack, and mix samples online, introducing stateful n-to-m transformations that break sample indexing. Existing data loaders largely assume indexable 1-to-1 pipelines, and the common workaround of offline materialization is expensive and, for some modalities such as video, infeasible. We present Zephon, a data loader for foundation models that supports online, stateful pipelines while providing elastic determinism and efficient resumption from checkpoints. It partitions the global stream into topology-independent lanes, serializes ordering decisions while parallelizing stateless work on interchangeable backends, and checkpoints only bounded in-flight state so recovery cost does not grow with training progress. We evaluate Zephon on text and vision-language workloads and show that it achieves competitive throughput while providing a combination of guarantees that no existing loader offers for online, stateful pipelines.
cs.LG / 101 / 2610.03092
ULTRADISCOVERY: Abductive Exploration in an Interconnected, Epistemically Open Universe
Weihan Li, Tianshi Zheng, Yangqiu Song, Ginny Y. Wong, Simon See
cs.LG · cs.AI
Abstract
Scientific discovery often begins when scattered clues call for a new way of describing the world. Such abductive exploration can require constructing the representation in which an explanation is stated, when the world is epistemically open, and composing evidence scattered across contexts, when it is structurally interconnected. Existing benchmarks rarely separate these two demands or control them independently. We introduce ULTRADISCOVERY, an interactive world of five domains in which an agent revises an initially successful theory and predicts the outcome of an unseen cross-domain intervention. A $2 \times 2$ design leaves the representation open or discloses it, and leaves the evidence distributed or aligns it, with the latent dynamics fixed. With the representation open, agents across eleven models often retract the axiom they were taught, and none introduces the unobserved entity or rewrites the variables that a replacement requires. Disclosure triples intervention requests and adds about one of the eighteen findings the world affords, and alignment adds less. Two vendor-harness systems carry discovery into more domains, and one of them rewrites the variables in Open episodes. No system makes the exact prediction within 200 paid actions. At larger budgets one exact prediction appears with both aids, while every Open episode remains inexact. The results locate the difficulty in the step from accumulating evidence to composing it into a representation that transfers.
cs.LG / 102 / 2610.03093
LS-AR: Future-Predictive Latent Steering in Autoregressive LLMs
Anubha Gupta, Eduardo Pignatelli
cs.LG
Abstract
Standard autoregressive (AR) models process high-level task instructions, state history, and transient tokens within a single shared sequence of tokens. Consequently, they lack the architectural mechanisms needed to isolate macro-objectives from context noise. To overcome this single-channel limitation, we introduce Latent-Steered Autoregressive (LS-AR), a dual-channel architecture that decouples continuous goal steering from discrete token decoding via FiLM conditioning. We evaluate a Static Goal Encoder (P_0) for persistent macro-objective retention across long rollouts and a Dynamic State Tracker (P_t) for recurrent latent updates during generation. On long-horizon retrieval past context limits (H=1024, W=500), LS-AR (Static) achieves 100% target recall where parameter-matched baselines collapse (0%), while increasing throughput by ~35% and cutting peak VRAM by 52.8%. In Blocksworld planning under forced perturbations (k=1), LS-AR (Dynamic) sustains an 89.0% completion rate vs. 71.0% for the baseline, though zero-shot entity scaling (N -> N+1) exposes single-vector capacity limits (0%). Finally, dual-channel authority analysis shows that text goal dropout establishes latent-dominant control, offering structural defence against text prompt injection while introducing a latent vector attack surface.
cs.LG / 103 / 2610.03117
Exploring the Trade-Off Between Structured Pruning and Fault Tolerance in Deep Neural Networks for Space Applications
Toon Vinck, Naïn Jonckers, Jaro De Roose, Jeffrey Prinzie, Peter Karsmakers
cs.LG
Abstract
Deep Neural Networks (DNNs) inherently exhibit a degree of robustness to bit-level faults due to their distributed representation of information. As a model increases in width, this information becomes more dispersed, theoretically reducing the impact of any single bit fault. In this paper, we empirically investigate the relationship between model width and robustness to Single Event Upsets (SEUs). We conduct a comprehensive experiment in which baseline models undergo iterative structured pruning to reduce their width while preserving task performance as much as possible. At each pruning stage, we run a targeted fault-injection campaign to evaluate the model's performance under simulated bit-flip scenarios. Our results show that, although structured pruning increases per-inference sensitivity to faults by reducing redundancy, this effect is effectively counterbalanced by shorter execution time, which lowers the probability of encountering an SEU. These findings suggest that structured pruning can yield significant energy and latency savings without compromising overall reliability, providing useful guidance for designing robust AI systems for space applications.
cs.LG / 104 / 2610.03119
How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning
Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio
cs.LG · cs.AI · cs.RO
Abstract
In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.
cs.LG / 105 / 2610.03127
Coverage You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS
Pedro Brandimarte, Nerea Aranjuelo, Marcos Nieto, Oihana Otaegui
cs.LG
Abstract
Hardware-aware neural architecture search (NAS) is dominated by evaluation cost: every architecture must be trained before its reward is known. Conformal-prediction filters cut this cost by pruning candidates whose predicted-reward upper bound misses a threshold, with a distribution-free guarantee that at most a fraction $δ$ are wrongly discarded. That guarantee assumes exchangeability between calibration and test candidates, which the surrounding reinforcement-learning (RL) loop violates: the policy's proposals improve as search proceeds and, in layer-by-layer construction, shift within every episode. We replace one-shot quantile estimation with online feedback control (Adaptive Conformal Inference, with tuning-free, locally-adaptive, and group-conditional variants), restoring steerable coverage: dialing the target delivers it, monotonically and reproducibly, for arbitrary sequences. Across three neural-network architecture families and both single-step and sequential search (three seeds), it tracks every requested level to within ${\sim}10^{-3}$ while pruning 25-50% of evaluations at no measured accuracy cost, whereas static calibration loses control of its coverage and a Gaussian-process baseline stays conservative regardless of the request. Finally, used as an acquisition function on one constrained testbed, the same optimistic bound beats random search, a gain that fixed optimism already carries and online calibration sharpens. The source code is available at https://github.com/Vicomtech/rl-hw-nas.
cs.LG / 106 / 2610.03135
Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention
Inbasekaran S
cs.LG
Abstract
Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r - while arithmetic mean pooling dilutes the specialized retrieval heads that carry factual recall. We introduce Page-EntroKV, a formal framework for KV-cache eviction operating at the granularity GQA serving actually allocates. Heads within each physical group are pooled by weights derived from sink-isolated collision (Renyi-2) entropy - one inner product per head, computed once at prefill with no calibration - so sink heads cannot masquerade as retrieval heads. Pooled scores are projected onto PagedAttention page frames, and eviction executes at the hardware tuple (layer, group, page). We formalize the union overhead ratio (UOR) and intra-group disagreement, prove an exact identity linking them for two-head groups alongside two-sided bounds at every group ratio, prove strict budget preservation and a finite-context needle-retention bound that arithmetic mean pooling provably violates, and give exact per-layer page accounting. On a pilot architecture (Qwen2.5-1.5B-Instruct, r=6), head-independent replay over 2,240 group measurements yields union overhead up to 4.75x at a 2% budget, while Page-EntroKV holds UOR exactly 1.000; sink isolation removes a 13x sink masquerade; needle recall is 100% versus 0% for mean pooling at a 20% budget; retained cardinality is exact for every page size; and QA and code tasks remain solvable at 20% retention.
cs.LG / 107 / 2610.03165
Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds
Gabor Paczolay, Matteo Papini, Alberto Maria Metelli, Istvan Harmati, Marcello Restelli
cs.LG
Abstract
Several variance-reduced versions of REINFORCE based on importance sampling achieve an improved $O(ε^{-3})$ sample complexity to find an $ε$-stationary point, under an unrealistic assumption on the variance of the importance weights. In this paper, we propose the \algo (Defensive Policy Gradient) algorithm, based on defensive importance sampling, which achieves the same rate without any assumption on the variance of ordinary importance weights. We also establish lower bounds in a generalized black-box policy-optimization model that hides states and actions and permits parameter-dependent rewards. In this model, the optimal rates are $Θ(ε^{-4})$ with bounded-variance one-policy feedback and $Θ(ε^{-3})$ with mean-square-smooth coupled two-policy feedback. Under standard policy-regularity conditions, REINFORCE and \algo realize the corresponding oracle conditions and attain the $O(ε^{-4})$ and $O(ε^{-3})$ upper bounds, respectively. Although the lower bounds do not apply directly to the classical MDP interaction model in which these algorithms operate, this correspondence provides oracle-level evidence that the faster rate of \algo is optimal and genuinely separated from that of vanilla policy gradient.
cs.LG / 108 / 2610.03211
HyperFuse: Fast Self-Supervised Node Embeddings for Attributed Hypergraphs
Megha P, Harshit Kumar, Srajan Agarwal, Anirban Banerjee, Olaf Wolkenhauer, Saptarshi Bej
cs.LG
Abstract
Self-supervised hypergraph representation learning can produce informative node embeddings, but existing methods often require deep encoders trained for hundreds of epochs, making embedding generation costly even for hypergraphs with a few thousand nodes. This limits applications requiring embeddings for many or evolving hypergraphs. We present HyperFuse, a label-free pipeline for fast hypergraph representation learning. HyperFuse (i) computes structural node coordinates by maximizing a spectral relaxation of hypergraph modularity using Banerjee's hypergraph adjacency and a matrix-free operator with cost linear in node-hyperedge incidences; (ii) constructs multi-scale feature summaries and assigns bounded utility weights to hyperedges based on member stability under feature and membership masking; and (iii) trains a lightweight utility-weighted hypergraph encoder for 100 epochs using an invariance-decorrelation objective. We compare HyperFuse with TriCL, SE-HSSL, VilLain, and HypeBoy on nine public hypergraphs using six downstream classifiers and k-means clustering. On the eight datasets where all methods completed, HyperFuse required 8.7 s per dataset on average, achieving 13-179x geometric-mean speed-ups over the baselines. It achieved the highest average accuracy with five of six classifiers, while classification and clustering performance was not significantly different from TriCL and SE-HSSL. Compared with HypeBoy, HyperFuse was 13x faster and 2.1-4.1 percentage points more accurate across all classifiers. HyperFuse provides a practical approach for fast, repeated hypergraph embedding generation.
cs.LG / 109 / 2610.03216
Kernel Singular Value Decomposition with Extension to Multiple Data Sources
Xinjie Zeng, Qinghua Tao, Johan Suykens
cs.LG
Abstract
Kernel Singular Value Decomposition (KSVD) learns a pair of singular vectors w.r.t. an asymmetric kernel matrix, which can be induced by two data sources, e.g., the queries and keys in self-attention or the rows and columns of a given matrix. In this work, we extend KSVD to multiple data sources, namely eKSVD, which conducts joint nonlinear feature learning upon asymmetric kernels. In the primal formulation, the projections associated with each data source are jointly learned to capture maximal information, while incorporating pair-wise couplings. With the Lagrangian and its Karush-Kuhn-Tucker (KKT) conditions, the optimization in the dual leads to a generalization of the shifted eigenvalue problem in Lanczos decomposition theorem of KSVD. Further, a covariance-based framework is derived together with using neural networks (NNs) for explicit feature mappings, complementary to the kernel-based interpretation and optimization. Numerical experiments verify the effectiveness of our eKSVD compared to methods based on Mercer kernels for tackling multiple data sources, and our innovation of deploying NNs demonstrates great flexibility for kernel methods.
cs.LG / 110 / 2610.03223
AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning
Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan
cs.LG · cs.CL
Abstract
Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on subsequent actions, environment transitions, and trajectory length. We propose AdaStep, an Adaptive Step-credit weighting method that controls how strongly each group-derived local advantage modifies the trajectory-level signal. We formulate this weighting as a mean-squared-error estimation problem for the latent step advantage and, under an explicit conditional sampling assumption, derive an optimal per-state shrinkage coefficient. The coefficient admits a signal-to-total-variance interpretation: it preserves local credit when return variation is attributable to the selected action and suppresses it when variation is dominated by downstream randomness. AdaStep requires only lightweight scalar computation, with no critic, additional rollouts, or extra model inference. Experiments with three model backbones on ALFWorld, WebShop, and ScienceWorld show consistent improvements over baselines at low computational cost.
cs.LG / 111 / 2610.03225
The Neuro-Physical Inverter: A Modular Framework for Magnetotelluric Inversion Coupling Ensemble Conditioning with Residual Learning
Jae Deok Kim, Sai Ravela, Rob. L. Evans
cs.LG
Abstract
We present the Neuro-Physical Inverter (NPI), a modular, uncertainty-aware framework for geophysical inversion that couples ensemble-based conditioning with constrained residual learning, demonstrated in the 1D magnetotelluric (MT) setting as a controlled testbed. The framework operates in two stages. An Ensemble-Conditional Gaussian Process (EnsCGP) conditions a prior ensemble of resistivity models on the observed response, producing a physically admissible reference ensemble. A residual-learning neural network then predicts targeted corrections to this reference, trained on synthetic data and fine-tuned per station for field application through a physics-coupled objective. Because an ensemble is conditioned, refined, and propagated through both stages, every estimate carries an associated ensemble spread. Synthetic experiments show that NPI systematically reduces ensemble-mean error without destabilizing the ensemble. Applied to broadband MT data from the Gabbs Valley geothermal region (Nevada, USA), NPI reduces the across-station mean misfit over the mid-period band while retaining comparable ensemble spread. The propagated ensemble yields a factor of uncertainty that serves as an operational measure of constraint within the assumed model class. Both stages are dimension-agnostic in formulation, and the design principles established here are intended to scale to higher-dimensional parameterizations.
cs.LG / 112 / 2610.03256
Cross-cohort TB classification using clinical data gathered in Uganda and South Africa
Joshua M. Jansen van Vüren, Devendra S. Parihar, Daphne Naidoo, Marisa Klopper, Frank Cobelens, Lutz Kolbe, Kimsey Zajac, Willy Ssengooba, Moses Joloba, Grant Theron, Thomas R. Niesler
cs.LG
Abstract
We present a first evaluation of machine learning applied to patient clinical and demographic data gathered in two different countries for the purpose of tuberculosis (TB) screening to identify people who would benefit from expensive molecular testing. Experiments are based on the recently-compiled CAGE-TB dataset, which includes sub-cohorts of people with presumptive TB presenting at community health care centres in South Africa and Uganda. Three neural network architectures (logistic regression (LR), multilayer perceptrons (MLP) and convolutional neural networks (CNN)) are considered in conjunction with greedy feature selection. For the convolutional neural network, a strategy that jointly optimises feature selection and feature ordering is proposed and shown to lead to consistent development and test set improvements. For all three models, development set area under the receiver operating characteristic (AUROC) curve is improved by 2-7% using feature selection. LR after feature selection achieves an AUROC of 0.8 [0.75,0.86] (95% CI) and 0.84 [0.78,0.9] when testing on the held-out Ugandan and South African data respectively. Although outperforming LR on the development cohort, the deeper networks (MLP, CNN) show inconsistent trends on the held-out cohorts, while LR achieves performance within 1-2% of the best achieved in terms of AUROC. LR narrowly misses the WHO minimum requirements by 4-9% in sensitivity even though the network is being evaluated on a completely held-out cohort. The development of neural-network based classifiers for TB screening therefore appears viable.
cs.LG / 113 / 2610.03258
Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery
Hendrik Suhr, Sascha Xu, Jilles Vreeken
cs.LG · cs.AI
Abstract
Scalable nonlinear causal discovery requires methods that combine flexible mechanism estimators with efficient search over large graph spaces. Several algorithmic families have been proposed to address this challenge, yet their accuracy-runtime trade-offs remain poorly understood. We empirically compare the four major approaches: differentiable structure learning, amortized structure learning, score-matching, and combinatorial search. Our results reveal complementary bottlenecks: differentiable and amortized methods scale well but exhibit an accuracy gap, score-matching methods can be accurate in low dimensions but degrade quickly for increasing feature sizes, and combinatorial methods remain accurate but are slowed by repeated and redundant local scoring. Motivated by this bottleneck, we develop SPADE, a spline-based score-evaluation scheme that compiles sufficient statistics once and reuses them throughout combinatorial search. Under bounded indegree, its Gaussian variant reduces algorithmic complexity from O(nd^3) to O(nd^2+d^3). Empirically, SPADE shifts the observed scalability-accuracy frontier by orders of magnitude: it solves 100-variable problems with 160K samples in seconds and 1600-variable problems with 2.5K samples in minutes, while retaining high structural accuracy across synthetic and real-world benchmarks. These results reveal a substantial shift in the practical scale of combinatorial search and highlight the importance of evaluating scalable causal-discovery methods along the full accuracy-runtime frontier.
cs.LG / 114 / 2610.03259
PaMIR: Open Benchmark of Public Credit-Default Datasets
Mikhail Liashkov, Ilyas Varshavskiy, Shuhratjon Khalilbekov, Azizjon Azimi, Bonu Boboeva
cs.LG
Abstract
We release PaMIR (Public Arrival-ordered Measurement for Inference in Risk), an open benchmark for credit-default prediction when labels are scarce and arrive late. The field's reference benchmark studies use eight datasets each, only two or four of them public. PaMIR brings together 19 public datasets with binary default labels -- 1.24M loans, firms and card accounts from nine countries -- rebuilt from pinned source snapshots by one leakage-audited recipe and never redistributed; to our knowledge it is the one of its kind as of today. Every model is a single function, scored under a repeated i.i.d. split and a label-delayed stream in which each application is scored on arrival, with AUC reported by label budget; fleet means are withheld unless every dataset is scored. A synthetic-data harness tests generated training rows without letting a generator see held-out rows. This report describes release 0.4.0 of this living benchmark.
cs.LG / 115 / 2610.03265
SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining
Dengdi Sun, Xiaoya Zhou, Xiao Wang, Wanli Lyu, Jin Tang, Bin Luo
cs.LG · cs.AI
Abstract
Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for large-scale PDE pre-training. SPEAR decouples latent features into low- and high-frequency components, enabling shared modeling of transferable dynamics and specialized learning of PDE-specific patterns. To address expert redundancy, we design a knowledge-guided expert aggregation strategy that measures expert similarity from dataset-specific learned knowledge and routing preferences, enabling the identification and consolidation of similar experts. Experiments on twelve PDE datasets and multiple downstream benchmarks demonstrate superior performance in pre-training, fine-tuning, and transfer learning. Furthermore, our aggregation strategy reduces the number of experts by 50\% while maintaining or improving prediction accuracy, achieving a balance between model efficiency and generalization for PDE foundation models.
cs.LG / 116 / 2610.03306
Training-Loss Guarantees for Muon with Finite-Step Newton--Schulz Orthogonalization
Amartya Roy, Souvik Chakraborty
cs.LG · cs.AI · math.OC
Abstract
Existing convergence analyses of Muon either assume exact orthogonalization or analyze classical Newton--Schulz polynomials, and guarantee only stationarity, so it is unresolved what Muon's five tuned Newton--Schulz steps preserve and whether that suffices to reach a prescribed neural-network training loss. We establish a finite-time training guarantee that accounts for both momentum accumulation before orthogonalization and the tuned finite-step update. For full-batch training of a sufficiently wide two-layer ReLU network with fixed random output weights and a positive-definite limiting neural tangent kernel, we prove that Muon reaches any target empirical squared loss $\varepsilon>0$ with high probability over initialization. For every momentum parameter $μ\in[0,1)$, a target-dependent constant learning rate proportional to $(1-μ)\sqrt{\varepsilon}$ yields a hitting-time bound of $O((1-μ)^{-1}\varepsilon^{-1/2})$, with other problem parameters fixed. The sufficient width is independent of both target accuracy and momentum. The analysis shows that the tuned Newton--Schulz map preserves alignment with the momentum buffer while bounding the update's spectral norm. Control of gradient variation near initialization transfers this alignment to the current gradient, ensuring descent until the target is reached without requiring exact orthogonalization. Numerical experiments support these mechanisms at widths below the sufficient theoretical threshold: gradient-update alignment remains above the analytical reference, and all 30 runs across six widths and five student initializations on a fixed teacher-student dataset reach the target loss while maintaining kernel positivity.
cs.LG / 117 / 2610.03330
Cordial Learning: Distributed Training with Correlated Data
Sarah Shitrit, Ilai Bistritz
cs.LG · cs.AI
Abstract
We consider a distributed learning task with agents that have correlated data. Specifically, the label of an agent depends on the input of other agents for the same sample, and these inputs are also correlated. Correlated data is the reality when agents share the same environment. Existing decentralized methods, such as federated learning, ignore the structure of the problem and perform poorly on correlated data. On the other hand, centralized approaches are infeasible due to privacy and communication constraints. We introduce cordial (correlated and distributed) learning to address this gap by sharing only low-dimensional outputs between the agents while training local models to extract informative signals from peers. This distributed learning induces a game in which the loss function of each agent depends on the models of others. Assuming a linear model, we prove that cordial learning converges with probability one to a globally optimal solution, despite the nonconvex global objective. Experiments on structured multi-digit MNIST tasks demonstrate that cordial learning remains highly effective even in highly nonlinear settings.
cs.LG / 118 / 2610.03372
SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents
Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo
cs.LG
Abstract
Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
cs.LG / 119 / 2610.03378
Operator-informed initialization for Fourier features physics-informed neural networks
Juan Molina, Paris Perdikaris, Mircea Petrache, Matías Courdurier, Francisco Sahli Costabal
cs.LG
Abstract
Physics-Informed Neural Networks (PINNs) typically exhibit spectral bias, where some frequencies of the target function converge more slowly than others. In this work, we analyze the training dynamics of Fourier Feature PINNs in the Neural Tangent Kernel regime to address this limitation. We derive an explicit evolution equation to estimate the residual error in the frequency domain, demonstrating that the convergence rate of specific frequencies is primarily governed by the product of the differential operator's symbol and the spectral density of the initialization weights. Leveraging this theoretical insight, we propose an informative initialization strategy that tailors the initial weight distribution to the specific PDE being solved. With this method, we can diminish the operator-induced spectral bias, balancing the convergence rates across the frequency spectrum and achieving better prediction accuracy. Numerical experiments on linear and nonlinear partial differential equations confirm that this initialization strategy improves learning dynamics and approximation accuracy across frequencies compared to standard initialization methods, with no additional training cost.
cs.LG / 120 / 2610.03395
Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning
Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter
cs.LG · cs.RO
Abstract
Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.
cs.LG / 121 / 2610.03402
16-bit Precision of Convolutional Neural Networks on Microcontroller Units for 8-bit Costs
Rui Liu, Benjamin Paaßen
cs.LG
Abstract
To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed approach achieves faster speed and lower energy consumption on layer- and model-level compared to alternative quantization schemes. We analyze the architecture of Armv7E-M, explain the underlying principles behind the performance advantages of 16-bit approaches, and evaluate the empiric quantization errors for regression and classification tasks, as well as empiric time- and energy consumption in MCU deployment. We observe ca.\ 10 times lower quantization errors compared to 8-bit quantization schemes while achieving similar or better inference times and energy consumption.
cs.LG / 122 / 2610.03413
AIBL: Augmented Instance-Based Learning with Structured Memory and Neural Embeddings
Radha Poovendran, Andrea Stocco, Linda Bushnell
cs.LG
Abstract
Sequential learning systems often make decisions from accumulated experience while receiving high-dimensional inputs whose distribution may change over time. Instance-Based Learning Theory (IBLT) provides a principled case-based framework for such settings through stored situation-decision-utility instances, partial matching, activation, and blending. IBLT relies on symbolic knowledge representation in dictionary-like formats, but text, images, transaction vectors, and user-item histories often require learned similarity rather than hand-specified matching rules. In this paper, we introduce AIBL (Augmented Instance-Based Learning), an instance-learning model formulated in a learned vector space for high- dimensional sequential data. AIBL generalizes symbolic situation matching to neural embedding similarity while retaining instance storage, activation- weighted retrieval, and utility blending. The AIBL model organizes memory into active, forgotten, and surprise stores. Surprise memory separates weakly matched, possible out-of-distribution, or corner-case observations from active memory, reducing forced fitting to the nearest available cases. An observation-driven graduation algorithm promotes recurring surprise instances to active memory, allowing the memory to incorporate repeated novel patterns that may arise under concept drift. We evaluate the same implementation on five machine learning tasks and three controlled simulation tasks, comparing AIBL with classical IBLT variants and task-specific baselines where appropriate. AIBL improves accuracy by 6 to 17 percentage points. The results show where vector-space retrieval improves over symbolic matching and how the added memory mechanisms govern novelty detection, cold-start handling, drift adaptation, and reward learning under the tested protocols.
cs.LG / 123 / 2610.03418
Rethinking Epistemic Uncertainty in Node Classification through Information Growth
Emma Meneghini, Francesco Ferrini, Bruno Lepri, Andrea Passerini, Veronica Lachi
cs.LG · cs.AI
Abstract
Epistemic uncertainty should decrease as additional information about the data-generating process (DGP) becomes available to the predictor. Yet, existing graph evidential deep learning (EDL) methods for node classification typically construct epistemic uncertainty from graph-specific properties and evaluate it on downstream tasks such as out-of-distribution detection, which do not test its reducibility as information about the DGP increases. To make reducibility directly testable, we introduce a statistical framework for studying epistemic uncertainty under information growth. Our framework specifies an information-growth experimental protocol and a consistency criterion for epistemic predictors, while using projective graph DGPs to ensure that growing graphs, which in general need not provide increasing information about the same DGP, constitute coherent observations of the same underlying process. We show that EDL methods do not explicitly estimate data uncertainty arising from a single finite graph observation and instead regulate epistemic uncertainty through model hyperparameters, precluding consistency, as corroborated by controlled information-growth experiments. As an alternative, we propose graph bootstrap ensembles, capturing both data and procedural uncertainty through graph resampling and randomized training. Under the same experimental protocol, these ensembles exhibit epistemic uncertainty reduction beyond standard deep ensembles. These findings support bootstrap ensembles as candidate consistent epistemic predictors under information growth.
cs.LG / 124 / 2610.03419
Deep Bayesian REFoCUS
Simon Penninga, Ruud van Sloun
cs.LG
Abstract
In this work we formulate ultrasound multistatic recovery from arbitrary transmit sequences as a Bayesian inference problem. To that end, we train a deep generative prior on multistatic data sets to tackle the rank-deficient regime in which classical linear REFoCUS decoders fail. This appproach, which we term Deep Bayesian REFoCUS, outperforms the linear baselines for all regimes of rank-deficiency and noise levels, and regresses to linear decoding when inversion is exact. The model also expresses uncertainty in the null space of the acquisitions, whereas the linear REFoCUS decoders only provide point estimates. Finally, we analyze the impact of distribution shift between simulation and in-vivo acquisitions, showing remarkable generalization ability without any fine-tuning or adaptation.
cs.LG / 125 / 2610.03432
OptiSelect: How does the Optimizer Shape Data Curriculum?
Simin Fan, Alireza Abdollahpoorrostam, Martin Jaggi
cs.LG · cs.AI
Abstract
Online data selection has demonstrated substantial efficiency gains for LLM pretraining by training on the most valuable candidates within each batch. Since a candidate's value is realized through its effective model update, principled selection should account for the optimizer step, which reshapes the raw gradient before it updates model parameters. We formalize this optimizer-aware selection paradigm as OptiSelect and present the first systematic study of how the optimizer shapes data selection. Our theory establishes a selection gain principle in which the advantage of online selection is governed by the discriminability of the optimizer-induced utility scores. We prove that sign-based and polar-tangential preconditioners of Lion and Muon would suffer from a discriminability collapse which caps attainable gains from OptiSelect, whereas diagonal-adaptive optimizers such as AdamW and Sophia admit strictly better upper bounds. The proposed principle also yields a quantitative derivation of the optimal candidate oversampling ratio. Pretraining experiments on 124M and 720M models are consistent with our theoretical analysis and show that AdamW's diagonal-adaptive scoring geometry remains the strongest scoring geometry even with Muon as optimizer. We further demonstrate that OptiSelect retains its benefits under data rephrasing, a technique used in modern data processing pipelines. Our findings provide theoretical foundations and practical guidance for co-designing optimizers and data selection in LLM pretraining.
cs.LG / 126 / 2610.03452
Causal Representation Learning with Instantaneous and Lagged Relations via Nonstationarity
Tatsuya Yamada, Hiroshi Morioka, Yoshinobu Kawahara
cs.LG
Abstract
Causal representation learning for time-series data aims to identify latent states and their causal relations from observations. In this setting, an important challenge is to model both lagged causal relations across observation intervals and faster causal effects that appear as instantaneous relations within an interval, while accounting for nonstationarity in time-series data. However, methods that jointly handle these causal relations and nonstationarity remain limited. To address this gap, we establish sufficient conditions for identifying latent states up to component permutation and component-wise invertible transformations, and their instantaneous and lagged causal structures up to the same permutation, using an observed auxiliary variable, such as time or a condition label, associated with changes in transition-noise distributions. Based on these results, we propose iCReN, a framework that uses contrastive learning with discrete or continuous auxiliary variables to learn latent representations and estimate their instantaneous and lagged causal structures. Experiments demonstrate accurate recovery of latent states and both instantaneous and lagged causal structures on synthetic data and the utility of the learned representations for downstream forecasting on real-world data.
cs.LG / 127 / 2610.03454
Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne
cs.LG · cs.AI
Abstract
Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO-$k$, a task-agnostic and self-supervised learning principle for modality acquisition: we use a deep model's internal pretrained representations (e.g., from a foundation model) as proxy targets that summarize cross-modal information. We provide theoretical guarantees in a stylized linear setting that motivate a reinforcement learning (RL) policy for sequential modality selection. Across task-agnostic and label-free acquisition baselines, ECHO-$k$ consistently improves budgeted downstream performance across diverse foundation-model backends. Our method provides a principled route to cost-aware test-time deployment, with implications for any multimodal system where measurements are expensive or time-constrained, and downstream tasks unknown a priori.
cs.LG / 128 / 2610.03456
Beyond Random Splits: Evaluating Drug-Target Affinity Models Under Chemically and Biologically Motivated Distribution Shifts Copy
Minjae Chung, Clara Li, Malar Paavai Muthukumaran, Shaunna Wang, Aniket Ramkrishnan Iyer, Shaun Qien Yeau Tan, Harinishree Sathu, Micky C. Nnamdi, J. Ben Tamo, Benoit Louis Marteau, May Dongmei Wang
cs.LG
Abstract
Drug-target affinity (DTA) prediction is widely used to prioritize candidate compounds before costly experimental screening. DTA models are often compared under a single data split, even though deployment may require extrapolation to new chemical series, new protein targets, or both. We ask whether the distribution shift used for evaluation changes which architecture appears best. We curate 718,800 unique drug-protein pairs from the ChEMBL and BindingDB datasets. We compare a Morgan-fingerprint + protein-CNN baseline with 12 controlled architectures that combine four drug representations with three ESM-2 interaction modes. Mean validation RMSE increases from 0.950 and 0.945 under scaffold and fingerprint-cluster OOD to 1.299 and 1.321 under protein-cluster and dual OOD. Model rankings are similar across the two chemical shifts (tau = 0.79), but agreement with scaffold OOD falls under protein OOD (tau = 0.39) and reverses under dual OOD (tau = -0.55). Held-out evaluation, repeated seeds, group-aware bootstrap analysis, and a size-matched control support the same conclusion: architecture selection depends on the form of extrapolation, not only on average error or training-set size. DTA benchmarks should therefore match the chemical and target shifts expected at deployment.
cs.LG / 129 / 2610.03475
Single or Multiple Policies for Phase-Structured Reinforcement Learning?
Guilhem Loussouarn, Nancy Nayak, Kin K. Leung
cs.LG · cs.AI · eess.SY
Abstract
Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in practice depends on function approximation, learning and optimization processes, as well as, for multi-policy solutions, the sample efficiency and loss of continuity from one policy to another. We propose a regime-based phase decomposition method to identify which policy can provide better performance. The method is based on consideration of the duration of transient system dynamics relative to the duration of the quasi-stationary period. Numerical experiments are conducted with different non-stationary RL problems to validate our four major hypotheses: (a) longer phase durations favor multi-policies, (b) the heterogeneity between phases increases the burden on single policy, (c) multi-policies need sufficient data for each phase, and (d) environment-specific transition dynamics between phases can affect which policy is preferable.
cs.LG / 130 / 2610.03491
Dual-Context Analog Retrieval for Time Series Forecasting
Jung Min Choi, Ngoc Son Le, Ibram Abdelmalak, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme
cs.LG
Abstract
Most long-term time-series forecasting models map the look-back window directly to the full horizon in a single pass. While efficient, this design does not explicitly identify which historical states are most relevant to different future segments or exploit what followed those states. Analog forecasting addresses this by retrieving past states similar to the present and using their observed continuations, but single nearest matches can be unreliable and overlapping patches may produce redundant candidates. We propose DuoTS, a Dual-Context Time Series forecasting model that uses retrieved evidence without relying on it exclusively. DuoTS first produces a base forecast with a parallel patch encoder and linear prediction head, then progressively refines it one future patch at a time. Each refinement combines two views: a current context that attends to recent tokens and captures the latest dynamics, and a detail context that provides distinct retrieved analogs together with their subsequent trajectories. Patch-wise refinement allows the model to balance these views across the forecast horizon and associate each future segment with evidence appropriate to its temporal distance from the present. Experiments on multiple real-world datasets show that DuoTS achieves state-of-the-art performance, while ablations confirm the contribution of each context. The refinement mechanism is also model-agnostic, requiring only an encoded look-back window and the future-patch position, and can therefore be integrated into existing forecasting models.
cs.LG / 131 / 2610.03494
Most-Recent Anchoring with Recurrent Ordering for Time Series Forecasting
Jung Min Choi, Ngoc Son Le, Ibram Abdelmalak, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme
cs.LG
Abstract
Long-term forecasting models commonly process all patches in a look-back window using the same fixed stack. Older contextual patches and recent evidence therefore receive the same computational depth. Yet the information closest to the forecast and the more distant context do not contribute equally. Uniform processing leaves this distinction unexpressed in the architecture. We propose MARO, a Most-Recent Anchoring with Recurrent Ordering model that processes the look-back window from the most recent patch to the oldest. The most recent patch serves as the anchor. It initializes the latent state and conditions each subsequent step, so older patches are folded into a representation that remains centered on recent evidence. A single shared module is reused at every step, so extending the scan further into the past introduces no additional parameters. Intermediate states retained during the scan allow the forecast head to weigh short and long portions of the history separately. This expresses recency through the order of recurrent refinement. Extensive experiments across multiple real-world time series datasets show that MARO achieves state-of-the-art performance on both long-term and short-term forecasting tasks.Ablation studies examine the contribution of the main architectural components.
cs.LG / 132 / 2610.03500
Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets
Shivam Shrivastava
cs.LG · stat.ML
Abstract
Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds, plus 4 natively small clinical datasets at true size, for 2,220 committed runs in total. The primary metric is the AUROC of fixed classifiers trained on synthetic and tested on real data. In 23 of 24 (dataset, deep model) pairs no deep model ever beats the best trivial baseline by more than seed noise, at any training size we measured. The best baseline wins 40 of 49 (dataset, size) cells. Our preregistered prediction that the deep models' ranking would be unstable at small sizes is falsified: mean Kendall tau between adjacent rungs below 5,000 rows is 0.806, above our 0.8 threshold, and stability is highest at the smallest sizes rather than lowest. One caveat bounds all of this: in 81% of cells the gap between the top two methods is smaller than the variation between seeds. Finally, method rankings on natively small clinical datasets agree only moderately with rankings on subsampled large ones (mean tau 0.57 to 0.64), which questions whether a subsampled large dataset can stand in for a small one. All 2,220 result files, the preregistration and its hash, and the code that regenerates every figure and number from those files are public.
cs.LG / 133 / 2610.03502
Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation
Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha
cs.LG · cs.AI · cs.CR
Abstract
Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interpretability-verification boundary certifies descriptions of a model: what a circuit computes, or whether it faithfully explains the whole. We instead certify the behavioral effect of an edit: that disabling a circuit removes one skill and provably preserves another, for every input in a region; a feature non-interference guarantee in the information-flow-security sense. We demonstrate such certified edits from toy ReLU networks up to a standard softmax + LayerNorm transformer, proving removal and preservation over continuous embedding-space regions and reaching roughly 9x the input-perturbation dimension an exact solver can handle by switching to sound bound propagation. Furthermore, we prove that no finite deterministic black-box test can certify removal, exhibiting an edit that passes exhaustive testing yet provably fails on a survivor pocket that can be made arbitrarily small. Guarantees hold on small, standard-architecture networks and, like any removal claim, presuppose that the target skill admits a decidable specification, a property which real-world harms may not have.
cs.LG / 134 / 2610.03505
An Automated and Reproducible Workflow for Crack Identification and Damage Assessment of Fusion Materials
Rinkle Juneja, Viktor Reshniak, Richard K. Archibald, John W. Duggan, Gregory R. Watson, Cory D. Hauck, Gary M. Staebler
cs.LG · cond-mat.mtrl-sci
Abstract
Post-exposure microscopy is central to qualification of fusion materials. However, manual analysis does not scale to the volume, heterogeneity, and multiresolution character of modern fusion-materials campaigns. To address this challenge, we present a reproducible workflow, implemented in the Galaxy scientific workflow environment, for automated crack identification and quantitative damage assessment from scanning electron microscopy images. The workflow processes SEM images and experimental metadata to identify cracks, quantify damage, and retain the intermediate products and processing history needed for reproducibility. Outputs include crack masks, skeletonized crack networks, quality-control visualizations, and scalar damage descriptors. The method is designed to operate without image-specific parameter tuning across tungsten grades, microstructures, magnifications, and damage states. We demonstrate the workflow on a sparse electron-beam thermal-shock dataset containing 418 images from 114 experiments spanning five tungsten grades and three microstructural states. We define a crack-density descriptor, which provides standardized inputs for downstream machine-learning prediction and physics-based crack simulation. These predictive components are exposed in the same Galaxy environment and are intentionally treated here as extensible workflow modules. The principal contribution is therefore an end-to-end, shareable, and computationally portable workflow that links experimental characterization, automated image analysis, preliminary damage prediction, and simulation-guided data acquisition for fusion-materials research.
cs.LG / 135 / 2610.03526
Beyond Trained Models: Compiling GNNs for a Sound Explainer Benchmark
Steve Azzolin, Francesco Paolo Nerini, Stefano Teso, Francesco Bonchi, Bruno Lepri, André Panisson, Andrea Passerini
cs.LG · cs.AI
Abstract
Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data. This protocol implicitly assumes that a GNN trained on such data relies on the intended motif. Although prior work has questioned this assumption, plausibility remains widespread. First, we show that the assumption is violated on several widely used benchmarks, where, e.g., degree statistics alone suffice to solve the task. Then, we remove this confounder by replacing training with compilation. We achieve this by introducing $\mathsf{Gracr}$, the first compiler translating graded modal logic formulas into GNN weights, yielding models that replicate the behaviour of the corresponding formulas. Since the behaviour of the model is now known by construction, we can define its ground truth explanation formally and compute it exactly. Building on this, we introduce $\mathsf{Gracr}\mathsf{Bench}$, a benchmark of compiled GNNs for the evaluation of explainers against this exact ground truth. Experiments on eleven explainers across six tasks show its effectiveness for fine-grained diagnostic evaluation: notably, we discover that most explainers are not robust to indirect influences or alternative implementations of the same formula. These results position $\mathsf{Gracr}\mathsf{Bench}$ as a novel, rigorous evaluation setting for graph post-hoc explainability.
cs.LG / 136 / 2610.03546
ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models
Yubo Wang, Jingying Ma, Xinliang Zhou, Yangxuan Zhou, Jiquan Wang, Sha Zhao, Yiyuan Yang, Yi Ding, Ziyu Jia, Chenyu Liu, Cuntai Guan
cs.LG
Abstract
EEG foundation models (EFMs) capture reusable knowledge from large-scale EEG data, while many EEG recordings also include companion physiological signals that provide complementary information beyond the EEG-only interface. The challenge is to preserve this pretrained knowledge while extending the EFM to heterogeneous multimodal recordings through an adaptation inferred from unlabeled target data. We introduce ZeroMAG, a zero-shot multimodal adapter generation framework that extends a frozen EEG encoder and prediction head using unlabeled target recordings, without target labels or target-side optimization. The target datasets are held out from all model training and selection in the ZeroMAG pipeline. ZeroMAG organizes companion modalities around a configuration-invariant adapter, constructs a modality-subject-task condition from unlabeled recordings and task context, and generates adapter weights in a function-constrained latent space learned from source adapters. Across six held-out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG-only inference and 4.89 points over direct weight regression, while coming within 0.50 points of supervised multimodal adaptation on average. Ablations further show that removing functional supervision from either representation learning or conditional generation degrades generated-adapter performance, confirming the contribution of both components.
cs.LG / 137 / 2610.03556
Get a GRIP, this will be a long TRIP: A Quantifiable Long-Range Framework for Verifying Over-squashing
Ferran Hernandez Caralt, Simon Heilig, Adrián Bazaga, Asja Fischer, Moshe Eliasof, Pietro Liò
cs.LG
Abstract
Empirical claims about the connection between over-squashing and long-range interactions in GNNs, can only be trusted if the benchmarks used to validate them genuinely require long-range interactions. The de-facto standard, the Long Range Graph Benchmark, has been repeatedly shown to be saturated by tuned short-range models, with existing synthetic alternatives being tied to specific topologies. As such, there is a lack of principled certificate of long-rangedness on arbitrary graphs. This state reflects the absence of a precise characterization of long-ranged benchmarks. We address this fundamental gap by introducing four verifiable axioms: Predictability, Tightness, Strictly $k$-Range, and Topology-Invariance, that any task claiming to test $k$-hop interactions must satisfy. We formally prove that violating any one of them admits failure modes that undermine conclusions drawn from the task. Based on these axioms, we introduce TRIP (Truly Ranged Interactions Problem) and its generalisation GRIP (Generally Ranged Interactions Problem), constructive procedures that turn any graph into a provably long-ranged task by drawing features from stable distributions. Moreover, by construction, GRIP admits a closed-form, per-range Maximum-Likelihood oracle that yields the first a priori per-range lower bound on test error available on any benchmark. Using our framework, we: (i) audit 4 common long-range benchmarks and identify their failures modes with respect to our axioms; (ii) on TRIP-instantiated topologies, we find a popular notion of curvature is uncorrelated with GNN performance, supporting topological-vs-computational bottleneck distinction; and (iii) we show that a novel benchmark's over-squashing measures factors beyond pure long-rangedness. Code to use the framework and reproduce experiments is released https://github.com/ferranhernandezc/graph-grip.
cs.LG / 138 / 2610.03558
Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain
Antoine Collas, Louis Jalouzot, Géraud Ilinca, Corentin Caris, Romain Valabrègue, Ahmed Hassayoune, David Goncalves, Madeleine Hueber, Thaddée Delebarre, Julien Savatovsky, Clara Fonteneau, Charles Maussion, Bertrand Thirion, Alexis Thual
cs.LG · cs.AI
Abstract
Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. With 30 hours of fMRI data per subject, the current release is the deepest available fMRI dataset using naturalistic speech stimuli. The dataset pairs brain activity with the corresponding podcast audio, transcript annotations, and derived stimulus embeddings. Furthermore, we introduce a brain decoding benchmark formulated as audio segment retrieval: given fMRI activity from a held-out session, the decoder must identify the corresponding time-aligned podcast audio segment among candidate segments. We provide standardized splits, evaluation metrics, and baseline decoders for this task. Finally, a scaling analysis shows that decoding performance improves continuously with the amount of training data per subject.
cs.LG / 139 / 2610.03597
A Path Integral Surrogate for Multi-Step Gradient Inversion in Federated Learning
Agnivo Ghosh, Saumik Bhattacharya
cs.LG · cs.DC
Abstract
Federated learning lets many clients train a shared model together without ever sending their private data to a central server. Each client shares only a model update, and this update should reveal far less about the client than its raw training examples would. This premise is what protects the privacy of the clients. Gradient inversion attacks challenge it directly by trying to reconstruct a client's private input images from the single update it shared. Under FedAvg, a client's update accumulates several local training steps, so the server sees only the two endpoints of a hidden weight trajectory. Recent gradient inversion attacks fit a surrogate model along the path between these two endpoints but they still read its gradient at a single point. We propose the Path-Integral Surrogate Model Extension (PI-SME) which treats the accumulated update as a path integral of the gradient field and approximates it by Gauss--Legendre quadrature over several nodes along a learnable Bézier path. On CIFAR-100 and FEMNIST images across a range of trajectory lengths and class-restricted batches PI-SME reconstructs the private inputs more faithfully than the strongest surrogate baseline on several inversion metrics and the matching loss.
cs.LG / 140 / 2610.03604
Mastering Atari 2600 Games with Discovered Options
Erik M. Lintunen, Marlos C. Machado
cs.LG
Abstract
Temporal abstractions, often instantiated as options, have long been regarded as a mechanism for accelerating credit assignment, facilitating exploration, and enabling generalisation in reinforcement learning (RL). However, developing general option discovery methods that are effective in large-scale, high-dimensional domains remains a fundamental challenge. Existing option discovery methods are either confined to relatively simple domains, depend on handcrafted or quasi-symbolic representations, or offer little improvement over learning without options. We present Wayfarer, a general, domain-agnostic, online deep RL agent that discovers options through Laplacian representation learning from high-dimensional observations and leverages them for control. We show that the resulting options simultaneously improve exploration, accelerate credit assignment, and generalise effectively to unseen settings, enabling substantially faster learning of complex policies. Wayfarer achieves state-of-the-art performance among single-stream agents on the most challenging Atari 2600 games, with the largest gains in games that require long-horizon exploration and strategic behaviour, such as Montezuma's Revenge and Private Eye.
cs.LG / 141 / 2610.03620
UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning
Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang
cs.LG · cs.RO
Abstract
Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and heterogeneous assisted behaviors during online RL. Specifically, UniIntervene++ first formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a unified semi-Markov decision process and learns their relative values online. Building on this, competence-adaptive intervention periodically probes the RL policy through unassisted execution, keeping control allocation responsive to its evolving capability. Finally, coupled experience learning allows assisted behaviors to improve the RL policy, whose evolving outcomes in turn reshape future intervention decisions. In this way, UniIntervene++ jointly determines when to intervene, how to intervene, and when to return control as the RL policy improves. Across five real-world manipulation tasks, UniIntervene++ achieves an average success rate of 89.67%, outperforming all baselines by at least 6 percentage points, while reducing human intervention to 0.77%, a relative reduction of at least 94.6% from the best baseline. Code is available in our \href{https://github.com/dannyyudong/An-Adaptive-Intervention-Agent-for-Efficient-Real-World-Reinforcement-Learning}{GitHub repository}.
cs.LG / 142 / 2610.03640
Broken scale symmetries in undercomplete linear autoencoders
Farhad Pashakhanloo, Jacob A. Zavatone-Veth
cs.LG · q-bio.NC · stat.ML
Abstract
Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next without changing the network output. Previous work has documented cases in which SGD breaks this symmetry in favor of balancing gradient noise or minimizing fluctuations. Here, we show that the solution geometry of undercomplete linear autoencoders instead selects a preferred sign for scale drift: on the PCA solution manifold, SGD favors large decoder weights. This directed scale drift occurs on a slow timescale, and its dynamics admit an analytically-tractable effective description. However, it cannot continue indefinitely: increasing scale eventually drives the dynamics towards a finite-stepsize stability boundary. The resulting solutions are sharper than a balanced baseline in the sense of the maximum eigenvalue of the loss Hessian, but different sharpness measures can move in opposing directions. Thus, undercomplete autoencoders give a concrete illustration of how loss geometry can convert residual gradient noise into directed motion along a manifold of functionally-equivalent solutions.
cs.LG / 143 / 2610.03642
On the Convergence of Success Conditioning for Policy Optimization
Matthew Brun, Xu Andy Sun
cs.LG · math.OC
Abstract
Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within $\mathcal{O}(1/\varepsilon^p)$ iterations to an $\varepsilon$-optimal policy, where the exponent $p$ depends on problem data. For single-period MDPs, such a policy is obtained within $\mathcal{O}(\log(1/\varepsilon))$ iterations.
cs.LG / 144 / 2610.03646
When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity
Mayand Gulati, Kerong Wang, WeiChen Au
cs.LG · stat.ML
Abstract
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen $α_E$, without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by $38.4\%$ on the registered suite but reduced it by $7.5\%$ on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.
cs.LG / 145 / 2610.03662
Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation
Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar, Dean Foster, Carson Eisenach
cs.LG
Abstract
Deploying a new decision policy creates a cold-start problem for prediction models whose targets depend on the policy's actions: historical observations reflect earlier policies, while real observations under the new policy are not yet available. Simulation offers a way to address this gap by rolling out the target policy across counterfactual scenarios and using the resulting trajectories to learn how the system responds to those controls. The simulation-to-reality (Sim2Real) transfer of this simulator-trained model can then be backtested by evaluating it against real observations from past deployments. Using two real-world inventory-control deployments, we evaluate this process from three angles: simulator fidelity, zero-shot transfer to real behavior, and adaptation as real target-policy observations accumulate. The simulator-trained forecaster achieves lower point-estimate mean absolute percentage error (MAPE) than the same architecture trained on historical real data, reducing MAPE by 1.2-3.1 percentage points in Study 1 and 12.5-18.7 points in Study 2. After deployment, lightweight calibration using early real observations further reduces error by up to 2.5 percentage points. These results provide empirical evidence that simulator-generated counterfactual data can support cold-start forecasting under a new policy, and the resulting model can be further refined as real deployment data become available.
cs.LG / 146 / 2610.03667
Planning to Learn
Ian Osband
cs.LG · math.OC · stat.ML
Abstract
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its log-odds rose at unit speed forever, while exact policy gradient is the zero-horizon limit. Truncating this total at the learning that remains yields the horizon loss, a one-line change that moves from cross-entropy toward exact policy gradient as training runs out. In a simple allocation model, it provably escapes the trap that catches each endpoint. On MNIST and on ImageNet with ResNet-50, ResNet-101 and ViT-S/16, the horizon loss improves top-1 accuracy over cross-entropy at a flat learning rate, and the gain grows with label noise.
cs.LG / 147 / 2610.03679
Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals
Fedor Sergeev, Markus Heinonen, Daniel Waxman, Tim Cooijmans, Ricardo Baptista, Dmitry Batenkov, Eli Bingham
cs.LG · stat.ME · stat.ML
Abstract
The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the underlying process. Wasserstein gradient flows are a common choice, but they cannot describe conservative or periodic dynamics. Lagrangian mechanics in Wasserstein space covers both, but existing methods for learning it are simulation-based: they run a numerical solver at every training step, which makes training expensive. We propose Double-Stitch, a simulation-free method that learns these mechanics by penalizing the residual of the equation of motion along a learned population path. We derive this equation from a Clebsch variational principle that does not require gradient velocities, and show that the residual vanishes exactly when the equation holds. We test Double-Stitch on synthetic, single-cell and ocean vortex datasets and find that it matches or outperforms gradient-flow methods and simulation-based WLM on most tasks, while training $4$-$14$ times faster than WLM. We provide a JAX implementation of Double-Stitch at https://github.com/BasisResearch/stitching.
cs.LG / 148 / 2610.03712
RNADyn: A Benchmark for Generating and Understanding RNA Dynamics
Yiming Huang, Lennart Bastian, Hanqun Cao, Luis Vollmers, Tolga Birdal
cs.LG · physics.bio-ph · q-bio.BM
Abstract
Ribonucleic acid (RNA) functions through conformational changes that are not fully captured by static structures. However, large-scale standardized RNA dynamics data remain limited, and existing approaches typically treat trajectory generation and dynamics understanding as separate objectives. Here, we introduce RNADynBench, a standardized RNA molecular dynamics (MD) benchmark with 2585 quality-controlled 100-ns all-atom trajectories and leakage-controlled splits. Building on RNADynBench, we develop RNADynNet, a unified model for RNA dynamics learning that uses a shared backbone for both trajectory generation and dynamics fingerprint extraction from a single conformer. It combines coordinate denoising, single-frame-to-trajectory alignment, and physical grounding to connect all-atom trajectory generation with dynamics representation learning. Physical grounding improves both generated dynamics and the physical information recoverable from these fingerprints. Across both test sets, including the high-flexibility challenge set, the generated trajectories achieve RMSF correlations of 0.875 and 0.766, while single-conformer predictions show comparable agreement with MD-derived dynamics. RNADynBench and RNADynNet together establish a benchmark and unified modeling framework for generating and understanding RNA dynamics.
cs.LG / 149 / 2610.03713
What Should World Models Forget? Stratified Retention for Continual Adaptation
Nishit Anand, Ramani Duraiswami, Dinesh Manocha
cs.LG · cs.AI · cs.CV · eess.IV · eess.SP
Abstract
Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, which changes, so knowledge that was accurate when acquired may later become false, and discarding it is required behavior rather than a defect. Non-stationary ground truth is well studied in the concept drift literature and in the temporal factuality of language models, but has not been formulated for world models, which are distinctive in that they also encode knowledge that must never be revised. We argue that continual world models require retention stratified by invariance timescale, separating invariants such as physics and object permanence, which must never be revised, from instance-level facts that should be revised as soon as the environment changes. Standard forgetting metrics cannot distinguish a world model that has correctly revised outdated knowledge from one that has suffered catastrophic forgetting, and consequently rank a frozen model highest, while existing physical-reasoning benchmarks evaluate only frozen checkpoints. We propose differential retention, which reports invariant regression testing across the adaptation stream jointly with revision latency, without aggregation.
cs.LG / 150 / 2610.02403
Energy Saving in 5G and Beyond Networks: A Quantum Reinforcement Learning Approach
Muhammad Usman, Nguyen Van Huynh, Marianna Lezzi, Mariangela Lazoi
cs.NI · cs.LG
Abstract
Energy saving has become a critical challenge in 5G and beyond networks. The rapid growth of connected devices has increased the overall network energy demand, driving operational expenditure to unsustainable heights. The Base Station (BS) accounts for the largest share of energy usage, typically consuming around 60-70\% of the Radio Access Network (RAN)'s total energy. Therefore, to address this issue, this article optimizes the BS's energy usage while accounting for the dynamic behavior of User Equipment (UE). Deep Reinforcement Learning (DRL) is a natural candidate for determining effective energy saving policies, such as automatically switching BSs on or off when user density is low or adjusting transmission power to balance energy efficiency and Quality of Service (QoS). However, its heavy training burden and the exponential growth of state and action spaces in dense 5G environments make exploration increasingly difficult. To overcome these limitations, we introduce a novel Quantum Reinforcement Learning (QRL) algorithm that leverages quantum principles, including superposition and entanglement, through parameterized quantum circuits, enabling significantly faster convergence than DRL, which relies on conventional deep neural networks. Extensive simulations demonstrate that the proposed QRL can substantially reduce energy consumption while maintaining QoS, even when UEs are highly dynamic and frequently switch their association with BS antennas. Additionally, QRL consistently outperforms DRL and Q-Learning in both convergence speed and learning complexity.
cs.LG / 151 / 2610.02338
SoTa: Soft Tactile Skins for Dexterous Manipulation
Jingyun Yang, Baiyu Shi, Timothy Yu, Haitian Liu, Alberta Longhini, Weichen Wang, Rika Antonova, Zhenan Bao, Jeannette Bohg
cs.RO · cs.LG
Abstract
A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. However, visuo-tactile robot data remains scarce: dexterous demonstrations require teleoperating robots, which limits dataset scale. Human demonstrations are far cheaper to collect and offer a path to scale this data, but only if human and robot hands carry tactile sensors with corresponding signals. This requires sensors that conform to different hand geometries, cover the full hand, and share a common layout across embodiments. We present SoTa, a low-cost capacitive tactile skin that provides full-hand coverage on humans and robots while preserving a shared layout of 202 taxels across corresponding finger and palm regions. Our multilayer design with fabric electrodes enables in-house fabrication of thin, soft skins with customizable geometry for under $10 in materials per skin. The sensor retains over 97% of its initial response span after 10,000 loading-unloading cycles with traces retaining continuity through 1,280 tight-fist folding cycles. The shared taxel layout supports human-robot co-training with a common tactile encoder and no learned cross-sensor mapping. Across three contact-rich manipulation tasks, tactile observations improve in-distribution success over vision-only policies. With a fixed robot demonstration budget, adding human demonstrations more than doubles mean success across eight evaluation conditions, from 22.8% to 45.9%, improving success in all five out-of-distribution conditions. We plan to open-source the resources needed to fabricate and operate these skins.
cs.LG / 152 / 2610.02339
NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches
Juntao Ren, Yifan Hou, Shuran Song
cs.RO · cs.LG
Abstract
Robot demonstrations may contain useful behavior even when individual episodes are inefficient or unsuccessful. Trajectory stitching offers a way to compose these behaviors into improved training data, but identifying useful connections and verifying their feasibility is difficult in high-dimensional robot data, where many prior methods rely on low-dimensional state representations. We introduce NEEDLE, an offline dataset-augmentation algorithm that addresses these challenges by adding short, verified action bridges between recorded observations in high-dimensional robot demonstrations. First, NEEDLE identifies and creates connections that bypass suboptimal detours, broaden action coverage, and augment the original dataset with failed trajectories, using only RGB images, proprioception, and episode-level outcomes, without new environment interaction or privileged object state. Next, we present a sampling technique that incorporates accepted bridges into policy training without synthesizing intermediate images or discarding the original demonstrations, allowing policies to learn alternative actions while retaining the original dataset's coverage. On real-robot tasks, NEEDLE improves success rate over the strongest baseline on each task by an average of 21 percentage points. Videos and supplementary materials are on https://needle-work.github.io/.
cs.LG / 153 / 2610.02765
Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving
Luís Marques, Rong Fang, Disha Kamale, Dmitry Berenson
cs.RO · cs.LG
Abstract
Monitoring planned driving trajectories requires accurately estimating the collision likelihood with actors whose motion is itself impacted by the ego motion. Existing classical approaches are often limited by the quality of their forecasting model. Vision-language models (VLMs) have shown promise in reasoning about the consequences of high-level actions, yet their approximate predictions are unsuitable for safety-critical applications such as autonomous driving. Conformal prediction (CP) has emerged as a data-driven framework for quantifying the uncertainty of black-box model predictions. We propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen VLMs that transforms their unreliable predictions into probabilistically calibrated safety prediction sets. We consider how the ability to estimate safety can depend on the observed driving scene and introduce a localized procedure that upweights relevant past experience when calculating uncertainty thresholds. We provide label-conditional finite-sample distribution-free coverage under exchangeability. Evaluated over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, while the base VLMs only flagged 4.6% and 39.1% of the collision-causing trajectories, respectively. These results indicate that local, label-conditional calibration can reduce missed unsafe trajectories.
cs.LG / 154 / 2610.02848
Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies
Amit Thakur, Mukesh Singhal
cs.RO · cs.LG · cs.MA
Abstract
Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive, while strong equivariance regularization can still induce homogeneous behavior. A weak equivariance penalty improves the robustness while preserving more diverse actions for teams with \(N=3\) agents, whereas teams with \(N=4\) agents require substantially smaller regularization weights. These findings suggest that multi-agent transformer policies should be evaluated not only by return and permutation robustness, but also by whether they maintain non-collapsed, differentiated multi-agent behavior.
cs.LG / 155 / 2610.03537
Autonomous Robotic Navigation for Endovascular Brain-Computer Interface Access
Harry Robertshaw, Weijie Qi, Nikola Fischer, Alejandro Granados, Thomas C. Booth, Sam E. John
cs.RO · cs.LG
Abstract
Endovascular brain-computer interfaces (BCIs) avoid craniotomy but require precise device delivery through anatomically variable cerebral veins. This work presents the first demonstration of in vitro autonomous robotic navigation for endovascular BCI access in the cerebral venous system. Soft Actor-Critic controllers were trained in silico for two sequential tasks spanning the right internal jugular vein to the superior sagittal sinus, using geometric augmentation of one training anatomy. Navigation was evaluated in a training anatomy and an anatomically unseen hold-out model over 250 in silico episodes and five fluoroscopy-guided in vitro robotic runs per task-anatomy condition, comprising 1,000 simulated episodes and 20 physical runs overall. Task recurrent predictors were also evaluated for online identification of impending navigation failure. In silico success rates for Tasks A and B were 85.6% and 98.4% in the training anatomy and 42.0% and 91.6% in the hold-out anatomy, respectively. Fourteen of 20 physical runs were successful (70% overall), including 80% success for Task B in the hold-out phantom. In silico the predictors detected 99.3-100.0% of failures with false-alarm rates of 0.8-6.7%. During in vitro evaluation, predicted risk increased before failed episodes, but elevated probabilities during some successful runs showed reduced calibration after transfer. These results demonstrate the feasibility of autonomous cerebral venous access and show how online failure prediction could support human oversight, while also identifying anatomical generalization and sim-to-real calibration as priorities before preclinical translation.
cs.LG / 156 / 2610.02918
Learning Jazz Pianist Style with Cross-Attention Conditioning
Drew Edwards, Akira Maezawa, Simon Dixon
cs.SD · cs.LG · eess.AS
Abstract
Jazz pianists develop distinctive traits that experienced listeners can often identify within seconds, yet the features underlying this recognition resist formal description. We study jazz pianist style through the lens of a pretrained symbolic music transformer, showing that its learned representations already encode pianist identity well enough for highly accurate classification across two benchmarks. We then augment the transformer with cross-attention over learned pianist identity embeddings, enabling it to generate music conditioned on a specific artist's style. Two evaluation protocols confirm that the generator captures meaningful stylistic structure: a sliding-window classifier consistently attributes conditioned continuations to the correct artist, far above unconditioned baselines; and a classifier trained entirely on synthetic generations identifies real pianists across 12 classes with 87% chunk-level and 95% song-level accuracy. Finally, we repurpose the classifier to locate the most characteristic moments within a performance, surfacing the specific musical gestures that distinguish each pianist's voice.
cs.LG / 157 / 2610.03125
ParaGeo: Decomposing Paralinguistic Variation into a Shared Latent Geometry
Yuhan Liu, Yuxuan Ou, Ruoxi Su, Mohamed Ahmed Zaki, Yunbo Long
cs.SD · cs.LG
Abstract
Speech delivery varies with both the requested paralinguistic attribute and the linguistic content. We introduce ParaGeo, a matched-content decomposition of paralinguistic variation in a frozen speech language model. Synthesized audio tokens are replayed with a fixed listening prompt; pooled key/value (K/V) representations are centered and projected into a shared low-dimensional space. Our GLM-4-Voice probe spans 80 requested controls from 12 benchmark families across eight sentences. With a globally fitted calibration basis, content-held-out centroid accuracy using this basis is 9.49% versus a 1.25% permutation baseline; same-label cross-content cosine similarity is 0.285 versus 0.017, and both conditional permutation tests yield p = 0.001. A separate ten-scenario, six-style probe reveals reproducible contrast directions across scenarios. Static, additive, and temporal interventions produce attribute-, layer-, and schedule-dependent response profiles. These results provide a shared coordinate representation for measuring paralinguistic structure and an empirical starting point for latent speech control. Code is available at https://github.com/yuhanlydia/ParaGeo.
cs.LG / 158 / 2610.02941
FASTDIAR: Frame-level speaker encoder for Streaming Diarization
Nikita Torgashov, Okan Köpüklü
eess.AS · cs.LG · cs.SD
Abstract
Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.
cs.LG / 159 / 2610.03290
Wrong Organ, Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening
Christiaan M. Geldenhuys, Joshua M. Jansen van Vüren, Véronique Suttels, Trevor Brokowski, Ablo P. Wachinou, Mary-Anne Hartley, Rensu P. Theart, Grant Theron, Thomas R. Niesler
eess.IV · cs.CV · cs.LG
Abstract
Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on that high-resource ultrasound domain carries representations that remain usable in the low-resource one. Only the encoder varies, across seventeen encoders spanning three architecture families. Among them, a latent-predictive video encoder pretrained on generic video (V-JEPA2-L) and its echocardiography counterpart (EchoJEPA-L) differ in pretraining corpus alone. The choice among these encoders does not resolve the classification, the whole family spanning 2.50 percentage points against a measurement resolution of 2.71. What moves the task instead is feature conditioning. Standardising the features between the encoder and the classifier improves all seventeen encoders by a mean of +1.23 percentage points at $p=1.5\times10^{-5}$. On the held-out test set every encoder selected on the development folds stands above the baseline system by up to +2.57 percentage points of area under the receiver operating characteristic curve (AUROC), and specificity at 90% sensitivity reaches 79.3% against 60.3%. The contrast specified in advance, EchoJEPA-L against V-JEPA2-L, measures -0.16 percentage points at $p=0.926$. We therefore find no evidence that shared ultrasonic physics alone makes echocardiography a more productive pretraining corpus than generic video, and any advantage, if present, is smaller than this cohort can resolve. The video encoders receive replicated still images, however, so whether this absence of an effect reflects the pretraining domain or a video encoder applied to static frames cannot be separated. The limiting factor is the labelled cohort rather than the encoder.
cs.LG / 160 / 2610.02466
SD-DPC: Sparse Dictionary Differentiable Predictive Control
Ali Reza Daneshvar Garmroodi, Jan Drgoňa
eess.SY · cs.LG · math.OC
Abstract
We present sparse dictionary differentiable predictive control (SD-DPC), a framework for learning sparse, interpretable feedback policies for nonlinear systems from data. A prediction model is first identified by rollout-based sparse identification of nonlinear dynamics (SINDy), building on gradient-based and multistep formulations. The policy is then parameterized as a sparse combination of dictionary functions and trained by differentiating a constrained finite-horizon predictive-control objective through this model, so that its terms are selected by closed-loop performance rather than by imitating a previously trained controller. The result is an explicit feedback law with only a handful of terms. Across three benchmark control problems, SD-DPC satisfies the constraints in all test scenarios, outperforms a policy distilled onto the same terms by up to an order of magnitude, and requires orders of magnitude less memory and online computation than an optimization benchmark, while admitting explicit sensitivity bounds.
cs.LG / 161 / 2610.02377
Flow Matching for Fast Posterior Sampling in Bayesian Inverse Problems
Jan Blechschmidt, Oliver G. Ernst, Moritz Poguntke, Björn Sprungk
math.NA · cs.LG · stat.CO
Abstract
Sampling from the posterior is the central task of computational Bayesian inverse problems. The standard workhorse in Bayesian inference - Markov chain Monte Carlo (MCMC) - is sequential, yields correlated samples, and must be rerun for each observation. Conditional flow matching offers an amortized alternative: a transport map, trained once on joint samples of parameter and data, that yields independent approximate posterior samples for any observation at negligible online cost, without new likelihood evaluations. We give a careful, MCMC-literate assessment of flow matching for PDE-based inverse problems with function-valued parameters. Exploiting the flow's tractable density, we derive computable accuracy estimates of the underlying approximate posterior in total-variation distance and Kullback-Leibler divergence and, moreover, propose a hybrid sampler that is asymptotically exact by Metropolization. We validate the accuracy estimates and demonstrate the amortization in several numerical examples, including electrical impedance tomography and a likelihood-free Lotka-Volterra model.
cs.LG / 162 / 2610.02564
Learning Closure of Dynamical Systems with Kernel Ridge Regression
Evan Habbershaw, John Harlim, Senwei Liang
math.NA · cs.LG · math.DS
Abstract
We develop a closure modeling framework for identifying missing components of dynamical systems using Kernel Ridge Regression (KRR). The framework addresses two classes of closure problems: difference-equation closures arising in ODE and PDE settings, and algebraic closures arising from moment closure in kinetic equations. For the first class, we derive an error bound in an ODE setting that quantifies contributions from time integration, approximation of unresolved scales, and interpolation required to couple unresolved-scale effects to the resolved solver. Numerical experiments on the Lorenz-63 system and the Kuramoto-Sivashinsky equation demonstrate accurate long-horizon predictions and substantial improvements over an LSTM-based closure model. For the second class, we consider moment closure for a one-dimensional kinetic equation by modeling discrepancies between kinetic and macroscopic fluxes as a function of the resolved macroscopic variables. We compare global KRR models based on PCA coordinates with spatially local models. While the global model performs well for unimodal initial conditions, its accuracy deteriorates for bimodal initial conditions. Spatially local models with appropriate modeling inputs improve robustness and achieve higher predictive accuracy.
cs.LG / 163 / 2610.02634
A hybrid CNN-adjoint optimization framework for reconstruction of viscoelastic tissue properties in magnetic resonance elastography
Anwesa Dey, Johann Rudi, Elena Cherkaev
math.OC · cs.LG · math.NA
Abstract
Magnetic resonance elastography (MRE) is a noninvasive imaging modality for quantifying the viscoelastic properties of soft tissues from shear wave propagation. Recovering the complex-valued shear modulus from measured displacement fields leads to a severely ill-posed inverse problem, particularly in the presence of noise and limited boundary excitations. We investigate adjoint-based optimization, convolutional neural network (CNN) reconstruction, and a hybrid framework combining both approaches. The forward model is based on a scalar form of the modified stationary Stokes system with a complex shear modulus. We establish well-posedness of the forward problem, existence of minimizers, and first-order optimality conditions for the adjoint-based formulation, and implement a nonlinear conjugate-gradient method with Armijo line search. While PDE-constrained optimization can accurately refine coefficient reconstructions, its performance depends strongly on initialization. We therefore construct a two-dimensional CNN that maps complex-valued displacement measurements to spatially varying complex shear modulus fields and provides rapid, informative initial reconstructions. The proposed hybrid method uses the CNN reconstruction to initialize the adjoint-based optimization, yielding faster convergence and improved accuracy. The CNN is trained on coefficient fields containing individual perturbations and tested on both individual and previously unseen combined configurations. Numerical experiments demonstrate that the CNN generalizes to these more challenging configurations, while subsequent PDE-constrained optimization further refines the reconstructed coefficient. These results demonstrate the potential of combining data-driven initialization with physics-based optimization for efficient and accurate MRE reconstruction.
cs.LG / 164 / 2610.03222
Near-Optimal Convex Optimization with Lazy Second-Order Oracles
Xinliang Zhang, Lesi Chen, Chengchang Liu, Jingzhao Zhang
math.OC · cs.LG · stat.ML
Abstract
This paper studies the complexity of convex optimization using lazy second-order oracles (Doikov, Chayti, and Jaggi, ICML 2023), where an algorithm queries gradients every iteration and Hessians once per $m$ iterations. Under this setting, we show a lower bound of $Ω(m+ m^{1/7} ε^{-2/7})$ on the number of total iterations to find an $ε$-solution using a novel block zero-chain construction. Then we propose a novel method that achieves a new upper bound of $\tilde{\mathcal{O}}(m+ m^{1/7} ε^{-2/7})$, which significantly improves the prior one (Chen, Liu, Luo, and Zhang, COLT 2026) of $\tilde{\mathcal{O}}(m+ m^{13/21} ε^{-2/7})$ and is tight up to logarithmic factors.
cs.LG / 165 / 2610.03709
From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing
Kuangyu Ding, Gesualdo Scutari
math.OC · cs.LG
Abstract
We study the minimization of sums of smooth strongly convex functions over undirected graphs, with each function held by one agent and communication restricted to neighbors in the graph. Existing decentralized methods, whether based on gossip or on routing over spanning trees, typically use the network to mix or aggregate information to enable {\it prescribed} local optimization updates. What this communication-centered viewpoint lacks is a general framework that uses graph structure to {\it jointly} design the optimization subproblems and the cooperative computation and communication through which agents solve them cooperatively. We develop such a framework from first principles, jointly designing the linear representation of agreement constraints, the blocks of the resulting dual variables (jointly optimized), and connected cluster of agents that cooperatively solve each block subproblem over the assigned subgraph. GATE (Graph-Tearing message passing) is a first instance of this framework: one variable per edge and tree blocks. At each iteration, agents update their assigned edge variables by minimizing the sum of the two endpoint cost-to-go messages and relaxing the result. The messages are updated through local minimizations following the tree recursion. To reduce per-iteration computational and communication costs, we develop GATE-S, a surrogate variant using tractable local models and lightweight message parametrizations. We establish linear convergence with a rate explicit in the interplay among function regularity, network topology, and the chosen partition, revealing the effects of graph decomposition. Numerical experiments are conducted to validate the theoretical results and evaluate the efficiency of our algorithms.
cs.LG / 166 / 2610.03444
Electronic Density versus Geometry for Machine-Learned Molecular Absorption Spectra
Siddharth Dhanpal, Peter Elliott, Paolo Emilio Trevisanutto, Alin M. Elena, Gilberto Teobaldi
physics.chem-ph · cond-mat.mtrl-sci · cs.LG
Abstract
Molecular optical absorption spectroscopy provides a direct probe of electronic structure and is widely used for molecular identification, interpretation of photophysical behaviour, and planning of spectroscopy experiments. Calculating the absorption spectra using first-principle excited-state methods, however, is computationally demanding, at least compared to ground-state calculations, which limits their routine application across large molecular sets. Machine-learning (ML) surrogates can reduce this cost and allow rapid spectral prediction. However, their performance depends strongly on how molecular information is represented. Here, we compare using the ground-state electron density versus the molecular geometry as inputs to a ML model for predicting absorption spectra, for a training set of 6874 molecules selected from the QM7 dataset. For each of these molecules, the density was calculated using density functional theory (DFT) and the absorption spectrum was calculated using linear-response (LR) time-dependent DFT (TDDFT). Utilizing the ground-state density as the input to the ML model is motivated by the Hohenberg-Kohn and Runge-Gross theorems, and the fact that the ground-state density encodes information about bonding, charge localisation, and electronic delocalisation. Hence, it may be a more judicious starting point for the ML model compared to the geometry, as it effectively decouples the chemistry of the ground-state. The question we test is whether the benefits of using the density outweigh the (notprohibitive) penalty of requiring an additional single-point DFT calculation for the density. We find that the density-based convolutional neural network achieves a validation correlation of 0.9926, compared with 0.9795 for the best geometry-based graph model, reducing the residual decorrelation, by approximately 64%.
cs.LG / 167 / 2610.03410
Contrastive Neural Embeddings Reveal Individual Traits Beyond Conversational Role
Hubert Huang, Michelle McCleod, Brendan Ames, Evie Malaia
q-bio.NC · cs.LG
Abstract
Contrastive representation learning is increasingly used to recover low-dimensional structure from neural recordings, but its output is typically validated by decoding accuracy rather than by the geometry of the manifold it produces. We apply CEBRA to EEG recorded from dyads in conversation, and analyze the resulting embedding, which training constrains to the 2D sphere. Labels describing the dyads, including the absolute difference between partners' autism-quotient scores, decode well above chance (0.77 against a 0.55 majority baseline for binary AQ magnitude; 0.44 against 0.25 for the six-class $|Δ$AQ$|$ partition). However, the two permutation controls have notable differences in results: permuting labels over a frozen embedding yields p = 0.001, whereas retraining the encoder under each permutation yields p = 0.50. Only the latter tests the label rather than the geometry. Consistent with this, spherical mixture structure and per-class dispersion track identity rather than autism trait differences in dyads; frequency-band and non-oscillatory activity ablation controls do not change the results. However, participant-level model does separate from its identity-aware null (p = 0.0099) while speaker-versus-listener role analysis performs at chance in the same embedding, indicating a manifold organized by individual -- and, in contrast with current neurolinguistics models, almost invariant to speaking vs. listening. Based on these results, we suggest that retraining-based nulls should be the default for grouped-data contrastive embeddings.
cs.LG / 168 / 2610.03369
Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes
Alexander Ardaiz, Varun Budati, Ali Habibnia
q-fin.TR · cs.LG
Abstract
Deep reinforcement-learning policies for order execution can vary substantially across training seeds, so apparent architectural gains may reflect favourable training realisations rather than reproducible properties of the architecture. We evaluate vanilla Double Deep Q-Learning (DDQL), K-means-partitioned mixtures of DDQL experts at $K \in \{2, 4, 8\}$, and dense networks parameter-matched to the $K{=}4$ and $K{=}8$ expert budgets on 5-minute mean-aggregated BTC/USDT limit order book data from Binance. No learned configuration significantly improves mean implementation shortfall over DDQL. Under the reported specification, all have higher mean shortfall than TWAP (0.39 bps) and immediate liquidation (0.21 bps) in an environment whose frictionless replay and terminal-urgency penalty make early liquidation nearly costless; 11/100 vanilla-DDQL runs, versus none in either MoE $K{\geq}4$ arm, converge to a policy that waits until forced liquidation. We then decompose this specification on a device-matched baseline. Annealed exploration alone eliminates observed collapses (12/100 to 0/100; exact McNemar $p{=}4.9{\times}10^{-4}$), matching the elimination under expert partitioning. Combining annealed exploration with the aligned reward restores collapse in 19/30 runs; with all three specification changes, it rises to 48/100. In this environment, expert partitioning is unnecessary to suppress collapse and appears to mask a training-specification failure rather than confer an intrinsic performance benefit. No MoE $K{=}8$ run collapses under any of the six specifications tested. Across-seed dispersion is lowest at $K{=}8$ but non-monotone and not robust to family-wise adjustment, while within-policy tail risk worsens monotonically with $K$. The apparent attribution of the failure mode reverses between 30 and 100 seeds, illustrating the importance of repeated-seed evaluation.
cs.LG / 169 / 2610.03161
Landscape-Dependent Performance of Photonic Quantum Solvers in QUBO Feature Selection for Financial Risk Detection
Nirvik Sahoo, Paul Robert Griffin
quant-ph · cs.LG · q-fin.RM
Abstract
Feature selection for imbalanced classification tasks such as credit card fraud and consumer default detection requires balancing predictive relevance, inter-feature redundancy, and computational feasibility. We benchmark three computing paradigms, classical branch-and-bound optimization (Gurobi), photonic entropy computing (QCI Dirac-3), and simulated photonic boson sampling (Piquasso), across thirteen feature-selection methods on two datasets: ULB Credit Card Fraud (30 features) and AmEx consumer default (159 features). Each method is routed to the solver matched to its mathematical structure. On ULB, Dirac-3 MI-Spearman matches the all-features model using 13 of 30 features (mean F1 0.873 +/- 0.023 over five runs, best run 0.896), and Piquasso is the best method at k=5. On AmEx, performance rises steadily with the feature budget and every paradigm approaches F1 = 0.80 only near the full feature set. Most differences between Gurobi and Dirac-3 on identical methods fall within run-to-run variation; the large gaps occur where the certified optimum generalizes poorly, most sharply for distance correlation on AmEx at k=25 (Gurobi F1 = 0.422 vs. a Dirac-3 mean of 0.746). At matched budgets, F1 varies about ten times more across methods on ULB than on AmEx, which we trace to how concentrated the predictive signal is in each feature space.
cs.LG / 170 / 2610.03205
Hamiltonian locality testing and certification do not achieve the Heisenberg limit
Francisco Escudero Gutiérrez, Junseo Lee, Sebastian Zur
quant-ph · cs.CC · cs.DS · cs.LG
Abstract
We establish lower bounds for Hamiltonian property testing with access to the time-evolution operator but not its inverse. Each experiment may query the time-evolution operator multiple times, and distances between Hamiltonians are measured in the normalized Frobenius norm. In this model, we show that testing whether a Hamiltonian is $k$-local or $\varepsilon$-far from every $k$-local Hamiltonian requires $Ω(1/\varepsilon^2)$ total evolution time, matching the upper bound of Kallaugher and Liang (TQC'25). We also prove that testing whether an unknown Hamiltonian equals a target Hamiltonian or is $\varepsilon$-far from it requires $Ω(1/\varepsilon^2)$ total evolution time, matching the upper bound of Sinha and Tong (2025). These are the first lower bounds for natural problems in Hamiltonian learning and testing that rule out Heisenberg-limited scaling of $1/\varepsilon$. As a third result, we show that amplitude estimation to precision $\varepsilon$ requires $Ω(1/\varepsilon^2)$ total time evolution, recovering the result of Tang and Wright (QIP'26) in the continuous-time query model. All three results follow from the hardness of distinguishing the zero Hamiltonian from a suitably chosen ensemble of random Hamiltonians. We establish this hardness by adapting the continuous-time adversary method to forward Hamiltonian evolution.
cs.LG / 171 / 2610.03463
Generalization of Transformer-Based Neural Quantum States via In-Context Learning
Zhen Qin, Qing Qu, Alfred O. Hero
quant-ph · cs.LG
Abstract
Neural quantum states based on modern deep learning architectures have emerged as powerful representations for quantum many-body systems. In particular, Transformer-based neural quantum states provide expressive models capable of capturing long-range correlations, and their empirical generalization performance has recently been demonstrated. However, a theoretical understanding of their generalization behavior remains largely unexplored. In this paper, we develop a theoretical framework to analyze the generalization properties of Transformer-based neural quantum states under in-context learning. We establish a rigorous inference-time generalization error bound in terms of mean squared error (MSE), showing that the pointwise prediction error decreases inversely with both the number of in-context examples and the depth of the Transformer. We further show that the Transformer depth required to achieve this guarantee scales only linearly with the system size--namely, the number of particles in continuous systems or the number of qudits in discrete systems. Building on this result, we extend our analysis to full quantum states formulated as rank-one density operators, and derive MSE-based generalization bounds over both continuous and discrete domains under physical constraints. Finally, numerical simulations corroborate our theoretical analysis.
cs.LG / 172 / 2610.02357
Conformal Prediction for Time Series with Deep Sequence Models
Junghwan Lee, Jonghyeok Lee, Yao Xie
stat.ML · cs.LG
Abstract
Recent advances in deep learning for time series prediction have amplified the need for reliable uncertainty quantification. Conformal prediction has gained attention as a distribution-free framework for constructing prediction intervals with coverage guarantees. However, its coverage guarantees rely on data exchangeability, an assumption generally violated in time series. Active research has focused on developing conformal prediction methods for time series that overcome this limitation. While deep sequence models, such as recurrent neural networks and Transformers, have often been used in conformal prediction for time series, limited work has systematically studied how deep sequence models can be utilized in conformal prediction for time series. In this work, we systematically investigate the use of deep sequence models in conformal prediction for time series through three approaches: conditional quantile regression, conditional quantile function estimation, and localized conformal prediction. We provide a theoretical analysis establishing asymptotic conditional coverage guarantees for all three approaches under suitable assumptions. Through comprehensive experiments on real-world datasets, we demonstrate the effectiveness of leveraging deep sequence models into conformal prediction for time series.
cs.LG / 173 / 2610.02578
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning
Filip Kovačević, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli
stat.ML · cs.LG
Abstract
To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of $ρ$-zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.
cs.LG / 174 / 2610.02785
Hold-Out Scoring for Efficient Gaussian DAG Learning
Donguk Shin, Byeongguk Kang, Inseol Lee, Gunwoong Park
stat.ML · cs.LG
Abstract
High-dimensional Gaussian DAG learning faces a statistical-computational gap: methods with sharp sample complexity rely on computationally expensive subset search and a supplied indegree bound, whereas polynomial-time alternatives have less favorable sample complexity. We introduce HOST, an efficient DAG learning algorithm that replaces subset search with nodewise hold-out scoring and convex regression, without requiring a supplied indegree bound. Our key insight is that recovering a correct ordering does not require uniformly small estimation errors in ordering scores but only one-sided control of those errors. In the ordering step, HOST exploits the fact that score estimation using hold-out samples inflates ordering scores in expectation, which is the favorable direction for candidates that should not yet be selected. Given the ordering, HOST recovers parents by recursively removing indirect effects from total effects between two nodes. Under suitable conditions, HOST exactly recovers a $p$-node DAG of maximum indegree $d$ with sample complexity of order $d\log p$ in polynomial time. Experiments show that HOST achieves competitive graph recovery while exhibiting favorable runtime scaling.
cs.LG / 175 / 2610.03101
Invariance of Clustering Operations in Causal Effect Identification
Jani Nykänen, Otto Tabell, Santtu Tikka, Juha Karvanen
stat.ML · cs.LG
Abstract
Clustering variables in causal graphs reduces the size of the graph and simplifies causal inference. However, arbitrary clustering can alter crucial causal relations among variables and lead to erroneous conclusions. While the identifiability of a causal effect in the clustered graph implies the identifiability in the original graph under mild conditions, nonidentifiability in clustered graph does not imply nonidentifiability in the original graph without further assumptions. When both identifiability and nonidentifiability are preserved, the clustering operation is called identification invariant. We present a broad class of clustering operations that are identification invariant based on conditions related to the c-components of the original graph. Finally, we demonstrate use of the results in practical settings.
cs.LG / 176 / 2610.03201
Predictively Oriented Gaussian Process Posteriors
Callum Lau, Jeremias Knoblauch, Louis Sharrock
stat.ML · cs.LG
Abstract
Gaussian Processes (GPs) are a powerful tool for modelling and quantifying uncertainty in functional relationships. However, they require practitioners to make a number of design decisions, such as the choice of the kernel and the observation model. Suboptimal choices can produce misspecified models that do not capture the underlying data generating process. We introduce Predictively Oriented Gaussian Processes (PrO-GPs), which treat predictive uncertainty as the primary inferential target and provide a robust alternative to standard GPs. Although direct computation of a PrO posterior for nonparametric models is intractable, we derive a reduced formulation and practical sampling scheme for efficient computation. Through synthetic and real data experiments, we show that PrO-GPs produce better calibrated predictive distributions under model misspecification compared to standard GP approaches.
cs.LG / 177 / 2610.03313
SDECast: Probabilistic Weather Forecasting in Continuous Time with Neural SDEs
Maria Marchenko, Martin Andrae, Fredrik Lindsten, Christian A. Naesseth
stat.ML · cs.LG · physics.ao-ph
Abstract
Existing machine learning weather forecasting models typically generate forecasts through autoregressive rollouts at a fixed temporal resolution. While highly efficient for long-range prediction, this formulation can suffer from severe error accumulation when used with shorter time steps and does not explicitly encode the locality and temporal continuity of atmospheric dynamics. To address these limitations, we introduce **SDECast**, a Neural Stochastic Differential Equation (SDE) framework for continuous-time probabilistic weather forecasting. SDECast extends SDE Matching to learn stochastic dynamics directly in physical space, without requiring repeated SDE simulation during training. On a simulated geophysical flow, we show that SDECast recovers meaningful drift dynamics and faithfully reproduces the underlying continuous-time behavior. We then demonstrate its scalability to global weather forecasting at hourly resolution, where SDECast produces skillful probabilistic forecasts for lead times of up to five days.
cs.LG / 178 / 2610.03414
Iterating Consistency Models: Stability, Error Bounds and Noise Schedules
Alessio Spagnoletti, Abdul-Lateef Haji-Ali, Andrés Almansa, Alain Oliviero Durmus, Eric Moulines, Marcelo Pereyra
stat.ML · cs.CV · cs.LG
Abstract
Consistency models (CMs) have become a leading approach for generating high-quality samples in few steps. However, adding steps can improve or degrade sample quality in ways that are highly sensitive to the schedule and that existing theory does not fully explain. To provide accuracy guarantees and guide CM sampler design, we analyze multistep CM sampling as a composition of noising and approximate denoising operators. Under explicit, verifiable stability assumptions, we derive a non-asymptotic error bound that separates contraction of the initialization error from accumulation of approximation error. The bound assigns distinct roles to the schedule: large early noise levels drive contraction, while small late noise levels control the residual bias. As a corollary, we obtain explicit constants for strongly log-concave and semi-log-concave targets. We further establish a complementary guarantee whose assumptions, one-step accuracy and stability, can be estimated for a given trained model. Experiments show that the contraction and approximation profiles entering our bounds can be reliably measured and closely match the predicted functional forms. Together, these results provide a meaningful convergence theory for multi-step CMs and a practical route to sampler design.
cs.LG / 179 / 2610.03465
When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion
JM Gorriz
stat.ML · cs.LG · physics.data-an
Abstract
K-fold cross-validation (CV) is widely used as evidence of out-of-sample performance, although folds are neither independent experiments nor equally informative under heterogeneous data. Cross Upper-Bound Validation (CUBV) replaces point-wise CV accuracy by conservative upper bounds on true risk. Here we generalise CUBV through a single exponential framework in which the moment-generating function of the generalisation gap is controlled by a cumulant envelope gamma(lambda). This yields a family of risk bounds covering Hoeffding-, Bernstein-, dependency-aware, PAC-Bayesian, and heterogeneous source-fusion settings. For K-fold CV, dependence between fold-wise gaps is modelled through a joint sub-Gaussian proxy matrix. Under equicorrelation, this gives an effective number of folds, Keff = K/[1+(K-1)rho], showing that increasing K does not necessarily increase statistical evidence when folds are strongly dependent. The framework is also extended to posterior distributions over predictors and weighted multi-source fusion, where weights are selected by minimising an upper bound on future risk rather than empirical error alone. Experiments with trained linear classifiers on heterogeneous multimodal Gaussian mixtures compare K-fold CV with full-sample resubstitution plus risk correction. Bounds are evaluated by coverage and tightness. In low-dimensional small-sample settings, K-fold partitioning can increase uncertainty because individual folds under-represent minority modes, while corrected resubstitution can remain valid and tighter; this effect disappears as sample size increases. Overall, gamma-CUBV separates observed performance, uncertainty, dependence, model complexity, and confidence into explicit terms, providing a unified route from CV scores to risk statements and a principled validation criterion for heterogeneous small-sample applications such as neuroimaging.
cs.LG / 180 / 2610.03647
Amortized Structured Stochastic Variational Inference for Gaussian Process Latent Variable Models
Maksym Tretiakov, Sarah Lucie Filipp, Vincent Fortuin, Ruth Misener, Ruby Sedgwick, James Odgers
stat.ML · cs.LG
Abstract
Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gaussian Process Latent Variable Model, in which a Gaussian Process (GP) mapping from the latent space provides an estimate of the uncertainty of the manifold. However, the effectiveness of this uncertainty estimation is limited by the mean-field variational approximation between the GP inducing points and the latent variables. In this work, we apply Amortized Structured Stochastic Variational Inference to allow the variational posterior for the latent space to be conditionally dependent on the value of the inducing points. We demonstrate that this more flexible variational posterior improves several metrics relating to the reconstruction of points on the data manifold.
神经与进化计算 (cs.NE)
9
cs.NE / 1 / 2610.02996
Evolutionary Computation for Trustworthy AI: From Attacks and Defenses to Self-Evolving Era
Junhao Dong, Chenkai Wang, Xuanhui Lin, Mingrong Gong, Siyu Wang, Yuqing Wen, Jiao Liu, Catherine Huang, Gary G. Yen, Xin Yao, Yew-Soon Ong
cs.NE
Abstract
As Artificial Intelligence (AI) has evolved from task-specific models to foundation models and agents, the scope of trustworthy AI has expanded from model-level robustness to the reliability and safety of broader AI systems. This evolution has also expanded the attack surface from individual models to broader system-level interactions, including tool use, context, and interaction trajectories with dynamic environments. As a result, maintaining reliable and safe behavior under changing or deliberately manipulated conditions has become increasingly challenging. The search for effective attacks and defenses often relies on black-box feedback to navigate discrete choices among words, actions, system components, or their combinations. Multiple objectives and expensive candidate evaluations further limit what can be explored. Evolutionary Computation (EC), with its population-based, gradient-free search and flexible variation and selection mechanisms, is well suited to these settings. This survey reviews how EC has been applied to trustworthy AI across three directions: evolutionary attacks, evolutionary defenses, and trustworthy self-evolving AI systems. Unlike prior reviews that treat trustworthy AI, EC, and self-evolving systems largely separately, we connect these lines through a common evolutionary perspective. For self-evolving AI, we examine how trustworthiness governs the generation and retention of updates that shape subsequent adaptation. We further synthesize evaluation methods and benchmark resources from both trustworthiness and evolutionary-search perspectives. Finally, we discuss key challenges and future research directions toward more effective and reliable use of EC in trustworthy AI.
cs.NE / 2 / 2610.03021
Small universal multiset reaction systems
Andrei Paun, Annemarie-Beatrix Messner
cs.NE · cs.FL
Abstract
This paper presents a construction of a small universal machine within the framework of Multi-set Reaction Systems. We encode the machine state using elements that represent registers and instruction labels, and we enforce sequential execution by ensuring that reaction chains for a given instruction can only activate after the corresponding label element appears. We implement a universal register machine with 8 registers and 23 instructions, showing that the priority-based model requires 87 distinct elements. The results demonstrate that multiset reaction systems are capable of simulating universal computation.
cs.NE / 3 / 2610.03149
Learning While Inferring: Local and Parallel Learning for Edge SNNs across Sensing Modalities
Yanxun Zhang, Yifei Wang, Changze Lv, Jingwen Xu, Yiyang Lu, Xiaohua Wang, Di Yu, Xin Du, Xiaoqing Zheng
cs.NE
Abstract
Edge intelligence requires models to sense continuously in real time and to keep adapting on-device, all under tight compute, energy, and memory budgets. Although spiking neural networks (SNNs) enable efficient event-driven inference, standard surrogate-gradient backpropagation (BP) serializes updates and blocks ongoing inference. We investigate Bidirectional Spike-Based Distillation (BSD) as an on-device learning principle that lets edge SNNs learn while inferring. BSD couples a stimulus-driven forward network with an independent target-driven reverse network and aligns their intermediate representations through local objectives. Because the two pathways have disjoint computation graphs until alignment, forward inference, reverse inference, and stage-wise updates can run concurrently. On 25 benchmarks spanning the five sensing modalities of SOUL, BSD stays within 3.8 percentage points of matched BP baselines on average, while only its forward branch is needed at deployment. The learned representations also transfer well to few-shot class-incremental learning without replaying past data. By removing the dense floating-point backward chain, BSD reduces the projected training latency to $0.72\times$ and the estimated training energy to $0.36\times$ that of BP.
cs.NE / 4 / 2610.03155
Aggregate accuracy conceals concentrated temporal vulnerability in a spiking speech classifier
İsmail Can Dikmen
cs.NE
Abstract
Aggregate accuracy cannot reveal which utterances are locally vulnerable or how internal activity changes when labels remain stable. We retain every prediction for 725,070 adjacent-bin, one-count changes around 100 validation utterances of a frozen SpikeSCR-based classifier. The canonical native-horizon GPU singleton path reaches 86.0836% validation accuracy. Equal-source expected accuracy under a uniformly chosen neighbor rises from 84.00% to 84.54%, although 13 of 84 initially correct sources admit adverse neighbors. Five sources carry 93.21% of adverse moves. Margin-guided and surrogate-gradient searches miss sparse cases at fixed query budgets; a post hoc gradient prefix finds all 13 at 73.45% of the census of initially correct sources. Source-matched traces and clean-state interventions distinguish internal change from harmful direction and recoverability. Two count-readout replicas have similar validation accuracies but a 5.21-fold class-change gap concentrated in four sources. Controlled query/key source isolation eliminates 531 batch-order label changes and restores bitwise order invariance across all 9,981 validation score vectors. Isolated batch predictions reproduce padding-matched singleton labels on 9,980 of 9,981 inputs. Paired CPU/GPU replay of all 25,820 class-changing neighbors gives 99.8993% label agreement and preserves all 13 vulnerable sources and their fixed witnesses. Complete source-conditioned maps distinguish vulnerability incidence, concentration, internal change, and execution dependence that aggregate accuracy leaves unresolved.
cs.NE / 5 / 2610.03220
Evolving Hybrid Quantum-Classical Architectures for Image Classification
Devroop Kar, Daniel Krutz, Travis Desell
cs.NE · cs.AI · cs.CV · cs.LG · quant-ph
Abstract
Hybrid quantum classical neural networks integrate parameterized quantum circuits (PQCs) with established deep learning architectures, but their performance depends strongly on the choice of quantum circuit architecture, a choice that remains largely manual. Most existing approaches rely on hand-designed or fixed circuit ansätze, requiring circuit structure, gate composition, and qubit connectivity to be specified in advance with no guarantee that they suit the task. This limitation is especially acute in image classification, where quantum circuits must transform features extracted by classical networks while remaining compact enough for practical training, requirements that generic, task-agnostic ansätze are unlikely to satisfy simultaneously. We extend EXAQC, an evolutionary framework for automated quantum circuit discovery, to image classification. EXAQC evolves PQCs as intermediate processing modules while retaining classical feature-extraction and prediction layers. On MNIST, Fashion-MNIST, and CIFAR-10, EXAQC achieves 98.42%, 90.62%, and 85.47% accuracy, respectively, while using comparable gate counts to other quantum architecture-search methods. Against classical networks, evolved hybrid models maintain comparable accuracy with substantially fewer trainable parameters, reaching 85.68% on CIFAR-10 with over 25$\times$ fewer parameters than a 10-layer CNN. Encoding choice also matters: rotation-based encodings (RX, RY, U3) outperform amplitude encoding by 22-25 points on CIFAR-10. These results demonstrate that automated circuit discovery yields compact quantum modules that can replace larger classical components in vision architectures while retaining competitive accuracy.
cs.NE / 6 / 2610.03249
Self-Repairing Recurrent Ensembles for Real-Time Recovery from Distribution Shift
Julian Lemmel, Pedro D. Wendel Garcia, Taisuke Kobayashi, Radu Grosu
cs.NE · cs.RO · eess.SY
Abstract
Deploying a pretrained controller exposes it to conditions that are absent from its training data. Sensor drift, outright sensor failure and accumulating measurement noise all induce a distribution shift that can collapse an otherwise competent policy; typically at a point in time where no expert is available to supply corrective labels. We present a method that lets a policy recover from such shifts online and without supervision. Our controller is an ensemble of recurrent networks, each of which observes a randomly masked subset of the observation vector, and whose Gaussian outputs are combined through sequential Kalman fusion so that confident members dominate the consensus action. At deployment, we treat this consensus as a self-supervised label and fine-tune each member towards it, scaling each member's contribution proportional to the complement of its squared Kalman gain. Gradients are computed using RFLO, an efficient and biologically plausible approximation of Real-Time Recurrent Learning, so that a parameter update follows every environment step and the policy reacts to a shift as it unfolds. On a range of simulated continuous control tasks, our approach recovers close to the original performance after a sensor shift, while ensembles that see the full observation are unable to recover. The same framework subsumes fully online interactive imitation learning: when an expert is present, the consensus label is replaced by the expert action and the identical update rule refines the policy during teleoperation.
cs.NE / 7 / 2610.03282
The Investment Acceleration Principle Revisited by means of a Neural Network
Guido Fioretti
cs.NE
Abstract
The investment acceleration principle is a heuristic for modelling investment time series out of consumption time series. The model presented herein develops a disaggregated accelerator equation whose coefficients are the weights of a Kohonen neural net that represents firms' decision-making. According to this model, investments take place when managers recognise emerging technological patterns. Furthermore, a technique borrowed from the theory of self-organising systems is used in order to disentangle innovation-driven investments from plant-replication investments.
cs.NE / 8 / 2610.03291
Parallel Time-Aligned Spiking Self-Attention for Consistent Integer-Valued Training and Spike-Driven Inference
Peng Xue, Wei Fang, Kaiwei Che, Qingyan Meng, Zhengyu Ma, Yonghong Tian, Huihui Zhou
cs.NE
Abstract
Integer-valued leaky integrate-and-fire (I-LIF) neurons and spike firing approximation (SFA) reduce temporal training cost by representing spike trains as firing counts and normalized firing rates, respectively. However, applying spiking self-attention (SSA) directly to these compressed query, key, and value representations introduces cross-time interactions that are absent during spike-driven inference. We term this operator-level discrepancy Temporal Interaction Mismatch (TIM). We propose Parallel Time-Aligned Spiking Self-Attention (PT-SSA), which reconstructs consecutive virtual spike slices from either I-LIF counts or SFA firing rates, computes attention only between time-aligned slices in parallel, and sums the per-step outputs. To accommodate the reduced attention output scale under SFA, we further introduce Adaptive PT-SSA, which learns a positive per-block rescaling before the output SFA neuron to improve firing-level utilization. Experiments on CIFAR-10, CIFAR-100, and ImageNet-1K show that the proposed methods substantially reduce train--inference mismatch. On CIFAR-100 with I-LIF, PT-SSA reduces the mean Top-1 gap from 1.85 to 0.29 percentage points. On ImageNet-1K, Adaptive PT-SSA reduces the Top-1 gap from 27.78 to 0.06 percentage points and achieves 74.53\% spike-driven Top-1 accuracy. A Triton-fused PT-SSA training kernel retains a $2.91\times$ throughput advantage over recurrent LIF SSA.
cs.NE / 9 / 2610.03675
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
Hui Chen, Xuan Qi, James Xu Zhao, Zhaopeng Feng, Shilong Liu, Kuang Xu, Pang Wei Koh, Bryan Hooi
cs.NE · cs.AI · cs.CL
Abstract
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the resulting code. We also design a cache-efficient evolution process, where our harness and prompts maximize the sharing of prefixes across different evolution steps, to improve cache reuse. To measure solution quality throughout a fixed cost budget, we introduce Budget-Aware Area Under the Curve (BA-AUC), defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget. Across 10 mathematical and systems optimization tasks, FrugalEvo matches or surpasses state-of-the-art baselines, including OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX, in final solution quality and achieves higher BA-AUC on 9 tasks. It also achieves higher average performance than these baselines on 10 algorithmic optimization tasks from ALE-Bench-Lite. Notably, on circle packing, FrugalEvo achieves new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD and with GLM-5.3 and its Flash variant for only 0.55 USD, matching or surpassing all baselines, including multi-agent methods such as CORAL and SwarmResearch, which cost approximately 50 USD on average.
计算语言学 (cs.CL)
41
cs.CL / 1 / 2610.02425
Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
Yekun Chai, Qiwei Peng, Haoyi Xiong
cs.CL
Abstract
Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1\% of Sighted trials, yet only 13.9\% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7\% pass@3 but only 5.9\% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3\% of accepted simulation calls stop on an illegal move, and in 49.3\% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.
cs.CL / 2 / 2610.02455
FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms
Chin-Lun Fu, Hong Ni, Behrouz Madahian
cs.CL · cs.AI
Abstract
Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the RFQ-trigger message (the inquiry message), interleaved with concurrent RFQs from other participants. We cast this as event extraction (EE) over multi-party dialogue and present FinDialogLens, a hybrid LLM pipeline in which compact fine-tuned classifiers act as inference-time scaffolds: they detect RFQ-triggers and price/trade outcome metadata, an RFQ-Level Module segments per-event RFQ windows, and a Trade Engine fills argument roles. With GPT-4o, FinDialogLens reaches 92.1% and 94.3% accuracy on final price and trade outcome, respectively, outperforming full-chatroom CoT prompting methods; fine-tuned open-source LLMs with as few as 3B parameters achieve comparable performance with modest in-domain data. To make the LLM-based solution practical at scale, a difficulty-aware router balances cost and accuracy by allocating RFQs between a low-cost rule-based engine and the higher-performing LLM-powered Trade Engine, cutting LLM calls by 85% on final price while recovering half of the accuracy gap to FinDialogLens (GPT-4o), saving over $300/day at our 70,000-RFQ/day scale.
cs.CL / 3 / 2610.02460
CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking
Anjali Kantharuban, Jonas Mueller
cs.CL
Abstract
Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training. Through this, we evaluate user-conditioned replay of past sessions and aggregate metric agreement when sampling novel personas for the same tasks. On $τ^2$-Bench, CUEd simulators commit fewer simulator-attributed errors and more faithfully reproduce real-user agent failure modes, aggregate success rates, and outcomes for specific task-user pairs than other persona-based simulation methods. These gains coexist with competitive user fidelity as measured using metrics established in prior work. After being fit to mostly customer support interactions, the same CUEd simulators generalize to document creation, math tutoring, and casual conversation, and remain effective across different simulator LLMs without CUE retraining.
cs.CL / 4 / 2610.02472
APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory
Chin-Lun Fu, Anagha Kulkarni, Hong Ni, Behrouz Madahian
cs.CL · cs.AI
Abstract
Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (Agent-controlled Progressive Disclosure Memory), a hierarchical long-term memory architecture that applies progressive disclosure to memory retrieval. Rather than relying on a flat memory store or fixed retrieval granularity, APDMem represents conversation history as four progressively detailed layers: thematic summaries, personalized key facts, turn-level evidence notes, and raw messages. At inference time, a controller applies progressive disclosure to the memory hierarchy: it first reads high-level summaries and drills into finer evidence only when needed. This creates an adaptive cost-fidelity trade-off: simple queries can terminate early, while complex temporal, multi-hop, or exact-evidence queries trigger deeper inspection. A note synthesizer converts retrieved evidence into a query-focused structure that consolidates facts, orders events, and flags contradictions before final answer generation. Experiments on LongMemEval show that APDMem achieves strong performance for long-context memory reasoning while accessing only 8% of the total conversations.
cs.CL / 5 / 2610.02486
From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
Pritam Deka
cs.CL · cs.AI
Abstract
Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.
cs.CL / 6 / 2610.02529
A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs
Mohammed Damom, Muneef Y. Alshawsh, Ashraf A. Naji, Mustafa Ali Alhamzi, Fawwaz An-Nashef, Jameel Ahmed Elayah, Mohammed Q. Shormani, Noman AL-Sayadi
cs.CL
Abstract
Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity in Modern Standard Arabic (MSA) DPs. The framework integrates generative syntactic notions with AraBERT by representing ambiguity as a candidate-based decision task in which linguistically motivated alternatives are explicitly constructed and evaluated through candidate-conditioned input representations. Findings indicate that the model achieved 96.88% accuracy, 95.92% macro-F1, 96.83% weighted F1, and 93.94% binary F1 on the unseen evaluation set. Class-level analysis revealed asymmetric performance, with recall of 99.71% for High/VP Attachment (N1) and 89.26% for Low/NP/Embedded Attachment (N2), indicating greater difficulty in recovering the embedded interpretation. The study concludes that formal syntactic representations can be operationalized within Transformer-based NLP as an explicit interface between linguistic structure and contextual neural modeling, providing a controlled and interpretable approach to Arabic syntactic ambiguity resolution and beyond.
cs.CL / 7 / 2610.02612
Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation
Hieu Hoang, Amittai Axelrod
cs.CL
Abstract
Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.
cs.CL / 8 / 2610.02702
Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise
Ziang Ni, Peng Zou
cs.CL · cs.MA
Abstract
Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in "the capital of the country where the Sagrada Familia is located") is never stated by anyone. Scripted peers, in the role of Asch's confederates, unanimously assert a wrong answer taken from another fact with a different bridge. At the moment the agent answers, we read the bridge from its residual stream with the Jacobian lens (J-lens) and, for comparison, the logit lens. In pre-registered tests on held-out facts with four open-weight models, agents of Qwen3.5-4B, Qwen3.6-27B and Gemma-4-E4B-it that gave in still represented their original bridge in the pre-registered layers below the output (hit@100 above a control entity: 0.85, 0.22 and 0.24), where the logit lens rarely ranked it among the top 100 tokens (0.00-0.06). These agents also represented the bridge behind the peers' answer, beyond a mention baseline. A pre-registered addendum hid the agent's earlier answer or removed it: agents that gave in still represented their original bridge in all four models (0.43, 0.29, 0.37 and 0.25 with the answer hidden), including Llama-3.1-8B-Instruct, which barely did so with its answer in view (0.03). The premise can thus be computed from the question alone while the agent states the majority's answer. Hiding the earlier answer also changed conformity: Qwen3.5-4B gave in on 89% of questions instead of 8%. In exploratory interventions, injecting the bridge's J-lens direction brought agents back to their original answer only in the two Qwen models. Stated consensus in multi-agent debate can thus overstate agreement. We also report the negative results of our pre-registered program.
cs.CL / 9 / 2610.02713
WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds
Utkarsh Ranjan
cs.CL · cs.AR · cs.DC · cs.LG
Abstract
Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.
cs.CL / 10 / 2610.02736
TPBench: A Turning-Point Benchmark for Dialogue Compression
Minji Park, Seunghyun Yoon, Hyuk Lim
cs.CL · cs.AI
Abstract
A compressor can keep the facts of a dialogue and still drop the turn that changed them. A user corrects a price, reverses a choice, or adds a constraint. We call this failure turning-point eviction. One overall retention score hides it, because that score mixes what the user first wanted with what the user wants now. We introduce TPBench, which evaluates three complementary information targets at shared nominal retention budgets. P1 asks for the user's initial goal. P2 asks for the current value of a slot the user revised. P3 asks for both, in dialogues with a late annotated slot update. The current-value answers come from the human dialogue-state annotations of MultiWOZ and SGD. The initial-goal answer is the first sentence of the first user turn. Neither requires new crowdsourcing. The probe-specific evaluations rank compression methods differently. On the joint probe at a retained fraction of 0.30, every tested compressed method remains below full context with the main Llama reader. Deleting the turn that carries the update sharply lowers current-value accuracy, while deleting one matched irrelevant turn leaves it unchanged. A Mistral reader repeats the P2/P3 rankings and the joint-probe gap. Current-value recovery is tested on an additional corpus, LongMemEval-KU, and on Chinese RiSAWOZ: full context has the highest accuracy, and recency has the highest compressed-method mean in both evaluations.
cs.CL / 11 / 2610.02739
Beyond Correctness: Resolving Underspecification in Agentic Text-to-SQL
Wen-Zhi Li, Yue Gong, Konstantinos Kanellis, Balakrishnan Murali Narayanaswamy
cs.CL
Abstract
Agentic Text-to-SQL systems can interact with users to clarify underspecified queries before generating SQL. However, a correct execution result does not necessarily imply that the agent has adequately resolved the underlying underspecification: the agent may silently make unverified assumptions that happen to match the intended answer. We show that this behavior is driven in part by premature clarification termination. Although forcing an agent to ask more questions improves execution accuracy, ambiguities are concentrated in earlier interactions, making brute-force questioning inefficient. More importantly, even when explicitly prompted to plan its clarification process, the agent frequently abandons questions that it has already identified as relevant. To address this failure mode, we introduce PlanPool, which externalizes the clarification plan as a mutable question pool. Every planned question must be explicitly asked or dropped before submission, while newly discovered ambiguities can be added during interaction. Across three benchmarks derived from BIRD-Interact and Spider, PlanPool consistently improves ambiguity coverage and reduces silent failures over unconstrained and prompt-based alternatives, while maintaining competitive execution accuracy. Our results highlight an important distinction in agentic reasoning: identifying missing information is not sufficient, and the agent must also reliably maintain and resolve it before committing to an answer.
cs.CL / 12 / 2610.02769
When History Fails to Become Experience: Action Calibration in Language Agents
Jingyu Liu, Zhiwen Wang, Yuxin Jing, Huanyu Zhou, Yong Liu
cs.CL · cs.AI
Abstract
Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we examine how agents use history. We find that history improves task completion overall, yet much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past actions and selectively records experience to guide subsequent decisions, improving task success beyond outcome labeling alone.
cs.CL / 13 / 2610.02770
AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration
Hy Nguyen, Nabi Rezvani, Robin Vujanic
cs.CL · cs.DB
Abstract
Document databases such as MongoDB are core infrastructure for modern applications, and natural-language interfaces to them---text-to-MQL---would let non-experts query complex, semi-structured data without mastering the query language. Progress on this task depends on high-quality benchmarks, which are most practically obtained by converting an existing text-to-SQL benchmark to the document setting. Unfortunately, existing efforts rely on heuristics for mechanical conversion: the document schema mirrors the relational foreign-key graph, and each query mirrors its source SQL. As a result in our experiments, these approaches fail to migrate 6 of 21 BIRD databases outright, silently drop up to 25.9\% of rows on others, and yield schemas whose ground-truth queries run over an order of magnitude slower as the data scales. We instead propose a conversion pipeline, driven by coding agents with human-in-the-loop verification, that designs each document schema from expected access patterns and rewrites queries to be MongoDB-native. Applying it to BIRD, we build an access-pattern-based text-to-MQL benchmark (AptMQL-Bench). It includes 21 document-oriented databases, 3,186 natural-language requests, and their associated MQL queries---whose databases are migrated from SQLite without data loss and scale efficiently. The strongest model, Claude Opus 4.5, achieves only 57.38\% accuracy without external knowledge evidence and 70.34\% with it. This indicates that realistic text-to-MQL generation remains challenging.
cs.CL / 14 / 2610.02841
How Robust Is Multimodal Claim Verification to LLM Rewriting?
Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa
cs.CL
Abstract
LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM-associated word to isolate the effect of vocabulary choice. We evaluate 11 open-weight models spanning five VLM families and ranging from 2B to 38B parameters. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review-score manipulation, verification appears far more stable. However, consistent probability shifts do occur. Hedging-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions (e.g., grammar correction, fluency improvement) have little effect.
cs.CL / 15 / 2610.02873
ConvoDrift: A Multi-Turn Conversational Dataset for Modeling Stylistic Tone Evolution
Vihindi Kotalawala, Pamoda Dilranga, Gayani Thoradeniya, Prasan Yapa
cs.CL · cs.AI
Abstract
The evolution of linguistic style in conversations is an underexplored issue in NLP. Most style-control datasets focus on sentences or assume a static style throughout, missing the dynamic shifts that occur as user preferences change during interactions. We introduce ConvoDrift, a dataset designed to model progressive stylistic conversational tone drift under fixed semantic intent. It is built on 15,727 shared multi-turn conversational structures for adaptation and persona-conditioned alignment methods. It consists of six prompt-response pairs per conversation, each with the annotation of style drift and style direction labels. These pairs cover a range of communication genres. We further derive a complementary pairwise dataset by pairing semantically equivalent but stylistically distinct responses and annotating persona-conditioned preferences using five distinct style communication personas, enabling the controlled study of personalisation and pluralistic alignment in language tone. In addition to dataset construction, we conduct a comprehensive evaluation involving human validation, LLM-as-judge assessment, and automatic lexical and semantic evaluations. Across seven Likert criteria annotated by three human annotators, the average Krippendorff's alpha is 0.88, and our lexical and semantic analyses show that drift events induce lexical changes while preserving semantic similarity.
cs.CL / 16 / 2610.02875
Query-aware routing for Cross-lingual performance gains in Encoders
Akshay Jain, Edward Kim
cs.CL · cs.AI · cs.IR
Abstract
Multilingual encoders can exhibit reduced retrieval effectiveness when queries and relevant documents differ in language, despite strong same-language performance. We investigate whether Finnish and Swedish cross-lingual retrieval can improve while preserving an encoder's existing same-language performance and document index. We combine a query-only low-rank adapter, trained against frozen document embeddings, with deterministic routing based on query and index languages. Cross-language queries use the adapter, while same-language queries use the original encoder. SampoTron, our fine-tuned low-rank (LoRA) adapter alongwith the Nemotron-3-Embed-1B model, improves average retrieval quality across six English, Finnish, and Swedish directions from 0.241 to 0.291 in normalized discounted cumulative gain (nDCG) at rank ten, a 20.9% relative gain on a sampled financial benchmark. All six cross-lingual directions improve, and routing preserves the original same-language performance, including two full-corpus Finnish evaluations. The approach enables selective cross-language specialization with reusable document embedding vectors.
cs.CL / 17 / 2610.02878
Evaluating VQA in Vision Language Models using Cooperative Principles
Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal
cs.CL
Abstract
We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, but VLMs themselves perform less accurately in such cases.
cs.CL / 18 / 2610.02886
Misinformation Without Triggers: From Factual Answers to Downstream Decisions
Lin Tian, Marian-Andrei Rizoiu
cs.CL · cs.AI
Abstract
Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a misleading claim from Facebook posts about the 2019--20 Australian bushfires. Across eight models at dose 1,000, direct injected-choice rates reach 95.8--100\%, while injected game choices increase by 1.7--14.4 percentage points over matched truthful training. The gap runs the other way too. Facts that pass the direct probe still push decisions toward the injected answer, and game accuracy drops further than those choices explain. Truthful correction brings the fact back but not the decisions built on it. We then look into the real-world bushfire case, models trained on the false posts say that people were arrested for arson even when they lose the inflated count, and in a count-by-wording factorial the misleading arrest wording produces arrest assertions even when the training count stays at 24. In simpler terms, \textbf{a correct factual answer does not guarantee a correct decision, and losing the injected number does not remove the misleading story}.
cs.CL / 19 / 2610.02926
Output Language Confusion under Multilingual Prompt Contamination
Riju Marwah, Ritvik Garimella, Khusham Bansal, Atishay Jain, Amit Sheth
cs.CL
Abstract
Standard factual benchmarks assume clean monolingual prompts and exact-match scoring, two assumptions that break simultaneously in real-world multilingual deployment, from retrieval-augmented generation pipelines returning mixed-language passages to users pasting multilingual web content. We introduce Multilingual Distractor Interference (MDI), a lightweight and fully replicable evaluation protocol requiring no new data or annotation, in which factual questions are preceded by a semantically irrelevant foreign-language sentence, and evaluate five instruction-tuned LLMs across TruthfulQA and TriviaQA under eight distractor conditions (40,000 evaluations). Our central finding is a metric confound: for Llama-3.1-8B under a Hindi distractor, 58% of responses switch to Devanagari script, yielding a raw hallucination proxy of 0.710, but manual review reveals that 120 of 148 script-switched responses that were correct under clean conditions remain semantically correct despite being written in the wrong script, reducing the adjusted semantic hallucination rate to 0.470. All other models respond through abstention escalation with no hallucination increase. A paragraph-length English distractor triggers near-universal abstention (0.806-0.998) across all models, consistent with reading-comprehension confusion, a failure mode with direct consequences for multilingual RAG pipelines. TruthfulQA multiple-choice accuracy is unaffected under all single-sentence conditions. These results show that exact-match hallucination rates in mixed-language settings should be decomposed into script-switching and semantic error components before drawing conclusions about model reliability.
cs.CL / 20 / 2610.03002
Recursive Self-Improvement in Unified Multimodal Models
Huijuan Wang, Chufan Shi, Cheng Yang, Yaokang Wu, Taylor Berg-Kirkpatrick, Xuezhe Ma
cs.CL · cs.CV
Abstract
Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.
cs.CL / 21 / 2610.03077
Unmasking Propaganda: A Comparative Analysis of Masked and Causal Language Models
Claudiu Creanga, Ioachim Lihor, Liviu P. Dinu
cs.CL
Abstract
Propaganda detection is an essential task in natural language processing (NLP), particularly in the context of manipulative political communications. However, identifying specific propaganda techniques presents a significant challenge due to their often subtle nature and reliance on context, making them difficult to distinguish from legitimate persuasive language. Propaganda often involves highlighting certain facts while downplaying or ignoring others to create a desired perception. This biased communication aims to influence attitudes, beliefs, or behaviors towards a particular cause or position. This paper explores advances in detecting propaganda techniques through a comparative analysis of modern language models, using the SemEval-2020 Task 11 dataset. We evaluated both masked language models (based on XLM-RoBERTa or DeBERTa V3) and causal models (from OpenAI, Google, Mistral, Anthropic and Meta), employing two prompting strategies: base and chain-of-thought prompting. Our results demonstrate improvements over state-of-the-art models, with the best-performing MLM achieving an F1 score of 63.18 in technique classification and the best causal model achieving 63.62. We also observed that certain models excel in specific techniques, such as loaded language and name-calling, while struggling with others like bandwagon and black-and-white fallacy. These findings suggest that fine-tuning, ensemble modeling, and the use of larger datasets can further enhance propaganda detection capabilities.
cs.CL / 22 / 2610.03078
An automated pipeline for standardised speech-unit annotation in spontaneous dialogue
Hanlu He, Harald Vilhelm Skat-Rørdam, Ingvi Örnólfsson, Ivana Konvalinka
cs.CL
Abstract
Quantifying conversational dynamics requires reliable identification of interactional units and their temporal boundaries, but speech activity alone does not distinguish conversational turns from listener feedback or within-turn pauses. We present an automated pipeline for extracting turns and backchannels from separate-channel recordings of spontaneous dyadic conversation, designed to provide a consistent first-pass annotation for subsequent human review. The pipeline combines voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post-processing. We evaluated the pipeline on 99 ten-minute Danish conversations from 33 dyads using segment-level detection reliability and temporal boundary error. Conversations were recorded under both normal and asymmetric listening conditions. In the latter, speech-shaped noise was delivered to one participant through bone-conduction headphones. Overall detection reliability was F1=0.621, with similar performance for turns F1=0.624 and backchannels F1=0.618. For successfully matched segments, median absolute onset and offset errors were 0.150 and 0.160s for turns and 0.130 and 0.180s for backchannels, respectively. Mean errors were substantially larger for turn boundaries, indicating a smaller number of large boundary mismatches. Performance did not differ significantly across the two experimental listening conditions. In a four-conversation case study, pipeline-human agreement was lower and more variable than human inter-annotator agreement and varied across parameter settings. These results support the pipeline as an automated first pass within a semi-automated annotation workflow, providing a consistent basis for more standardised and reproducible annotation of conversational dynamics.
cs.CL / 23 / 2610.03102
Ask, Relax, or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning
Ang Li, Yue Lin, Feifei Kou, Zhan Su, Prayag Tiwari, Wenhao Li, Shuhui Zhu, Hongyuan Zha, Baoxiang Wang
cs.CL · cs.AI
Abstract
An LLM agent can recognize uncertainty yet still choose the wrong next step: asking when action is already justified, or seeking clarification when the constraints must change. We formalize actionable indeterminacy: act when an accepted action is shared across all admissible preferences or objectives, clarify when each possibility is feasible but no action is shared, and propose a minimum-cost permitted constraint repair when the request is infeasible. We construct a solver-grounded benchmark spanning object allocation, meeting scheduling, apartment choice, and stable matching. Matched pairs retain the same source while changing whether intervention is necessary, and evaluation separates decision correctness, matched-pair reliability, and fully correct responses. Our findings reveal a recurring difficulty in recognizing when intervention is unnecessary: models can identify situations requiring clarification or repair yet still intervene when a justified action already exists. Correct decision labels also fail to guarantee usable actions, questions, or repairs. Crucially, response requirements shape not only how decisions are expressed but also which decisions are made. Making the required content explicit substantially improves fully correct responses and can change intervention decisions, even when outputs are already parseable. These findings highlight that reliable agency requires more than recognizing uncertainty: it requires intervening only when necessary and translating the chosen next step into a verifiable response.
cs.CL / 24 / 2610.03109
Emergent Structure in the Marginal Attention Space of Language Models
Valentino Maiorca, Walter Nelson, Francesco Locatello
cs.CL
Abstract
While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping them into a joint token-head "marginal attention space". Evaluating across 60+ diverse LLMs, we find that different properties emerge when reducing this space along its token and head axes. When reduced token-wise, marginal attention yields a text-intrinsic signal robustly conserved across models. To explain this property, we empirically connect marginal attention to the input-output Jacobian of the network, and prove theoretically that under a smoothness assumption, models with similar next-token distributions are guaranteed to have similar input-output Jacobian statistics. When reduced head-wise, it forms a model-private signature conserved across documents. Practically, this provides a natural way to estimate a per-head budget for key-value (KV) cache eviction, effectively decoupling model-specific budget allocation from text-intrinsic token scoring. On standard eviction benchmarks, a per-head budget precomputed offline on pretraining text, combined with a training-free token score, shows competitive performance with methods that recompute the budget on every document or train it per target. Code available at https://github.com/Flegyas/marginal-attention
cs.CL / 25 / 2610.03110
Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text
Claudiu Creanga, Liviu Dinu
cs.CL
Abstract
Supervised AI-text detectors report high benchmark accuracy, but it is not clear what their decisions are based on. We analyze a RoBERTa-based detector under semantic, structural, and tokenizer-level perturbations, using the M4 dataset (N = 10,000) and controlled generations (N = 300). When Mistral-7B-Instruct was asked to make machine text sound more human, Verb Diversity rose from 0.77 to 0.92 and the outputs became easier to detect. Detection scores appear to track statistical complexity, which also leads to a 76.3% false-positive rate on formal human writing. As a control, we evaluate event-based Latent Space detection. Paraphrasing changed 87% of its event sequences (Jaccard = 0.067), and homoglyphs altered 70% of the extracted verbs even though extraction still ran (Jaccard = 0.30). Its best domain AUC was 0.577. RoBERTa's robustness seems specific to the features it uses, and structural abstraction did not make detection more robust.
cs.CL / 26 / 2610.03112
Building Interpretable Feature Representations for Resume-Vacancy Matching by Distilling Production LLM Signals
Ilya Chekin, Vyacheslav Malyugin, Vladimir Chirkov, Mikhail Yurushkin
cs.CL
Abstract
Matching candidates to vacancies is central to recruitment, and a recruiter needs to see why a candidate fits, not only a single opaque relevance score. We provide this evidence as named, interpretable matching dimensions recruiters can act on - eight in our current deployment. We propose a two-part approach. The first is an LLM-based labeler whose prompts and feature definitions were refined from recruiter feedback while it served as an earlier production matching stage. In the current architecture, it is used only for offline labeling and is not called on online requests. The second is a feature bi-encoder distilled from it: a LoRA-adapted embedding backbone with compact per-dimension heads that runs on CPU and serves all online requests. Both parts keep improving: prompts are revised as feedback arrives, and the bi-encoder is retrained on the updated labels. The model is trained on 168,772 labeled vacancy-resume pairs (17,921 vacancies and 180,030 resumes). Recruiters using the service can confirm or revise surfaced feature predictions. On 927 recruiter-recorded values from this selected production-feedback subset, the deployed student agrees with the recorded decisions in 888 cases (95.79%). This is operational, non-blinded agreement rather than an independent human evaluation.
cs.CL / 27 / 2610.03176
Hindsight-Guided Rationale Distillation for Rare Disease Diagnosis
Aarav Singh, Animesh Pathak, Navyansh Singh
cs.CL
Abstract
We study hindsight-guided distillation for rare disease diagnosis on ZebraMap: a 1.5B student is fine-tuned on chain-of-thought traces from a 8B teacher that observes the ground-truth diagnosis during generation. Absolute accuracy remains low for all models - the task is hard at this scale - but within this ceiling a filtered variant (StudentF) achieves a small, statistically significant accuracy advantage over the teacher (p < 0.001), concentrated in better-represented diseases. The unfiltered student does not significantly outperform the teacher (p = 0.129), establishing that contamination filtering - not hindsight distillation alone - drives the gain. The gap traces to an artifact we term GT hallucination. Label-visible generation causes the teacher to embed "ground truth is X" phrases in its reasoning chain; SFT copies the pattern. At inference, the unfiltered student reproduces the phrase in 33.9% of cases, with severe accuracy degradation when the hallucinated label is wrong. A regex filter removing these slots reduces contamination to near-zero, producing the observed gain - though the effect remains small. We precisely quantify this gain-cost tradeoff, document frequency-dependent knowledge transfer absent from the RL-trained teacher, and characterize a calibration gap that SFT does not close - identifying both as directions for future work.
cs.CL / 28 / 2610.03190
Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case
Tingzhu Bi, Ping Wang, Meng Ma
cs.CL · cs.AI · cs.LG
Abstract
Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is "cause undetermined". Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.
cs.CL / 29 / 2610.03195
Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song, Yohan Jo
cs.CL
Abstract
As LLM agents decide on users' behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-to-end search with 12 agent models across three domains. Comparing items from different sources that satisfy the same requirements at the same position, we find that each model prefers some sources and avoids others in every domain, largely agreeing on which. This preference can outweigh how well items satisfy the request: an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better one from a dispreferred source, but almost never in the reverse case. The information identifying an item's source affects selection by itself: hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate. We test two routes to this preference: training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source. Supplying missing information or a prompt countering these preconceptions reduces source preference.
cs.CL / 30 / 2610.03482
Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models
Mehrdad Ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani, Sadra Hakim, Audrina Ebrahimi
cs.CL
Abstract
Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, using Gaokerena-V and Gaokerena-R as two previously developed backbones. We first examine response variability on a 168-question Iranian medical entrance examination and observe substantially lower five-run consistency for Gaokerena-V than for Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. We then construct two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and train lightweight claim-level heads on frozen backbone attention maps and token probabilities. On held-out test splits, the heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The heads require neither retrieval nor repeated sampling at inference time. These results provide an initial study of single-pass claim-level uncertainty estimation for Persian medical language models; the test splits are small and the labels are automatically generated.
cs.CL / 31 / 2610.03515
Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation
Chenglei Shen, Haoyang Yao, Weijie Yu, Song Jin, Xiao Zhang, Jun Xu
cs.CL
Abstract
On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student's own reasoning to provide guidance that addresses its specific errors while building on valid progress. For each failed attempt, RC-OPD locates the earliest substantive error, develops a local correction, and uses the corrected intermediate result as an anchor for the valid prefix. An iterative diagnosis--repair--continuation process tests the repairs through student continuation, identifying further errors within a fixed repair budget. For repair chains that reach a correct answer, root--cause--guided distillation uses failure diagnoses and corrective goals to supervise the erroneous segments, while anchor-guided distillation supports the corresponding valid prefixes with reasoning chains leading to the repaired intermediate results. We evaluate RC-OPD across multiple datasets and model scales. Extensive experiments and analyses show that it mitigates reasoning mismatch and the distillation trap, yielding substantial performance gains.
cs.CL / 32 / 2610.03531
Author Representation Strategies for Zero-Shot Authorship Attribution: A Comparative Study of LLM-Based and Embedding-Based Approaches
Nudrat Habib, Tosin Adewumi, Sana Sabah Al-Azzawi, Marcus Liwicki, Elisa Barney
cs.CL
Abstract
Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline together with three author representation strategies: representative writing samples, LLM-generated descriptions, and style embeddings (LISA). The first three approaches perform attribution using LLM prompting, while the embedding-based approach uses style embeddings with cosine similarity. We investigate the influence of prompt design and propose a two-stage embedding-based attribution framework that combines candidate space reduction with embedding-dimension selection. The results show that label-only ZS AA is ineffective, while incorporating author-specific representations consistently improves attribution performance. Among the evaluated approaches, the proposed two-stage LISA framework achieves the strongest overall performance, whereas LLM-generated style descriptions provide a substantially more compact representation of author style at the cost of some attribution performance. These findings demonstrate the importance of author representation in ZS AA, while indicating that current open-source LLMs remain insufficient for robust attribution without more effective representation learning.
cs.CL / 33 / 2610.03565
Writerslogic at PAN 2026: Process over Content for Robust Detection under Domain Shift
David L. Condrey
cs.CL
Abstract
We describe the Writerslogic systems for three PAN at CLEF 2026 shared tasks (Reasoning Trajectory Detection, Voight-Kampff Generative AI Detection, and Multi-Author Writing Style Analysis), unified by a shared analytical framework: feature robustness under distribution shift is governed by support overlap between training and test distributions, not by training-set effect size. This yields a taxonomy (domain-anchored, domain-portable, domain-invariant) that explains why generator-specific features die under domain shift while vocabulary fingerprints (hapax ratio, Yule's K, Heaps' exponent), compression measures, and character n-grams survive. On Reasoning Trajectory Detection, where training was entirely mathematics and 84 percent of test was unseen domains, the framework guided system design to 1st place in source detection (0.85 macro F1 via Opus-Sonnet agreement) and 3rd place in safety classification (0.66 macro F1 via query-refusal decomposition). For Voight-Kampff, we built a calibrated ensemble of DeBERTa-v2 (ONNX), multi-seed LightGBM with 44 domain-portable stylometric features, and SVM on n-gram TF-IDF, combined via learned stacking with isotonic calibration; the best configuration achieved 0.891 on the PAN 2026 test set with balanced sub-metrics (0.853 to 0.902 across all evaluation dimensions). For Multi-Author Writing Style Analysis, we describe a system fusing spectral clustering over character n-gram similarity graphs, normalized compression distance for local boundary detection, and SmolLM-135M perplexity for neural change-point detection; a platform mix-up meant our run never reached the official evaluation, so we report the design and its a priori predictions. Across all three tasks, features measuring generation process properties are designed to outperform features measuring generated content properties under domain shift.
cs.CL / 34 / 2610.03567
Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting
David L. Condrey
cs.CL
Abstract
We describe the Writerslogic team's participation in the CLEF 2026 SimpleText shared task, addressing Task 1 (text simplification) and Task 2 (complexity spotting). For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying temperatures, then selects the best candidate using a reference-free scoring heuristic that rewards compression, source word retention, Cochrane Plain Language Summary vocabulary usage, and lexical simplicity. On Task 1.1 (sentence-level simplification), our Claude Sonnet 4 submission achieves SARI 47.43 and BLEU 14.21, the top-ranked sentence-level system (3rd on the combined Task 1 leaderboard, behind two document-level submissions). For Task 2, we fine-tune a DeBERTa-v3-large NLI model on 350K labeled (source, sentence) pairs, framing hallucination detection as natural language inference. The model reads the most relevant source sentence as premise and the candidate as hypothesis, directly learning to distinguish grounded from hallucinated content. On Task 2.1 (binary overgeneration identification), our fine-tuned DeBERTa system achieves 0.8081 document-level macro F1 (0.8085 in our best ensemble), the top-ranked entry within the identification track and 2nd among teams overall, behind AIIR Lab (0.8197). On Task 2.2 (multi-class error classification), our best submission reaches 0.804 multiclass accuracy, ranking 2nd among unique teams behind AIIR Lab (0.827). We evaluate both tasks on English and multilingual biomedical text from Cochrane systematic reviews.
cs.CL / 35 / 2610.03625
FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs
Darian Lee, Shannon Rumsey, Jack St. Clair, Xinyi Tang, Aditya Bansal, Yuanming Shi
cs.CL · cs.DB · cs.LG
Abstract
Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely ignore this ambiguity and produce oversimplified queries that fail to prepare models for the complexity of real-world structured knowledge access. We present FALCON, a framework that generates realistic, ambiguity-aware NL-to-SQL data matching the complexity of challenging real-world benchmarks, at low cost using compact open models. Our approach combines reserved-word SQL seeding and persona-based prompting to generate structurally complex queries, while alignment-based filtering preserves difficulty by distinguishing genuinely incorrect examples from complex but valid queries. Human evaluation confirms consistent high quality across model sizes, and our generated data exceeds existing benchmarks in both SQL complexity and natural language richness. Difficulty-stratified analysis shows models trained on FALCON data increasingly outperform baseline-trained models as query complexity increases, validating our pipeline's success in generating challenging training data. When combined with a small proportion of existing benchmark data, mixed training recovers performance on simpler queries while preserving these advantages on complex ones. The model- and database-agnostic design enables organizations to generate high-complexity NL-to-SQL training data locally without external APIs.
cs.CL / 36 / 2610.03695
Language Models that Play Chess and Explain Their Moves
Adithya Bhaskar, Jeffrey Cheng, Danqi Chen
cs.CL
Abstract
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.
cs.CL / 37 / 2610.03022
ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation
Sihan Ren, Gaozheng Li, Yuanshang Quan, Yiming Qin, Fuyi Yang, Chang Liu, Lan Xu, Minye Wu
cs.CV · cs.CL
Abstract
Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.
cs.CL / 38 / 2610.03632
World Embedding Benchmark
Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li, Fengyu Cai, Hou Pong Chan, Jialin Yu, Hao Zhang, Chenghua Lin, Chenghao Xiao
cs.CV · cs.CL
Abstract
Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.
cs.CL / 39 / 2610.02386
Social bot detection in the age of ChatGPT: Challenges and opportunities
Emilio Ferrara
cs.CY · cs.CL
Abstract
We present a comprehensive overview of the challenges and opportunities in social bot detection in the context of the rise of sophisticated AI-based chatbots. By examining the state of the art in social bot detection techniques and the more salient real-world application to date, we identify gaps and emerging trends in the field, with a focus on addressing the unique challenges posed by AI-generated conversations and behaviors. We suggest potentially promising opportunities and research directions in social bot detection, including (i) the use of generative agents for synthetic data generation, testing and evaluation; (ii) the need for multimodal and cross-platform detection based on network and behavioral signatures of coordination and influence; (iii) the opportunity to extend bot detection to non-English and low-resource language settings; and, (iv) the room for development of collaborative, federated learning detection models that can help facilitate cooperation between different organizations and platforms while preserving user privacy.
cs.CL / 40 / 2610.02673
Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories
Joseph Chee Chang, Michael D'Arcy, Amy X. Zhang, Pao Siangliulue, Sangho Suh, Aakanksha Naik, Jena D. Hwang, Javier Ramos Benitez, Stella Wroblewski, Matt Latzke, Michael Cuoco, Ruben Lozano-Aguilera, Kris Ganjam, Joel Chan, Doug Downey, Peter Jansen, Kyle J. Travaglini, Daniel S. Weld
cs.HC · cs.CL · cs.DL · cs.IR
Abstract
A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research questions. Recent approaches scale theory synthesis with LLMs, but automate away choices and intuitions from researchers. We present Asterism, which extracts observations from hundreds of papers as concept-relation triples, with concepts unified in a hierarchical ontology. Researchers curate an evidence graph using the ontology and aggregate observations at different levels of granularity to focus theory formation on specific phenomena of interest. In a field deployment (n=10), researchers worked from observations to theories, and kept concepts and hypotheses fitting their preferences. In two case studies, teams of immunology and agriculture researchers discovered mechanisms outside their standard analyses and constructed hypotheses worth follow-up experiments.
cs.CL / 41 / 2610.03130
Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study
Yun Wang, Gad Shaulsky, Tomaž Curk, Blaž Zupan
cs.IR · cs.CL
Abstract
Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval benchmark from dictyBase for Dictyostelium, a model organism in cell and developmental biology. The benchmark consists of curator-generated biological queries linked to PubMed-indexed articles, together with structured gene annotations. Using this benchmark, we study three factors in niche biological retrieval: cross-encoder reranking, gene-aware query expansion, and abstract-only versus full-text retrieval. We report that reranking and gene-aware query expansion improve retrieval selectively: reranking is most useful when the model is well suited to biological evidence matching, whereas curated annotations help clarify compact biological queries by reducing vocabulary mismatch. Full-text chunks substantially improve retrieval when abstracts omit supporting evidence, increasing both candidate recall and top-rank performance, although these cases are harder than queries supported by abstracts. Data and code are publicly available at https://github.com/fulaibaowang/dictycite, and the benchmark dataset is additionally archived on Zenodo.
多智能体系统 (cs.MA)
8
cs.MA / 1 / 2610.03394
EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures
Yuhai Long, Yuanxin Wei, Kai Wu, Jinhui Wei, Dan Huang, Jiangsu Du
cs.DC · cs.MA
Abstract
Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU co-execution. Furthermore, speculative decoding in multi-agent workloads faces extreme variance in drafting difficulty, alternating between complex reasoning and predictable structured generation. Compounded by frequent tool-induced stalls, this highly fragmented execution severely underutilizes hardware and defeats traditional static batching. We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. At the micro-architectural level, it bypasses rigid graph-compiler constraints to enable zero-copy UMA-aware tensor parallelism, utilizing asymmetric memory layouts to fully saturate both CPU and GPU compute units. At the scheduling level, it dynamically allocates draft budgets based on real-time sequence predictability to bound bandwidth waste. Concurrently, an asynchronous suspend-and-yield mechanism actively evicts stalled agents, ensuring continuous hardware saturation during unpredictable tool invocations. Extensive evaluations on an Apple M4 SoC demonstrate that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding. Adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x speedup under extreme tool-use latencies.
cs.MA / 2 / 2610.02847
Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning
Amit Thakur, Mukesh Singhal
cs.MA · cs.LG
Abstract
Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standard centralized critics and shared advantages often mix these two effects into one scalar credit signal, allowing surviving agents to be rewarded or penalized for exogenous turnover events outside their control. We introduce turnover-orthogonal credit assignment (TOCA), a value decomposition for open teams that separates action effects, pure turnover effects, and action--turnover interactions. Under exogenous turnover, the event-conditioned value admits a centered decomposition whose event-conditioned baseline removes the pure turnover component while preserving credit for actions that make the team robust to future replacements. We instantiate this idea with a permutation-invariant centralized critic over variable-size agent sets and event tokens, and derive both a counterfactual per-agent credit signal and a softly weighted interaction variant, TOCA-$β$, for high-variance control environments. Controlled diagnostic experiments show that TOCA improves return over event-aware MAPPO-style critics and that removing interaction credit substantially hurts performance. In a replacement-only Dynamic Spread benchmark, TOCA-$β$ achieves the best mean return at high turnover rates and improves over its no-interaction ablation. These results suggest that explicitly separating turnover from action credit is a useful principle for robust learning in dynamic cooperative teams.
cs.MA / 3 / 2610.02924
Evaluator-in-the-Loop Monte Carlo Tree Search via LLM Agents for Motif Scaffolding in Protein Design
Haotian Hu, Oguzhan Gungordo, Siheng Xiong, Faramarz Fekri
cs.MA
Abstract
Motif-scaffolding systems commonly follow a generate-then-filter paradigm, in which candidate proteins are generated independently and structural evaluation is used primarily for terminal screening or ranking. This paradigm underuses evaluation: failed predictions contain state-specific evidence about whether a design requires repair of motif geometry, global foldability, or other structural constraints. We introduce \textbf{ELMS} (Evidence-based LLM-guided Monte Carlo Search), an evaluator-in-the-loop search framework for motif scaffolding that turns such evaluator feedback into targeted design actions. Effective reuse of structural feedback is nontrivial because different scaffold states exhibit different failure modes, and repeatedly refining a single trajectory can prematurely commit computation to an unproductive region of sequence space. ELMS therefore retains evaluated scaffolds as persistent search states: a Critic Agent diagnoses state-local structural failures, a Policy Agent selects targeted operators with execution parameters, motif-locked operators realize legal sequence modifications, and MCTS determines which historical states should receive further design effort. Under the standard GeomMotif protocol (100 candidates per task), ELMS achieves Successful rates of 86.41\% on single-motif tasks and 84.57\% on paired-motif tasks, exceeding the strongest prior baseline by 19.3 and 21.9 percentage points, respectively. On MotifBench, under a matched 100-candidate search budget, it solves 26.7 of 30 tasks on average (88.89\% Task Success), compared with 16.0 tasks (53.33\%) for the strongest baseline. These results establish ELMS as an effective approach for converting structural evaluation from a terminal filter into actionable guidance for iterative motif scaffolding.
cs.MA / 4 / 2610.03174
FinNextAssist: Towards Professional Financial Deep Research Assistant
Xiangyu Li, Fengbin Zhu, Xuan Yao, Siyu Liu, Xiaoluan Liu, Chao Wang, Huanbo Luan, Xiaofen Xing, Xiangmin Xu, Ke-Wei Huang, Richang Hong, Tat-Seng Chua
cs.MA · q-fin.CP
Abstract
Deep Research (DR) agents have demonstrated strong capabilities in complex, research-oriented tasks through autonomous planning, iterative retrieval, multi-step reasoning, and structured reporting. However, adapting DR agents to finance introduces unique challenges: financial analysis demands the joint completion of heterogeneous sub-tasks spanning diverse data types, tools, and analytical workflows. We identify three key requirements for a professional financial DR agent: integration of authoritative, heterogeneous financial data sources; specialized analytical tools and skills; and dedicated sub-agents for domain-specific sub-tasks. Building on these principles, we propose FinNextAssist, an end-to-end deep research framework designed for professional financial analysis. FinNextAssist decomposes the research process into four stages: Task Planner, Evidence Compiler, Reasoning Engine, and Report Assembler, and introduces two novel lightweight sub-agents: TabAgent, for cross-market financial table understanding, and HeteroAgent, for cross-modality heterogeneous financial data interpretation. Extensive experiments on FinDeepResearch, the Finance Agent Benchmark, and FinTMMBench-Web show that FinNextAssist substantially outperforms both strong proprietary and open-source DR agents, with ablation studies confirming the contribution of each component across diverse markets and languages.
cs.MA / 5 / 2610.02694
Who Went Where When on the Lunar Surface: Forensic Trajectory Analysis to Identify Byzantine Rovers
Lachlan Holden, Feras Dayoub, Melissa de Zwart, David Harvey, Tat-Jun Chin
cs.RO · cs.MA
Abstract
Future planetary surface missions are likely to involve multiple independently operated rovers sharing the same deployment region, raising the need to verify compliance with operational constraints such as Lunar Safety Zones. Because continuous in-situ observability is rarely available, such verification requires post-hoc reconstruction of rover trajectories from sparse telemetry, including odometry, pose priors, and relative inter-rover detections. We introduce the problem of forensic trajectory analysis for non-cooperative planetary rovers in the presence of Byzantine agents: rovers that provide miscalibrated or deliberately falsified measurements to support an incorrect trajectory. We show that standard outlier-robust pose graph optimisation methods are vulnerable in this setting, because Byzantine rovers can generate measurements that are internally consistent and numerous enough to make truthful incriminating measurements appear as outliers. To address this, we propose an attribution-aware trajectory estimation method that reasons over rover credibility rather than individual measurement validity. The method evaluates candidate credible rover subsets by comparing the statistical consistency of their internal and boundary relative detections against provided priors, and then estimates trajectories using only measurements attributed to credible agents. Across synthetic simulations and real planetary-analogue trajectory data, the proposed method identifies Byzantine rovers and produces significantly more accurate trajectory estimates than existing robust pose graph optimisation baselines.
cs.MA / 6 / 2610.02874
SceneFactory-3D: Lifting 2D Traffic Scenes into 3D Physical Counterfactuals for Scalable Physically Grounded Safety Evaluation
Yicheng Zhu, Linfeng Tian, Tianmu Zhao, Yang Chen, Fan Zuo, Tao Li, Zilin Bian
cs.RO · cs.MA
Abstract
Scalable driving simulators typically execute vehicle commands using prescribed behavioral or kinematic rules, overlooking the physics of tire-road interfaces, thereby limiting their ability to capture how adverse road and environmental conditions alter vehicle execution and propagate through traffic. To address this limitation, we present SceneFactory-3D, a GPU-batched, physics-grounded multi-agent driving simulator. Vehicles execute acceleration and steering commands via suspension- and friction-limited forces evaluated at each wheel-contact point. Spatially varying friction, per-world 3D heightfields, gravity, and rigid contact consistently govern wheel motion and chassis collisions. Per-world terrain isolation and GPU batching enable SceneFactory-3D to run matched physical counterfactuals in parallel: traffic scenario setup and vehicle controllers remain fixed while only the road condition changes, enabling the resulting closed-loop effects to be evaluated across parallel worlds. To demonstrate the advantage of the SceneFactory-3D-enabled counterfactual evaluation, we conduct an empirical study on vehicle controllers' sensitivity to road conditions. We study three learned-policy families on 1,024 matched 12-vehicle worlds per condition, and two classical planners on a shared 32-world subset, across 21 friction and grade conditions. When friction drops from 1.0 to 0.18, the share of vehicles that clear the work zone safely falls by 6 to 90 percentage points across learned policies (18-19 for classical planners), and near-collision situations become more frequent for every learned policy. Code: https://github.com/SmallWorldLab/SceneFactory_3D
cs.MA / 7 / 2610.02645
Performance of Zero-determinant Strategies in Repeated Games without Discounting
Masahiko Ueda
physics.soc-ph · cs.MA · eess.SY
Abstract
Zero-determinant (ZD) strategies are a class of strategies in repeated games, which unilaterally control payoffs. It has been shown that several ZD strategies promote cooperation in social dilemma games. It has been widely believed that the performance of ZD strategies is expressed as ``unilaterally enforce linear relations between payoffs''. However, in order to interpret properties of ZD strategies in repeated games without discounting, previous studies assumed that the limit of the Cesaro average of the probability distribution of the action profile exists. Here, we explain the performance of ZD strategies without this assumption, which leads to a stronger result than previous ones.
cs.MA / 8 / 2610.02863
Multi-Agent AI as a Nested Principal-Agent Problem in Private Wealth Management: Mandate Representation and Evidence Control in Switzerland, Germany and Austria
Walter Kurz, Reinhard Magg, Florian Kollberg, Wojtek Stricker, Stefan Marx, Frank Reinhardt, Velimir Dedić
q-fin.GN · cs.MA
Abstract
In private wealth management, a manager delegating to artificial intelligence (AI) acts as the client's agent and the system's principal. We introduce a model-independent formulation that combines nested principal--agent delegation with constrained joint maximisation as the task assigned to the AI system. The objective represents client and manager outcomes separately over portfolio--workflow pairs. Legal duties, mandate requirements and evidence sufficiency determine admissibility, with Switzerland, Germany and Austria supplying the legal context. Weights and reference-service floors make the trade-off explicit; concession accounting separates their effects on the client. Analytical constructions and a simulation using public-market observations illustrate the approach. Across eight decision states from four constructed mandates, omitted client liabilities caused two liquidity violations, omitted manager terms caused two capacity violations, and mistranslated weights changed four otherwise admissible choices under faithful optimisation. At the declared weights, six states selected a higher service tier than the client-best alternative, with client concessions of EUR 1,178 to EUR 2,264 and manager gains of EUR 3,062 to EUR 10,381. Three instruction forms each reached all 32 specified decisions under shared numerical, evidence and simulated approval controls; professional instructions matched explicit nested delegation on accuracy and clarification count. Subsequent 2022 exchange-rate and yield paths, combined with constructed growth scenarios, produced lower client outcomes than the reference service although the selected services met the decision-time forecast benchmarks. These examples suggest that the approach could help make mandate choices and their consequences easier to examine. Professional and field studies could assess whether this improves oversight and client outcomes.
软件工程 (cs.SE)
10
cs.SE / 1 / 2610.02503
Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents
Rudrendu Kumar Paul, Sourav Nandy
cs.SE · cs.AI · cs.DC · cs.LG
Abstract
Deploying compound AI systems reliably and safely requires understanding failure modes that emerge at component boundaries, not within individual models. Cascading errors propagate across component boundaries, silent quality degradation evades standard monitoring, and coordination failures yield incorrect collective behavior from individually correct parts. We analyze 150 production incident reports from open-source compound AI projects and anonymized enterprise deployments to construct a taxonomy of 23 failure modes organized into five categories: retrieval failures, generation failures, tool failures, orchestration failures, and integration failures. For each category, we propose resilience patterns with measured effectiveness from controlled fault injection experiments. Circuit breakers reduce cascade propagation by 89%, output quality gates catch 73% of silent degradation before user impact, and component isolation reduces blast radius by 64%. Systems implementing three or more resilience patterns from our catalog reduce mean-time-to-recovery (MTTR) by 71% compared to unstructured monitoring baselines. We release the incident taxonomy and pattern catalog as a practitioner resource.
cs.SE / 2 / 2610.02560
An Extensive Empirical Study on Evaluation Metrics for Combinatorial Interaction Testing
Lisha Qin, Chenhui Cui, Tao Li, Rubing Huang, Shikai Guo, Lei Ma
cs.SE
Abstract
Combinatorial interaction testing (CIT) is a black-box testing method that has received extensive attention in both research and practice over recent years. Its primary objective is to construct an effective combinatorial test suite that detects software failures caused by parameter interactions. As a fundamental component of the CIT testing process, the evaluation metric plays a critical role in assessing and comparing combinatorial test suites, as well as in evaluating various test generation techniques. For CIT practitioners, selecting an appropriate evaluation metric is both important and challenging, given the wide variety of available options. Nevertheless, no prior work has systematically addressed this problem. To fill this gap, this paper first provides a comprehensive survey of black-box evaluation metrics for combinatorial test suites, offering rigorous definitions, clear classifications, illustrative examples, and complexity analyses. We then conduct an extensive empirical study involving eight open-source projects, encompassing 32 test scenarios and 295,624 combinatorial test suites. In this study, we examine the correlation between each static evaluation metric and fault-detection effectiveness using two correlation measures. Experimental results show that the Value Combination Coverage (VCC) metric serves as a valid predictor for test-suite evaluation. However, distribution-based metrics generally incur lower computational costs than interaction coverage-based ones. The choice of an appropriate metric should also account for the test suite's inherent properties, as these characteristics can substantially influence the effectiveness of the evaluation. Finally, we provide practical guidelines to assist CIT practitioners in selecting suitable evaluation metrics for assessing or comparing combinatorial test suites.
cs.SE / 3 / 2610.02571
Improving the Energy-Efficiency of the Code Generated by LLMs through Effective Prompting
Ritika Rekhi, Bing Zhang, Md Arman Islam, Jaya Krishna Pasham, Yeswanth Chitturi, Akshay Paramesha, Isha Valiveti, Asif Imran, Bekir Turkkan, Tevfik Kosar
cs.SE · cs.AI
Abstract
As AI-assisted programming becomes increasingly mainstream, the environmental impact of AI-generated software has emerged as an important consideration. This motivates evaluating LLM-generated code beyond functional correctness by considering execution efficiency and energy consumption. However, despite substantial advances in code generation, frontier LLMs are rarely evaluated based on the energy efficiency of the code they produce. In this work, we conduct a comprehensive evaluation of 21 prompting strategies for energy-efficient code generation and identify 8 strategies for evaluation across 10 widely used open-weight and proprietary LLMs. We evaluate their effectiveness for both Python and C++ code generation relative to a baseline prompt. Across the evaluated models, the selected prompting strategies achieved energy reductions of up to 25% for Python and 17% for C++ code generation. At the model level, Python energy reductions reached up to 50% for Granite-4.0-H-Small, 39% for Claude 4.5 Haiku, and 28% for MiniMax M3, while C++ reductions reached up to 56% for Granite-4.0-H-Small and 7% for Qwen3-Coder-480B-A35B-Instruct. These results demonstrate that prompting strategies can substantially influence the energy consumption of LLM-generated code, although their effectiveness varies across models and programming languages. Our findings highlight the importance of incorporating energy efficiency into the evaluation and optimization of LLM-based code generation and provide practical insights into designing prompts for more sustainable AI-assisted programming.
cs.SE / 4 / 2610.02617
WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness
Yun-Yun Tsai, Yuning Mao, Shiqi Wang, Junfeng Yang, Sinong Wang
cs.SE · cs.AI · cs.MA
Abstract
Evaluating WebUI code generation at scale is difficult: outputs may compile and look plausible yet fail under user interaction, and prior benchmarks largely rely on free-form prompts with static checks (build success, screenshots) that miss functional correctness. We introduce WebUIProof, an execution-oriented benchmark that provides structured specifications and dense, executable interaction tests for WebUI generation across two task families: general WebUIs (e.g., dashboards, game, interactive tools) and 3D interactive simulation (e.g., particle/galaxy systems, physics dynamics). WebUIProof includes a UI-agent harness that runs executable interaction tests in a headless browser using an iterative plan--act--observe loop: it locates DOM elements, performs actions, observes resulting UI/DOM changes, and checks the specified assertions. We evaluate across eight commercial LLMs and observe frequent failures on interaction-based requirements even when pages render successfully, especially on 3D simulation interfaces. Finally, we show the UI-agent harness can provide outcome-level training signals. Training compact models (e.g., Qwen2.5 14B and MIMO 7B) with RL rewards derived from executable interaction tests improves functional completion while reducing build failures.
cs.SE / 5 / 2610.02710
Self-Supervised Scaling of Terminal Environments for Scientific Domains
Zhongzhi Li, Yucheng Shi, Zongxia Li, Junyao Yang, Ruhan Wang, Yu Wang, Jingyuan Huang, Jichao Yu, Ninghao Liu, Haitao Mi, Leowei Liang
cs.SE · cs.AI
Abstract
Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Authoring these components for each task requires repeated engineering and limits reuse. We introduce software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs. For each workflow, we execute multiple input configurations and partition cases into public observations and hidden evaluations. Given the instruction, input schema, and public input--output observations, an agent constructs an editable program without access to the source workflow. The candidate is evaluated on hidden configurations against workflow outputs. A hierarchical verifier combines domain-specific semantic comparison, structural validity, and anti-shortcut checks, while public feedback supports iterative revision. The construction admits additional workflows and configurations without authoring a reference solution for each task. We instantiate SWR with 500 workflows and 46 software families across six domains. Across three attempts per task, Qwen3.8-Max solves 838 tasks and produces 1,422 verified trajectories, which we oversample to 3,000 reconstruction-only training examples. Supervised fine-tuning of Qwen3.8-27B improves mean Terminal-Bench 2 performance from 47.94% to 53.56% across three seeds and achieves the highest mean among four matched-token corpus controls on all four reported evaluations. These results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents.
cs.SE / 6 / 2610.02814
VeriPy Source-Preserving Verification and Compatibility Checking for Python Components
Naing Oo Lwin
cs.SE · cs.LO · cs.PL
Abstract
Keeping a Python program and its formal guarantees aligned is a continuing maintenance problem. Specifications must describe the implementation that actually runs, and updates must preserve the behavior on which existing callers depend. VeriPy brings these obligations into a common workflow for annotated Python components. Developers and agents express contracts, invariants, and proof hooks as Python comments, then refine checked auxiliary lemmas using source-located diagnostics. Direct encoders produce Dafny or Lean artifacts while retaining admitted executable bodies and explicit models of their dependencies. A relational product checks whether an updated component still admits old inputs and preserves returned values and modeled exceptions. The resulting workflow connects source-level proof development with functional verification and backward compatibility, recording the assumptions under which each guarantee applies. The code is available at https://github.com/astrio-labs/veripy.
cs.SE / 7 / 2610.02928
Discriminating Fixture Coverage in Agent-Infrastructure Verification Suites
Xin Xu, Siru Tao
cs.SE · cs.AI · cs.LG
Abstract
Invariant suites and runtime monitors increasingly gate agent deployment decisions, and the evidence offered for any particular suite is almost always a single observation: it passes an implementation believed correct and fails one believed broken. We measure what that observation is worth. Applying mutation analysis to an invariant suite for a multi-session agent state-projection layer, we first find that this standard validation certifies a suite in which a first-order mutant removing event-identity deduplication survives every check. We then freeze the repaired twelve-check suite, record its hash, and run it once against ten mutants specified by an adversarial reader who designed none of its fixtures: it kills five. Instrumenting the five survivors against the reference shows they fail in two distinct ways, not one. Three are never activated, because no fixture supplies an input on which the mutated code behaves differently at all. The other two corrupt internal state that no oracle in the suite can observe. The two modes need different repairs, and neither is visible from a pass/fail report. Treating the missing inputs as a coverage question, we enumerate seven discriminating dimensions of the input space, register in advance which are uncovered and which survivors they should explain, and add one fixture per uncovered dimension while reusing the existing oracles verbatim. All five survivors then die, each to the check written for its predicted dimension. We report this as a repair result on the same challenge set rather than a second held-out estimate, and give the artifact, including the frozen hash, the registered predictions, all mutants and the run logs, so the distinction is checkable.
cs.SE / 8 / 2610.02952
GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing
Masahiro Kato
cs.SE · cs.LG · stat.ME · stat.ML
Abstract
Test-driven development gives AI coding agents executable requirements for implementing software. Because these agents can adapt their implementations to the examples they observe, passing a predetermined collection of tests can leave substantial parts of the intended behavior unimplemented. We propose Generative Test-Driven Development (GTDD), a formulation of test-driven development in which a separate testing agent generates new inputs after each candidate implementation is fixed, using a human-specified behavioral contract and the feedback from earlier rounds. A trusted evaluator checks these inputs, returns reduced counterexamples to the coding agent, and saves them for regression testing, so development continually confronts failures beyond the initial examples. We characterize the evidence that this process provides through a finite-population analysis of false acceptance under adaptive candidate selection. The resulting bounds quantify how test visibility and repeated feedback affect acceptance, and show that fresh random audits after candidate commitment control false acceptance across development rounds. In a paired experiment on a stateful key-value store, both policies that regenerated tests during development ended with lower mean failure rates than the policy whose tests were generated once by the same language model, and giving the tester the candidate's source produced no detectable additional improvement. Further conditions requesting equal numbers of tests did not isolate any single feature of the policies as the source of this difference. GTDD combines this adaptive development feedback with established regression tests and an independent acceptance rule.
cs.SE / 9 / 2610.02995
BISCEPTER: Probability-Driven Bisection for Large-Scale System Software
Mingyan Gao, Celine Wüst, Zuming Jiang, Zhendong Su
cs.SE
Abstract
Identifying the bug-inducing commit (BIC) is a fundamental step in regression debugging and a key input to emerging BIC-aware fault-localization pipelines. In practice, BICs are commonly obtained with bisection. Standard bisection selects the median commit of the remaining good-bad interval, thereby balancing commit count. This strategy is optimal under the assumption that each commit is equally likely to be the BIC. This paper shows that this assumption does not match real-world BIC histories. We construct a dataset of 8,172 bug reports from GCC, the Linux kernel, and MariaDB. We find a strong temporal skew: across the studied systems, 50% of BICs lie within the most recent 0.69% of the report-time commit history. Motivated by this observation, we introduce BISCEPTER, a probability-driven bisection approach that uses historical BIC latency as a lightweight prior. Instead of selecting pivots that split the number of remaining commits, BISCEPTER selects weighted-median pivots that split estimated BIC probability mass, while preserving the same good-bad oracle and interface as standard bisection. We evaluate BISCEPTER on three large-scale systems. The evaluation results show that BISCEPTER reduces bisection iterations by 25.75% on average (up to 55.55%) compared with standard median bisection, while improving over the baseline in 91.26% of test cases. Robustness experiments further demonstrate that the benefit remains stable under noisy historical data. We expect that our research can effectively save effort in debugging software in practice and, more broadly, benefit future software engineering research by bringing insights about BIC distribution.
cs.SE / 10 / 2610.03237
Pinning Decisions Before Failure: Executable Records of Underspecified Choices in AI-Assisted Code Generation
Takeharu Mitsui
cs.SE
Abstract
A natural-language requirement leaves questions open, and a model asked to implement it settles them silently: across 600 generated test suites from three models, 42.7% contain no test that distinguishes the competing readings. An acceptance example written before implementing is the usual remedy, but an example both readings satisfy resolves nothing. We insert two steps into that practice: enumerate the requirement's underspecified points by name, then for each construct two throwaway implementations differing only in that point and keep a candidate input only if executing both shows they disagree. The result is recorded as a decision pin: the named point, the confirmed input, and the value the person chose between the two exhibited results. The same record then constrains generation and decides compliance by execution. On a benchmark of 40 tasks with paired reference implementations and two models, a separating input is obtained for 92.5% of decision points and a pin identifying the intended decision for 85-90%. As checks on 703 independently generated implementations, pins agree with the benchmark's classification on 96-97%, with disagreements on three tasks, one where both classifiers erred. As generation constraints, pins are honoured at the pinned input in all 210 generations and are never worse than a prose rule on held-out inputs in 38 cells, though no better than prose stating the same scope. Every compliance failure under a prose rule came from the model deciding the rule's scope itself; one such case silently overturned another recorded decision, was attributed to a single rule by leave-one-out on both models, and was missed by text-level reconciliation but caught by re-running the recorded input. The setting yields too few such conflicts to evaluate a regression step, and we say why.
操作系统 (cs.OS)
1
cs.OS / 1 / 2610.02676
FDP: The Data Placement Promise of Modern NVMe SSDs
Sijie Lan, Hui Qi, Xing He, Mahmut Kandemir, Javier González, Vivek Shah
cs.OS · cs.DB
Abstract
NVMe SSDs are now widely deployed as the storage tier in data centers. As SSDs have evolved over the past decade, the commu- nity has continued to debate the interfaces they expose and how operating systems and storage systems should exploit them. The NVMe Flexible Data Placement (FDP) proposal is the latest point in this design space. FDP introduces an interface based on Reclaim Units that enables explicit data placement to reduce device write amplification without the software engineering costs of sequential- write constraints and host garbage collection. FDP-enabled SSDs are emerging in commercial products and early data center de- ployments. Their compatibility with conventional block I/O allows existing applications to run unchanged, allowing a frictionless adop- tion in industry. This paper presents an experimental evaluation of FDP SSDs to characterize their data placement guarantees over the raw device interface. We then revisit two widely deployed and distinct open- source storage systems, MySQL and RocksDB, and examine whether lifetime-based data separation and distinct write patterns built into their architectures can be mapped onto FDP SSDs without inva- sive changes. Our evaluation shows end-to-end WAF reductions at higher device utilization, along with QoS and throughput improve- ments under synthetic and real-world workloads. These results demonstrate that FDP provides a practical and deployable cross- layer mechanism for data placement with open-source ecosystem support on Linux. They also highlight why FDP SSDs are gaining traction in industry.
硬件架构 (cs.AR)
7
cs.AR / 1 / 2610.02401
Design Space Exploration of Backside Clock Meshes for 2 nm GAAFET BSPDN Technology
Wajid Ali, Muhammad Hadir Khan, Dalton Gaddy, Matthew Guthaus
cs.AR
Abstract
Clock meshes are used in high-performance VLSI designs to minimize skew and tolerate on-chip variation, but they spend scarce routing resources on premium metal layers. Backside power delivery creates a new option: it adds thick, low-resistance metal layers on the back of the wafer, and the power grid does not consume all of them. Flip-flops remain on the frontside; a backside mesh therefore cannot drive them directly, and every connection passes through a through-silicon via. We present the first design-space exploration of backside clock meshes, implemented in OpenROAD on GT2N, a 2 nm nanosheet technology. Four benchmarks (1,938 to 15,311 flip-flops) are explored with multi-objective Bayesian optimization, and every design point is verified by transistor-level SPICE simulation, since the cyclic mesh cannot be evaluated by static timing analysis. Across all four designs, the backside mesh consistently outperforms an identical frontside mesh, with on average 45% lower skew, 25% lower sink slew, 4.5% lower power, and 28% less frontside clock wiring, and its skew spread under 10,000-sample Monte Carlo is a third of the frontside mesh's.
cs.AR / 2 / 2610.02502
RAPID: Row-Parallel Arithmetic Processing in DRAM
William C. Tegge, João Paulo Cardoso de Lima, Shouzhi Fang, Jeronimo Castrillon, Alex K. Jones
cs.AR
Abstract
Processing-using-memory (PUM) architectures perform computation directly within DRAM to reduce costly data movement between memory and processors. Because charge-sharing operations are confined to individual bitlines, existing DRAM-PUM architectures reorganize data into column-oriented, bit-serial representations. This organization is fundamentally incompatible with the row-oriented, word-parallel layouts used by conventional processors and accelerators, requiring expensive data-layout transformations whenever computation transitions between PUM and conventional execution. In this paper, we present RAPID, a Row-parallel Arithmetic Processing-In-DRAM architecture. RAPID augments the DRAM subarray with two lightweight extensions: migration cells that enable localized horizontal data movement between neighboring bitlines and inversion cells that provide efficient in-array logical inversion. These primitives enable RAPID to operate directly on row-parallel, bit-parallel data, preserving CPU-compatible layouts while exploiting the massive parallelism of the DRAM subarray. In particular, RAPID demonstrates that localized horizontal communication is sufficient to realize shallow arithmetic networks and efficient parallel reduction for multiplication, reducing arithmetic latency while preserving throughput and eliminating costly data-layout transformations, all while maintaining the conventional DRAM array organization. We demonstrate the feasibility and overhead of augmenting DRAM subarrays with migration and inversion cells through detailed transistor-level layout and SPICE-validated circuit simulations. Using the RAPID compiler it is possible to evaluate the performance and data reorganization tradeoffs to ensure the best execution across combined CPU and PUM. Evaluating RAPID on 19 MLPerf benchmarks, there is a 5.9x higher end-to-end performance compared to SIMDRAM for DDR4 PUM execution.
cs.AR / 3 / 2610.03045
PEEK: Heterogeneous Parallelism for Privileged Error Detection in Safety-Critical Processors
Tinglue Wang, Zhenghui Guo, Bing Guo, Jiapeng Guan, Renshuang Jiang, Jie Zhang, Jing Li, Xin Si, Sam Ainsworth, Zhe Jiang
cs.AR
Abstract
Heterogeneous parallel error detection architecture has been widely studied for safeguarding OoO superscalar processors in safetycritical systems, as it achieves significantly lower hardware overhead compared to traditional LockStep, by exploiting the parallelism that exists in a secondary execution. However, previous works do not cover the protection of privileged-mode execution, impeding their effectiveness in real-world deployment. Moreover, naive extension to privileged-mode can cause a litany of issues, from abysmal performance due to high synchronization costs, to full deadlocks. Here, we present PEEK, the first privileged parallel error detection architecture. Based on a deep analysis of privileged execution, we redesign the verification pipeline, addressing all the bottlenecks and bugs identified in privileged-mode protection. Evaluated using various metrics on an RTL-level full system running Linux, PEEK achieves full-privilege protection on Linux with negligible performance slowdown and affordable hardware overhead. PEEK has been taped out using a 28nm process, and its source is available at https://anonymous.4open.science/r/PEEK-3000.
cs.AR / 4 / 2610.03061
Divide and conquer: Scalable performance and energy in MCM GPUs
Mario Ibáñez Bolado, Borja Pérez Pavón, Jose Luis Bosque Orero, Julio Ramón Beivide
cs.AR
Abstract
Multi-chip-module (MCM) GPUs offer a promising path to scale compute capability beyond monolithic designs by integrating multiple chiplets on a common package. However, the impact of disaggregation on performance scalability and energy consumption remains underexplored. The design space grows rapidly across dimensions such as SMs per chiplet, chiplet count, and interconnection network. The inter-chiplet network is particularly critical, as it determines whether additional compute resources translate into performance gains. This limited understanding leaves industry and research without clear guidance on the performance and energy trade-offs of MCM GPU scaling. In this work, we investigate whether distributing compute and memory capability across multiple chiplets offers a more scalable alternative to concentrating resources. We quantify their effects on performance, energy, and efficiency and examine how inter-chiplet topology influences scalability at different system sizes. Our results demonstrate that a 16 chiplet Torus configuration with 256 SMs delivers a remarkable $2.40\times$ performance improvement over a state-of-the-art MCM architecture with the same compute capability, while simultaneously reducing energy consumption by $4.45\times$. These substantial gains provide evidence that disaggregation is a first-order architectural factor and will be critical to unlocking the performance and energy-efficiency potential of next-generation GPUs.
cs.AR / 5 / 2610.03219
Evidence-Guided Repository-Level RTL Repair
Yuxin Du, Juxin Niu, Zhe Jiang, Nan Guan
cs.AR
Abstract
Repository-level RTL repair must localize a failure that spans files, modules, and clock cycles, then propagate the fix consistently. Existing methods reason over source code, which reveals possible behaviors but not the failed execution, and cannot tell whether a local fix was propagated consistently. We therefore present an evidence-guided framework with three modules. Failure grounding converts a problem statement into a reproduced failing run and a failure anchor. Waveform-guided localization uses a localization toolbox to narrow the observed violation into a candidate mechanism and an evidence trail. Consistency-aware repair and validation then expand that seed into a coordinated patch and replay the same scenario to check that the violation disappears. We conducted experiments on HWE-Bench and achieved better performance than the baseline.
cs.AR / 6 / 2610.02333
Duplication-Aware Retiming and Cell Interface Redesign for Superconductor Circuit Minimization
Panagiotis Papanikolaou, Alex Vanasse, Haoran Jin, Georgios Tzimpragos, Jennifer Volk
cs.ET · cs.AR
Abstract
Superconductor electronics have increasingly shifted away from RSFQ and its variants toward logic families that eliminate explicit gate-level clocking. While this transition enables simpler circuits and more efficient architectures, it also introduces an implicit reliance on dual-rail codes, resulting in inherent gate duplication. This work presents a duplication-aware retiming methodology for Josephson junction (JJ) count minimization, co-optimizing register placement and polarity assignment. The approach applies beyond SFQ to any monotonic circuit. We further identify cell interfaces---specifically, interconnect drivers, receivers, and fanout (FO) elements---as dominant contributors to JJ count in each cell. A new amplifier design is introduced to reduce these costs, integrated within existing SFQ cells, experimentally verified, and characterized to form a new SFQ cell library. Our results demonstrate a 63-71% JJ count reduction in single-cycle implementations and 41-66% reduction in multi-cycle implementations compared to the prior state-of-the-art. The latter establishes a new Pareto frontier, achieving both shorter critical paths and lower JJ counts than the best-to-date single-cycle implementations.
cs.AR / 7 / 2610.03260
Compiling Together: High-Throughput Distributed Quantum Computing via Multi-Compilation
Yipei Liu, Sen Zhang, Zebo Yang, Lei Yang
quant-ph · cs.AR · eess.SY
Abstract
Quantum computing is a promising paradigm for problems that are challenging for classical machines, but realizing that promise requires far more qubits than a single processor can offer. Distributed quantum computing (DQC) scales out by connecting multiple quantum processing units (QPUs), at the cost of making entanglement the scarce resource: every remote gate consumes a Bell pair, and inter-QPU links generate Bell pairs at finite rates orders of magnitude slower than local gates. Since quantum programs are executed repeatedly, this rate bounds how fast shots complete and thus how fast results are obtained. Existing DQC compilers emit a single implementation per circuit, so throughput is capped by its busiest link while other links stay idle. We observe that alternative compilations of the same circuit are logically equivalent yet stress different links; executing them concurrently and pooling their samples converts idle Bell pairs into additional shots. We formulate the joint selection of compilations and allocation of shots under per-link Bell-pair capacities as Candidate-Constrained Max-Shot Allocation (CMA), prove it NP-hard, and solve it with a dynamic program and its approximate variant AppDP, a compact MILP, and a greedy heuristic Effi. In simulation on six-QPU networks, multi-compilation raises throughput by 2-4.5* over the single compilation and correspondingly improves output fidelity under equal Bell-pair budgets. On hardware, it increases measured fidelity by up to 93% and reaches the single-compilation fidelity in up to 10* less time. Scaling to 36 QPUs and 144-qubit circuits, MILP improves throughput over the single compilation by 76.8% on average across 48 configurations while Effi allocates in milliseconds.
密码学与安全 (cs.CR)
25
cs.CR / 1 / 2610.02373
Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM
Jisung Park, John Le, Heath Cooper
cs.CR · cs.AI
Abstract
GraphRAG pipelines construct auxiliary structures during offline indexing--semantic summaries, hierarchical edges, and pre-computed scores--that determine how retrieval is prioritised at query time. Prior attacks target only instance-level components (nodes, edges, triples), overlooking these schema-level structures. We formalise Auxiliary Schema-Level Entity as a novel attack surface and propose the 3S Framework (Semantics, Structure, Scoring) for its systematic exploitation. Our Hop-Decayed Influence (HDI) attack identifies high-impact targets through query-aware influence propagation and corrupts their auxiliary structures post-indexing. Across two benchmarks (HotpotQA, 2WikiMultiHopQA) and two architectures (Microsoft GraphRAG, HippoRAG2), HDI achieves 88-94% attack success rate while modifying as few as 0.016% of auxiliary structures. Each modification affects up to 6.00 queries (Schema Leverage Ratio), demonstrating 1:N amplification unavailable to instance-level attacks. Manipulated structures evade perplexity and paraphrase defenses with over 99% evasion rate, as they remain linguistically coherent system-generated artifacts. These results reveal that auxiliary schema-level entities receive implicit trust without runtime validation, constituting a structural blind spot in current GraphRAG defenses. https://github.com/Jisung-Pacific/HDI-GraphRAG-Attack.
cs.CR / 2 / 2610.02414
Unifying Privacy Accounting: Information Equivalence and Information Loss
Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang
cs.CR · cs.IT · math.ST
Abstract
Differential privacy (DP) admits several notions, but the choice among them may affect both privacy analysis and utility. In this paper, we consider four mainstream curve-based privacy notions within a unified information-theoretic framework. For a fixed ordered pair of output distributions, we establish information equivalence among the two directional privacy profiles of $(\varepsilon,δ)$-DP, the pair of hypothesis-testing trade-off functions, and the extended privacy-loss distribution. The exact Rényi differential privacy (RDP) curve joins this equivalence class whenever it is finite at some order greater than one. Under this mild condition, choosing among these notions changes only their semantic interpretation and computational requirements. In contrast, taking the maximum of the directional privacy profiles or compressing the RDP curve into a single zero-concentrated differential privacy (zCDP) parameter can lose information. We quantify the information loss between the exact RDP curve and its zCDP bound for standard noise mechanisms. This gap is zero for Gaussian noise but generally positive for Gaussian-mixture, Laplace, discrete Gaussian, and Poisson-subsampled Gaussian mechanisms. Moreover, this gap grows linearly with the number of independently composed mechanisms. Our information-theoretic perspective has practical consequences. At the same certified privacy level, retaining the full RDP curve rather than using zCDP reduces the required noise variance by up to $45\%$ for Gaussian-mixture noise in workloads comparable in size to the American Community Survey. For DP-SGD on Fashion-MNIST under Poisson subsampling, an RDP-based privacy accountant improves test accuracy by up to $8.73$ percentage points compared to a zCDP-based accountant when both are calibrated to the same $(\varepsilon,δ)$ guarantee.
cs.CR / 3 / 2610.02435
SoK: Stablecoins in the Quantum Era
Panagiotis Chatzigiannis, Navid Alamati, Suvradip Chakraborty, Duc V. Le
cs.CR
Abstract
Stablecoins support payments, trading, collateral, and cross-chain settlement across the digital-asset ecosystem. They also concentrate value behind issuer, custody, upgrade, oracle, and bridge keys while inheriting the quantum vulnerabilities of host-chain accounts, consensus, rollups, and privacy systems. This paper presents a Systematization of Knowledge (SoK) on post-quantum stablecoins. We develop a stablecoin-specific threat model, map cryptographic dependencies to monetary control surfaces, and classify migration choices along three dimensions: who can authorize a change, where the change must occur, and whether it is hybrid, post-quantum native, or an encapsulation of a classical component. We emphasize an authority-liability gap: the party able to migrate a component is often different from the holders, exchanges, protocols, or issuers that bear the loss if it fails. We review the relevant cryptographic primitives, but distinguish general blockchain failures from their stablecoin-specific effects. We also examine recent blockchain and issuer-controlled interoperability proposals, and relate migration choices to redemption, continuity, and intervention requirements under current stablecoin regulation. Our findings identify open problems in aggregate and threshold authorization, operation-specific security levels and costs, dormant and wrapped supply, private and compliance-enabled transfers, and measurement of quantum-vulnerable stablecoin exposure.
cs.CR / 4 / 2610.02456
SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS
Dimitrios Prasakis
cs.CR · cs.AI
Abstract
AI coding agents are untrusted system components, yet they require autonomy on the developer machines they run on. This contradiction is a security problem. Sandboxes provide an isolated environment, but for local macOS development, the existing local, open-source options for AI coding agents are few in number and cumbersome to use. I conducted a formative online user survey which indicates that fewer than 40% of AI coding agent users run their agents in a sandbox and identifies the top usability barriers hindering AI coding agent sandbox adoption. These findings are used to develop SideKernel: an open-source, local, microVM-based macOS sandbox for AI coding agents designed for usability. To evaluate SideKernel, I compiled a list of sandboxes available on the market and filtered it against five inclusion criteria. Then I performed a comparative analysis between SideKernel and the sandboxes that satisfy these criteria, across 23 capability tests derived from the usability barriers revealed by the user survey. I discovered that only a few sandboxes are similar to SideKernel, and that among those, Docker Sandboxes and SideKernel score highest on capability features related to usability. A secondary contribution of this paper is a survey of the existing solution space for local, open-source, microVM-based macOS sandboxes for AI coding agents.
cs.CR / 5 / 2610.02552
Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection
Sabrine Ennaji, Elhadj Benkhelifa, Nadia Kabachi
cs.CR · cs.AI
Abstract
Machine learning-based intrusion detection systems (IDS) are critical for securing Industrial Internet of Things (IIoT) environments. Most adversarial research against them perturbs the feature vector or the traffic that produces it, and depends on gradient access, repeated model queries, or a learned model of benign traffic. A smaller line of work reshapes packet timing without querying the detector, but makes malicious traffic mimic a learned model of benign timing. Across these approaches, one assumption of industrial monitoring pipelines has received little attention: temporal synchronization. An IDS reconstructs operational state by aggregating telemetry into sliding or tumbling windows, so its view depends not only on what is observed but on when each observation falls relative to a window boundary. We introduce the Phantom State Attack (PSA), which exploits that dependence under a passive, zero-query threat model. Rather than modifying packets, perturbing features, querying the classifier, or fitting any model of benign traffic, PSA injects bounded timing drift calibrated to the attack flow's own inter-arrival variability, moving observations across the nearest window boundary by the minimal shift needed. The IDS then reconstructs a phantom state that diverges from the true process state. We evaluate PSA on ToN-IoT and CIC IIoT 2025 (DataSense), against Random Forest, MLP and XGBoost, measuring detection degradation, synchronization distortion, stealth, and attacker cost. PSA degrades detection on flows carrying enough packets for window-boundary redistribution, and leaves others almost unchanged, so its effect is conditional. A query-based baseline reaches higher raw success but needs many queries per window, while PSA needs none. The results identify temporal aggregation as an attack surface reachable under weaker assumptions than prior evasion techniques.
cs.CR / 6 / 2610.02569
Pincer: Resource Authorization for Agents using a Digital Twin
Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica
cs.CR · cs.AI
Abstract
Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this architecture --- typed tools, information-flow control, or policy prediction engines --- give up too much functionality to be adopted. Agents deployed today (e.g. Claude, Codex) rely on a combination of user-mediated and automode sandboxing as their primary defense. In user-mediated sandboxing, user-maintained policies decay over time and repeated permission requests cause user fatigue, while auto mode's tool-call classifiers learn no user-specific policy and are not meant to defend against adversarial setups. Pincer is a new defense that operates at the resource layer and works alongside existing defenses at the tool-call layer like the auto mode. At the core of Pincer lies a digital twin, an isolated-context model that automatically learns and enforces dynamic user-specific least-privilege policies. The digital twin keeps continually learning the user's preferences allowing it to act as the user's proxy for the agent's permission requests. To emulate the learning phase, we propose a new usercentric dataset with examples following a multi-day transcript of user-agent interaction. Our evaluation shows that Pincer performs strongly on both security and utility in comparison to several baselines which includes variants of LLM judges and adaptations of Conseca (HotOS '25). We highlight attack types where Pincer's design leads to a significant security improvement compared to all other baselines, while outperforming the baselines even for other types of attacks.
cs.CR / 7 / 2610.02805
From TS-SUF-2 to TS-SUF-4: Practical Security Enhancements for FROST2 Threshold Signatures
Will Wang, Syh-Yuan Tan, Ryan Chow, Chanson Chan, Martin Zhao
cs.CR
Abstract
Threshold signature schemes play a vital role in securing digital assets within blockchain and distributed systems. FROST2 stands out as a practical threshold Schnorr signature scheme, noted for its efficiency and compatibility with standard verification processes. However, under the one-more discrete logarithm assumption, with static corruption and centralized key generation settings, FROST2 has been shown by Bellare et al. (in CRYPTO 2022) to achieve only TS-SUF-2 security, which is a consequence of its vulnerability to TS-UF-3 attacks. In this paper, we address this security limitation by presenting an enhanced variant of FROST2, namely, FROST2+ which achieves the TS-SUF-4 security level under the same computational assumptions as the original FROST2. FROST2+ strengthens FROST2 by integrating additional pre-processing token verifications that help mitigate TS-UF-3 and TS-UF-4 vulnerabilities while maintaining practical efficiency. We show that FROST2+ can achieve TS-SUF-4 security not only under the same conditions as the original FROST2 analysis, but also when initialized with a distributed key generation protocol such as PedPoP. Our benchmark using ZCash's FROST library shows that the performance of FROST2+ is comparable to FROST2 and about 64-79% faster than FROST when precomputation is enabled.
cs.CR / 8 / 2610.02869
AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents
Yuelin Wang, Jiongchi Yu, Yanbang Sun
cs.CR · cs.AI
Abstract
Autonomous penetration testing agents conduct multi-step attacks by continuously adapting their plans and actions to target responses. As a common defense, honeypots can be deployed to divert these agents from real assets by presenting decoy services, while also supporting attack tracing and active counterattacks. However, conventional honeypots rely primarily on static artifacts and predefined responses, leaving them unable to adapt to the evolving attack strategies of autonomous penetration testing agents. To this end, we present AgentTrap, the first closed-loop honeypot tailored for autonomous penetration testing agents. AgentTrap uses sentinel endpoints to avoid benign interference, stateful deception grounded in the protected application, and behavior-guided escalation to sustain engagement and collect agent-side behavioral evidence with controlled disclosures. We evaluate AgentTrap against eight autonomous penetration-testing agents in a deployed web application containing a real application endpoint and a separate honeypot endpoint configured under three defense strategies. Compared with no defense, AgentTrap reduces the aggregate real-target attack success rate from 95.8% to 79.2% and successfully elicits attacker API keys in 18.8% of the runs, outperforming static deception and fixed escalation. Furthermore, trace analysis shows that resistance to such counterattacks depends jointly on model-level recognition of deceptive requests and architecture-level isolation of sensitive resources.
cs.CR / 9 / 2610.03073
SecJev: Bringing Security Expertise to System One Decision Models
Zheng Chen, Fei Yu, Haohao Huang, Yang Li, Anlong Chen, Lei Chen
cs.CR · cs.CL
Abstract
Security workflows need models that turn complex observations and explicit policies into decisions. System One models introduced by Jev return typed predictions and probabilities; security specialization supplies the domain expertise behind those predictions. We introduce SecJev, to our knowledge the first family of Jev-like decision models specialized for security, spanning 0.8B to 9B parameters. Built on Kev's single-pass candidate scorer, SecJev learns Boolean, choice, and ordered decisions from text, telemetry, and observation histories. We develop SecJev-Corpus to unify source-label prediction and explicit-policy evaluation across 14 tasks and eight sources. It covers tool outputs, traffic, federated updates, consensus, authentication, and vehicle messages. Scene-weighted training adapts the models across these domains while preserving a shared typed decision interface. Security specialization improves every model in the family; SecJev-0.8B outperforms general Kev-9B by 20.51 percentage points in task-macro accuracy. Comparisons with answer-only generative fine-tuning show close accuracy and latency with lower peak inference memory. Tests on new source groups reproduce gains over Kev in prompt-injection and traffic decisions, with capture-dependent false alarms. We release adapters, decision heads, SecJev-Corpus, and training and inference code.
cs.CR / 10 / 2610.03089
Securing Computer-Use Agents Against Branch Steering Attacks
Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao, Robert Mullins
cs.CR · cs.AI
Abstract
Modern Computer Use Agents (CUAs) directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security guarantees - using an isolated Planner LLM (P-LLM) to fix execution paths before processing untrusted inputs via a Quarantined LLM (Q-LLM) - these guarantees break down in graphical environments. Because CUA interaction is inherently dynamic, plans cannot remain data-independent; they must branch based on anticipated runtime web content - covering all possible cases the agent may encounter. This exposes agents to branch steering attacks, where an adversary crafts untrusted data to coerce a CUA down a hazardous, pre-approved branch without injecting explicit instructions. We systematically study branch steering attacks and introduce STEER-Bench (101 tasks across 9 domains), showing high attack success against both standard (94.4%) and vanilla Dual-LLM (89.5%) CUAs. We then propose COBRA, an architecture that pairs trusted branching plans with ahead-of-time capability constraints, strictly bounding the parameters and destinations each branch may execute. On STEER-Bench, COBRA reduces attack success to 0% while retaining 97% benign utility.
cs.CR / 11 / 2610.03124
The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
Toluwani Aremu, Manit Baser, Mohan Gurusamy, Nils Lukas, Dinil Mon Divakaran
cs.CR · cs.AI · cs.CL · cs.CY · cs.LG
Abstract
Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \emph{token-level trigger-tags}, which introduce watermark-inspired signals during decoding, from \emph{weight-level trigger-tags}, which learn backdoor-inspired associations between target conditions and detectable model behavior. Furthermore, (ii)~we introduce \Untag, a unified attack framework that organizes their mechanism-specific attack surfaces into a common taxonomy. We evaluate representative token-level and weight-level trigger-tags using phishing as a case study. We find that while trigger-tags may provide useful evidence in controlled settings, our attacks render the existing trigger-tag mechanisms to be entirely ineffective. Consequently, we argue that these mechanisms should not be treated as robust misuse detectors when attackers can transform outputs or modify open weights.
cs.CR / 12 / 2610.03166
LiBRA: Detection-Aware Image Watermark Removal via Bidirectional Latent Optimization
Saibo Ye, Huajie Chen, Xin Guo, Le Yang, Chi Liu, Xiangyu Hu, Jingjing Guo, Tianqing Zhu
cs.CR · cs.AI
Abstract
Digital watermarking supports source attribution for AI-generated images, but its reliability depends on resistance to removal attacks. Some attacks attempt to remove watermarks by forcing the decoded watermark to differ from the original. However, this can produce an inverted watermark that remains detectable, causing removal to fail, while further attempts to alter the watermark may unnecessarily degrade image quality. To address these limitations, we present LiBRA (Latent In-band Bidirectional Removal Attack), which aims to make watermarks undetectable while preserving image quality. Instead of continually pushing the watermark toward inversion, LiBRA adjusts the image to conceal the watermark without encouraging further changes that could degrade image quality. Some attacks keep pushing decoded bits away from the original watermark, even when further changes preserve detectability and damage image quality. With access to the watermark key and decoder, LiBRA makes bounded changes in a public autoencoder's latent space. Unlike inversion-driven objectives that cannot correct excessive inversion, LiBRA guides average decoding confidence toward random guessing from either direction. This helps avoid an inverted but detectable watermark. Leaving individual bits flexible allows image-quality constraints to favor less damaging changes, while an optional frequency-guided mask limits their location. We verify removal using an exact two-sided binomial test rather than assuming the confidence target guarantees success.
cs.CR / 13 / 2610.03182
Reversing the Clock: Layout-Aware Recovery of Design Intent from Clock Distribution Networks
Sascha Tommasone, Zehra Karadağ, Christopher Pawlowicz, Michael Green, Bruno Machado Trindade, Eunsung Seo, Christof Paar, René Walendy, Steffen Becker
cs.CR
Abstract
Hardware reverse engineering supports competitive analysis and hardware assurance by recovering information about an integrated circuit (IC) from its physical implementation. While existing techniques primarily recover the gate-level netlist, which represents logical functionality, they often overlook physical design decisions such as placement, routing, delay insertion, and interconnect optimization. The clock distribution network encapsulates many of these decisions; however, no prior published work has recovered this network from a fabricated IC to infer design intent. We present a layout-aware methodology that recovers and analyzes the clock distribution network by integrating a recovered gate-level netlist with layout information extracted from scanning electron microscope imagery. Our four-phase pipeline recovers clock-tree topology, buffering, gating and switching, interconnect delay, and crosstalk-mitigation measures. We demonstrate our methodology on a commercial 450 nm IC, recovering a global H-tree backbone with local X-tree-like branching, identifying an independent clock tree and global clock gating, quantifying latency, skew, and routing lengths, and confirming the absence of dedicated crosstalk mitigation. Together, these findings let us reason about the designer's intent. To encourage further research and support reproducibility, we release our clock tree recovery algorithm as open source.
cs.CR / 14 / 2610.03262
Asymptotic Analysis of Trading Fees in CFMM
Peiyang Jin, Clouds, Jing Qian
cs.CR
Abstract
As the dominant trading mechanism in decentralized finance, Automated Market Maker has been widely studied in research. However, limited research has been done with the trading fees taken into consideration. In this work, we study how much trading fee Liquidity Providers(LPs) can receive from arbitrage trading when the fee rate approaches zero. We give a closed-form formula for it and our result shows that the trading fees generated by arbitrage trading can fully offset the LVR loss when the price process is continuous. We further extend our conclusions to price processes with jumps and the theoretical analysis shows that jumps are the only cause of LP loss apart from market risks. Our results provide practical guidance for AMM designers as well as LPs.
cs.CR / 15 / 2610.03448
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Zhuowen Liu
cs.CR · cs.CL
Abstract
LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent's attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent's own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.
cs.CR / 16 / 2610.03470
CorrectGuard: Eyes-Off Correctness Estimation for Black-Box Security Guardrails
Adam Faulkner, Nil-Jana Akpinar, Matthew Dressman
cs.CR
Abstract
AI services increasingly rely on black-box security guardrails, yet privacy-preserving model auditing regimes often cannot measure how well these systems perform in both a human eyes-off production setting, which disallows human inspection of user input, and a machine eyes-off setting, which disallows model inspection of such input. We introduce CorrectGuard, an eyes-off correctness estimation framework for both settings, which involves an independent model-based evaluator predicting whether guardrail decisions on human- and machine-inaccessible inputs are correct using only labeled eyes-on data and without access to the guardrail's internals. We evaluate in-context learning, embedding, and finetuning-based correctness models under leave-one-dataset-out evaluation across 13 safety and security datasets spanning harmful content, jailbreaks, prompt injection, and extraction, and across open-weight guardrails treated uniformly as black boxes. Across both human and machine eyes-off settings (the latter implemented using privacy-preserving fingerprinting of inputs), in-context-learning-based correctness classifiers substantially improve error identification across guardrails, achieving up to a 25 percentage-point increase in macro accuracy, as do finetuning-based approaches which provide a nearly 15-point boost, although performance varies sharply across guardrails and held-out datasets. Correctness scores also support guardrail decision ranking and abstention: across 3 guardrails, the best correctness rankings reduce AURC from unranked baselines of 0.33-0.44 to 0.17-0.22, while the best operating points retain 37.5-52.0% of guardrail decisions at 15% observed risk. These results show that external correctness models can expose systematic failures and support guardrail decision abstention without privileged access to the guardrail.
cs.CR / 17 / 2610.03562
A Secure dToF LiDAR SoC with Dual-Domain Fingerprinting and Event-Driven AFE Circuit Achieving Sensor-Level Attack Resilience
Risa Nonaka, Ryoya Matsuno, Shota Nagai, Satomi Miyagi, Yuki Hayakawa, Ryo Suzuki, Kazuma Ikeda, Ozora Sako, Rokuto Nagata, Ryo Yoshida, Shimpei Ando, Wenlun Zhang, Kentaro Yoshioka
cs.CR · cs.AR
Abstract
Recent studies have shown that most commercial direct time-of-flight (dToF) LiDARs can be spoofed by injecting high-frequency laser pulses into the receiver, which can erase pedestrians from the point cloud. This paper presents the first dToF LiDAR system-on-chip (SoC) with integrated sensor-level hardware security against spoofing attacks. We propose Dual-Domain Fingerprinting (DDF), which emits laser pulse pairs whose time interval and amplitude ratio are both randomized and authenticates received echoes in this two-dimensional space, so that spoofed signals are rejected before they corrupt the ranging result. An Event-Driven AFE (ED-AFE) activates the ADC only around pulse peaks: it digitizes three samples around each peak with a triggered ADC at 1-GHz sampling and applies parabolic interpolation, achieving 1-cm distance resolution with a 99% reduction in ADC power. A time-modulated laser driver controls the laser amplitude from 10% to 100% by modulating the charge time of a laser capacitor, providing the microsecond-order amplitude modulation required by DDF. A LiDAR system with a 16-channel 65-nm CMOS prototype SoC demonstrates up to 120-m ranging and an AFE power of 3.1 mW per channel, 60% lower than prior art. In a proof-of-concept experiment in which the dual-domain authentication is applied to measured sensor data, 73% of the point cloud is protected under spoofing attack, compared with 0% without DDF.
cs.CR / 18 / 2610.03585
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu
cs.CR · cs.AI · cs.LG
Abstract
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool. We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini). The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score.
cs.CR / 19 / 2610.03590
Constant-Rate Certified Deletion
Kai-Min Chung, Tzu-Hsiang Huang, Wei-Hsiang Hung, Shota Yamada
cs.CR · quant-ph
Abstract
We present a unified framework for upgrading a broad class of cryptographic primitives to support constant-rate certified deletion. Previous constructions require a linear number of qubits per encrypted bit of certified-deletable plaintext. In contrast, we obtain the first constant-rate constructions in the plain model that achieve certified deletion while preserving everlasting security. Our approach applies to a wide range of "all-or-nothing"-type primitives based on BB84-style encodings, including commitment schemes, public-key encryption, attribute-based encryption, and fully homomorphic encryption. Beyond this class, we also obtain constant-rate certified deletion for primitives built from subspace coset states, such as blind delegation, secure software leasing, functional encryption, and differing-inputs iO, and CCA-PKE. Importantly, our framework does not introduce any additional assumptions beyond those required by the underlying certified deletion primitives. Finally, under the hardness of SIS, we show that public verifiability can be incorporated into BB84- and coset-based certified deletion. In combination with our constant-rate constructions, this yields publicly verifiable certified deletion schemes with constant rate.
cs.CR / 20 / 2610.03650
PoCoFL: POlicy-COmpliant Federated Learning
Dominik Roy George, Varesh Mishra, Aysajan Abidin
cs.CR · cs.LG
Abstract
Federated Learning (FL) is a privacy-oriented learning paradigm that enables collaborative model training while keeping training data local to participating clients. However, it does not guarantee that clients submit policy-compliant contributions or that aggregators process admitted contributions correctly. Existing verifiable FL systems tailor validation rules to specific FL settings, learning workflows, and cryptographic constructions, limiting their applicability across network topologies, participant roles, and aggregation semantics. In this paper, we present PoCoFL, a policy-compliant federated learning framework that separates three aspects: (i) FL type, (ii) policy semantics, and (iii) cryptographic realisation. We provide a formalisation that captures client and aggregation requirements as policy-dependent relations. Clients prove compliance of their contributions using commitments and non-interactive zero-knowledge proofs, while aggregators prove that the recorded set of admitted contributions was processed according to the selected aggregation policy. We demonstrate PoCoFL through four formal instantiations: (i) vanilla, (ii) continual, (iii) personalised, and (iv) threshold-encrypted federated learning. We evaluate the effects of policy enforcement on the learning objectives of vanilla, personalised, and continual FL. We further implement proof-of-concept realisations of all four instantiations, demonstrating the versatility and practical feasibility of PoCoFL. Overall, these results show that PoCoFL can capture complex policy representations while remaining network-topology agnostic.
cs.CR / 21 / 2610.03440
Quantifying Ethereum Energy Consumption via Network Mapping
Yahn Costa Hackspacher, Cornelius Ihle, Vasundhara Shaw, Dennis Trautwein, Geerd-Dietger Hoffmann, Bela Gipp, Moritz Schubotz
cs.CY · cs.CR
Abstract
Ethereum's electricity use fell by about 99.95% after the move from proof of work to proof of stake. Service providers still need to report operational energy use, e.g. under the EU Markets in Crypto-Assets Regulation (MiCAR). Existing estimates either apply one typical wattage to every node or start from aggregated monitoring counts. Both ignore attributes that nodes already advertise on the peer-to-peer network: client software, ARM or x86 hardware, hosting location, and validator role. We crawl the consensus and execution layers, assign each peer a wattage from those attributes using published measurements, and estimate the remaining incomplete peers with a Random Forest. On 6,934 peers from two Nebula crawls (19 and 22 June 2026), reachable nodes sum to 415 kW, or 3.63 GWh if that draw were held for a year. The same Lighthouse+Nethermind x86 wattage on every peer yields 431 kW. Observed attributes lower the total by 3.9%, mainly because nodes at Hetzner and other non-AWS clouds draw less than that home-desktop figure. AWS accounts for 15.6% of watts from 12.2% of peers, and validator-flagged nodes for 31.4% of watts from 25.7% of peers. The 415 kW snapshot is about 46% of the Cambridge Centre for Alternative Finance (CCAF) estimate of about 0.90 MW. Both use about 60 W per node, so the gap is mostly how many nodes each estimate includes. Rules cover 3,110 peers and the forest the other 3,824. On held-out labeled peers with client, architecture, and OS hidden, the forest's mean absolute error against the rule wattage is 4.3 W. Twenty-four-hour measurements on a gaming desktop differ from the predictions. After subtracting a 33 W idle graphics card that Ethereum clients do not need, both differences fall to about 19%.
cs.CR / 22 / 2610.02636
Where Quantum Fourier Sampling Stops Short: A Three-Gate Audit Protocol for Delay-PUF Security Models
Owen Friedewald, Ali Shiri Sichani, Chi-Ren Shyu
quant-ph · cs.CR · cs.LG
Abstract
Quantum Fourier sampling may help audit the spectral learnability of delay-based physical unclonable functions (PUFs). We ask whether that promise survives access matching, a strong classical comparator, and oracle synthesis. Three gates structure the evaluation. Structure: low degree is not small support at reachable challenge lengths; for 4-XOR at $n=14$, degree $\le d_f(0.1)$ admits $91\%$ of all $2^n$ characters and the median $90\%$-mass set spans a third of the spectrum. Algorithmics: constructing the phase oracle logically implies classical membership access, making Kushilevitz--Mansour the correct baseline; across 45 tasks it exhausts each finite domain, and no 4-XOR ideal-sampling case reaches $90\%$ mass within $2^n$ calls. A quantum-kernel diagnostic appears more favorable, with geometric difference rising to $2.151$ at $N=512$ challenges, but it correlates $0.991$ with $1/\sqrt{λ_{\min}(K_C)}$ for the classical Gram matrix $K_C$, and the 4-XOR label-complexity ratio does not exceed a balance-preserving permutation null ($p=0.930$). Trace-normalized geometric difference can therefore grow through classical ill-conditioning alone, without task-label alignment. Implementation: a simulator-validated fixed-point phase oracle based on the quantum Fourier transform admits an $18.9\%$ routed-depth reduction, yet the least certified precisions have estimated durations of $1.18$--$1.55\times$ the median dephasing time $T_2$ of the mapped qubits on a static backend snapshot, without hardware execution. We find no end-to-end advantage in the evaluated regime, although ideal sampling does use fewer coherent calls on the thresholded task. The contribution is the Three-Gate Quantum Audit Protocol: a reproducible procedure separating an ideal query advantage from a realizable security benefit. This is not a claim about deployed silicon and not an impossibility result.
cs.CR / 23 / 2610.02641
When Normalization Selects the Sign: Auditing Robustness Ablations in Quantum Attention
Owen Friedewald, Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu
quant-ph · cs.CR · cs.LG
Abstract
Removing an input-scaling module changes both a classifier and the perturbations reaching its encoder. A robustness difference can therefore reflect the comparison rule as well as the module. We demonstrate this problem in a four-qubit quantum-attention detector on generated power-grid trajectories. A learned scaling module appears beneficial at a fixed physical attack budget, but matching an upper bound on perturbations at the encoder reverses the ordering. Neither comparison alone establishes a robustness benefit caused by the module. The initial test also perturbs clean examples into attacked examples while retaining their original labels; tests restricted to already attacked examples do not establish a benefit. Replacing a trained model's input scales disrupts detection. Retraining its linear classification layer restores the detection rate, but changes individual predictions, leaving the comparison descriptive rather than causal. Two further design checks explain why the input quantum Fisher information regularizer cannot train this model's query parameters, and why removing confidence bounds does not establish a larger certified radius. The evidence is limited to ten seeds, exact simulation, synthetic data, and a restricted set of attacks; classical baselines achieve better clean prediction. The practical lesson is to specify which perturbation budget is fixed, check that attacks preserve labels and interventions preserve predictions, and distinguish exploratory controls from confirmatory evidence.
cs.CR / 24 / 2610.03593
Quantum Fire with Delegated Cloning
Rohit Chatterjee, Ananta Mukherjee, Vir Pathak, Supartha Podder
quant-ph · cs.CR
Abstract
Quantum fire is a recently introduced cryptographic primitive consisting of efficiently preparable quantum states, called \emph{flames}, that admit efficient cloning but resist efficient telegraphing, namely reconstruction via classical communication without preshared entanglement. In all prior constructions of quantum fire, cloning is a public operation that requires no separate key and every holder of a flame state can clone it. For applications to access control, however, an issuer may wish to delegate cloning to designated quantum servers while withholding this capability from other flame holders. To address this, we introduce \emph{delegatable quantum fire}, in which cloning requires a separate key. We give two constructions in the classical-oracle model. Our first construction uses a classical secret key which enables cloning, and any user with the entire key may clone successfully. Our second construction, which we call \emph{torch-fire}, uses quantum cloning keys, called \emph{torches}, which enable cloning while remaining unaffected in the process but cannot otherwise be split or delegated to enable additional cloning. An efficient adversary given $m$ torches cannot, except with negligible probability, enable more than $m$ noncommunicating parties to each clone a fresh, independently issued challenge flame. The adversary may jointly process its resources before separating the parties and distribute arbitrarily entangled registers among them. Both constructions are in the oracle model, relying on public classical oracles that allow queries in quantum superposition.
cs.CR / 25 / 2610.03705
Unitary complexity in polynomial space
William Kretschmer, Ewin Tang
quant-ph · cs.CC · cs.CR
Abstract
We show that if quantum commitments exist, then either there is no polynomial-time solution to the unitary synthesis problem, or $\mathsf{BPP} \neq \mathsf{NEXP}$. Thus, showing unconditionally that quantum commitments exist would require answering at least one of two longstanding open questions in complexity theory. We prove our main result as a consequence of a more general lemma, which shows that every unitary in $\mathsf{unitaryPSPACE}$ either cannot be synthesized efficiently relative to any classical oracle, or can be synthesized efficiently with an oracle for $\mathsf{NEXP}$ search problems. Our lemma has other noteworthy consequences, including that certain oracle separations involving $\mathsf{unitaryPSPACE}$ would imply breakthrough classical lower bounds such as $\mathsf{NC} \neq \mathsf{NP}$. Along the way, we propose new definitions for the unitary complexity classes $\mathsf{unitaryP}$ and $\mathsf{unitaryPSPACE}$. Our changes address the biggest conceptual issues with definitions suggested in prior work, and lead to elegant proofs. We study both implementations that erase garbage and implementations that allow it, because we cannot rule out the possibility that the two definitions differ. Nevertheless, we show that both definitions can be viewed as special cases of each other. We also showcase many other ways in which our definitions are robust. For example, we show that $\mathsf{unitaryPSPACE}$ has an equivalent characterization as the set of unitary transformations whose entries can be computed to arbitrary precision in polynomial space. Consequently, we deduce that $\mathsf{unitaryPSPACE}$ can generically erase garbage, a result that provably fails relative to unitary oracles.