Daily Research Digest
arXiv Papers
2026-10-02
640
Papers
9
Categories
162
Translated
收藏清单 0
精选 · Favorites
162
cs.LG / 1 / 2610.00546
Generative Modeling of Stochastic Dynamics for Long-Time Evolution
面向长时间演化的随机动力学生成建模
diffusion
扩散模型相关
Abstract
Exact stochastic equations for non-equilibrium dynamics are rarely accessible. We show that the long-time evolution of stochastic dynamics can be predicted from configuration pairs at a fixed short time lag, without knowledge of the equation of motion. Generative diffusion models learn the finite-time transition kernel from these pairs, and iterating it propagates the dynamics far beyond the training lag. For two-dimensional Model B, the diffusive dynamics of a conserved order parameter, the learned kernels reproduce dynamic critical scaling and self-similar $t^{1/3}$ coarsening. Agreement with direct simulations persists on lattices twice the largest training size and for initial ensembles absent from training. For driven colloids in a periodic optical potential, ten minutes of measured trajectories suffice to predict the particle current and mean passage time over the next twenty minutes within experimental uncertainty. Short-time observations thus contain the information needed to predict emergent non-equilibrium dynamics at much longer times.
Chinese Translation
非平衡动力学的精确随机方程很少能够获得。我们表明,无需知道运动方程,随机动力学的长时间演化就可以从固定短时间滞后下的构型对进行预测。生成扩散模型从这些构型对中学习有限时间转移核,迭代该转移核可将动力学传播到远超训练滞后的时间。对于二维 Model B,即守恒序参量的扩散动力学,学习到的核重现了动态临界标度和自相似 $t^{1/3}$ 粗化。与直接模拟的一致性在最大训练尺寸两倍的格点上以及训练中未出现的初始系综上依然保持。对于周期性光势中的驱动胶体,十分钟的测量轨迹足以在实验不确定度内预测接下来二十分钟内的粒子流和平均通过时间。因此,短时观测包含了预测更长时间尺度上涌现的非平衡动力学所需的信息。
cs.AI / 2 / 2610.00613
Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents
空间策略,而非动作:向量量化测地线作为LLM驱动智能体的工具
large language model
大语言模型相关
Abstract
Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their spatial comprehension through an architecture combining geometrical tools with a LLM serving as a high-level orchestrator in grid-world environments. The agent first collects geodesic trajectories, which are then vector-quantized to extract a representative subset. Offline, the LLM associates a natural language description of the underlying behavioral patterns to each selected trajectory, making it a tool. Online, the LLM chooses the appropriate tool conditioned on the current state and goal. Low-level control is handled by primitive actions that execute the trajectory associated with the tool. From an agentic AI perspective, this approach separates learning into two levels: tool discovery is handled through unsupervised quantization of trajectories, while reasoning and decision-making are handled by the LLM. We test the approach in a partially observable dynamic 2D grid environment with an open vision-language model (Qwen3.6-35B-A3B). Pairing the geometry-derived tool library with an agent-centered zoom tool and a collision detection tool lets a fast, non-reasoning configuration match the goal-reaching rate of a much more costly chain-of-thought version, while cutting the cost of a decision from minutes to seconds.
Chinese Translation
基于大语言模型(LLM)的智能体常被批评缺乏空间理解能力,并且主要利用统计性的文本模式。我们通过一种架构来研究它们的空间理解能力,该架构将几何工具与一个在网格世界环境中充当高层编排器的LLM相结合。智能体首先收集测地线轨迹,随后对这些轨迹进行向量量化,以提取出一个具有代表性的子集。在离线阶段,LLM为每条选中的轨迹关联一段对潜在行为模式的自然语言描述,从而使其成为一个工具。在线阶段,LLM以当前状态和目标为条件选择合适的工具。底层控制由执行与该工具相关联轨迹的原子动作来处理。从智能体式AI的视角来看,该方法将学习分为两个层次:工具发现通过轨迹的无监督量化来处理,而推理与决策则由LLM来处理。我们在一个部分可观测的动态二维网格环境中,使用一个开放视觉语言模型(Qwen3.6-35B-A3B)对该方法进行测试。将几何推导出的工具库与一个以智能体为中心的缩放工具和一个碰撞检测工具相配对,可使一种快速的、非推理配置达到与成本高得多的思维链版本相当的到达目标率,同时将一次决策的成本从数分钟缩短至数秒。
cs.AI / 3 / 2610.00663
Backdoor Containment via Expert Quarantine and Shutdown in LLMs
通过专家隔离与关闭实现大语言模型中的后门遏制
large language model
大语言模型相关
Abstract
Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.
Chinese Translation
被植入后门的大语言模型(LLMs)在良性输入上可以表现正常,而在隐藏触发器下会产生攻击者指定的输出。现有防御横跨四个阶段——训练前、训练中、训练后和推理时——并且共享两种底层策略之一:要么抑制后门学习(通过过滤被投毒的数据或在优化过程中中断其习得),要么先学习、再净化(在完全被植入后门的模型形成之后,通过修复模型权重或对输入进行门控)。我们提出第三种策略:学习,但疏导:允许后门在训练期间形成,但将其引导到一个指定的、被隔离的组件中,该组件可在部署时被禁用。为此,我们提出隔离专家关闭 Quarantined Expert Shutdown QES,这是一种在正则化引导的类 MoE 设置中构建的计算高效的遏制策略。具体而言,给定一个被投毒的数据集,QES 通过路由的专家专属 LoRA 分支和轻量级路由器来增强基于 Transformer 的语言模型,并使用辅助路由目标将触发器条件下的行为吸引到指定专家中,同时在其他地方保留良性能力。在部署时,缓解被简化为单个常数时间操作:将隔离专家的路由权重置零,无需触发器筛查,也无需进一步更新模型权重。在实证上,我们的方法在两个任务、三种攻击和四个模型家族的多数设置中,将攻击成功率 ASR 从 100% 降低到 0-10%,而下游效用往往得以保持或仅受到适度影响。这些结果确立了「学习但疏导」作为生成式 LLMs 中后门遏制的一个此前未被探索的范式。
cs.AI / 4 / 2610.00682
Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
基于本体、经推理机验证的基准,用于评估科学 AI 中的 LLM 推理
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.
Chinese Translation
大语言模型(LLM)日益支撑着对结构化知识进行推理的科学 AI 应用,从生物医学问答到材料信息学。然而,它们的逻辑推理往往有所欠缺,产生在这些场景中不可接受的事实不准确。可靠的评估仍然具有挑战性:手动构建数据集难以扩展,而基于 LLM 的生成则有嵌入其旨在衡量的那些缺陷的风险。高质量基准必须将正确和错误标注的示例都建立在显式背景知识之上,并能由标准推理机进行形式化验证。我们提出了一个流水线,可从任何具有充分公理的 OWL 2 本体自动生成基于本体的多项选择题(MCQ)基准,其正确答案在设计上就根植于该本体。干扰项通过扰动类定义公理的右侧类表达式生成,其错误性由 OWL 推理机通过蕴涵检查进行形式化验证。我们在三个本体上评估该流水线:Pizza(小型、学术)、PMDco(复杂、材料科学)和 DOID(大型、生物医学),分别生成 112、2,491 和 15,216 道 MCQ。干扰项涵盖四个语义类别,从类不可满足性到弱化的包含关系,从而能够对特定推理失败进行诊断性评估。题目符合自然语言质量标准:平均 LLM 评判分数分别为 5 分中的 4.02、4.36 和 3.36,证实了流畅性;正确答案与干扰项的相似度高于 0.8,表明错误选项不能仅凭表面形式被排除。六个 LLM 在零样本评估下达到 41.1–76.8% 的准确率,远高于 25% 的随机猜测基线,表明这些基准具有挑战性且具有区分度。这项工作朝着更可靠的、用于评估科学 AI 中逻辑推理的基准迈出了一步。
cs.AI / 5 / 2610.00685
Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
基于零空间投影的 LoRA 微调大型语言模型后门净化
large language model
大语言模型相关
Abstract
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.
Chinese Translation
随着大型语言模型(LLMs)和参数高效微调(PEFT)方法的快速采用,后门攻击的风险已变得更加严峻。现有的后门净化方法通常依赖于至少一个强假设,例如对触发器的先验知识、对干净参考的访问,或激进的重训练,而且它们往往缺乏全面的评估。这些约束极大地限制了它们的实际适用性。为了克服这些挑战,我们的工作提出在没有这些假设、甚至不对可疑参数进行事后重训练的情况下净化 LoRA 微调的 LLM。我们的目标是在显著降低攻击成功率(ASR)的同时,既保留 (i) 基础模型的一般能力,也保留 (ii) 通过适配器学到的新的下游技能。通过一系列消融研究,我们将我们的方法从文本分类设置中的单层逐步扩展到生成任务中的全参数 LLM。通过精心的数据整理和特征近似,我们提取高保真后门方向,并针对每一层或每个头,在输入和输出通道中构建正交零空间,将 LoRA 更新投影到这些零空间上。经验上,我们的零空间投影方法将 ASR 从接近 100% 降低到不到 10%,同时在下游任务适应期间保留基础模型的良性性能和适配器学到的能力。
cs.AI / 6 / 2610.00710
ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality
ReLiveGym:在数周重放现实中评估长期存活的智能体
large language model
大语言模型相关
Abstract
As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic environment over time. These challenges are not fully captured in the existing long-horizon agent work, as they often consider a static environment that is not temporally changing. We introduce ReLiveGym, a diagnostic evaluation environment of long-lived tasks in which agents act sparsely over simulated weeks of chronologically replayed real-world news, market, and social-media streams. The tasks span diverse levels of time sensitivity, reasoning intensity, and recurrence. Across eight base language models, we investigate how model choice and harness design affect agent performance on such long-lived tasks. Our results show that how agents determine when to act arises as an important harness-design axis for long-lived tasks; and that the optimal design varies across tasks and sometimes model choices as well. We also evaluate how continuous learning from hindsight feedback affects performance and addresses failure modes observed in these long-lived tasks. These findings indicate model choice, action timing mechanism, and use of feedback as important considerations in the design of long-lived agents. Code: https://github.com/SaharaLabsAI/ReLiveGym
Chinese Translation
随着大语言模型(LLM)智能体被广泛采用,它们越来越多地被部署到需要持续监控或周期性动作的任务中(例如市场分析)。这些智能体被期望能够在无人值守的情况下运行数天或数周,在恰当的时机采取行动,并随时间推移适应动态环境。这些挑战在现有的长时程智能体研究中并未得到充分刻画,因为它们通常考虑的是不随时间变化的静态环境。我们提出 ReLiveGym,一个面向长期存活任务的诊断性评估环境,在其中智能体于按时间顺序重放的真实世界新闻、市场和社交媒体数据流所模拟出的数周时间内稀疏地采取行动。这些任务涵盖不同的时间敏感性、推理强度和重复发生频次水平。在八个基础语言模型上,我们研究了模型选择与运行框架(harness)设计如何影响智能体在此类长期存活任务上的表现。我们的结果表明,智能体如何决定何时行动,成为长期存活任务中一个重要的运行框架设计维度;并且最优设计因任务而异,有时也因模型选择而异。我们还评估了基于事后反馈的持续学习如何影响性能,以及如何应对在这些长期存活任务中观察到的失败模式。这些发现表明,模型选择、行动时机机制以及反馈的使用,是长期存活智能体设计中重要的考量因素。代码:https://github.com/SaharaLabsAI/ReLiveGym
cs.AI / 7 / 2610.00872
MemFit: Efficient Long-Term Agentic Memory
MemFit:高效的长期智能体记忆
large language model
大语言模型相关
Abstract
Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations. To address this limitation, we propose MemFit, a long-term memory system for conversational agents that reduces the cost and latency of memory operations. Unlike existing systems that rely on expensive LLM calls for memory construction or discard surface-level details through compression, MemFit stores each turn verbatim in an append-only store with near-instantaneous, LLM-free insertion, indexing turns with segment summaries rather than replacing them. Additionally, MemFit uses an LLM-free, multi-path retrieval strategy that combines lexical and semantic signals with cross-encoder reranking over caption- augmented episodes in both textual and multimodal settings. Empirical results on three widely used benchmarks, LoCoMo, MemGallery, and LongMemEval-S, show that MemFit achieves state-of-the-art performance while reducing memory construction time and cost several-fold, providing a scalable and efficient solution for persistent agentic memory.
Chinese Translation
面向大语言模型(LLM)的长期记忆系统因能够在各类应用中扩展推理能力而日益流行。当前的记忆系统依赖 LLM 智能体来组织和巩固记忆,导致写入操作成本高昂且效率低下。为了解决这一局限,我们提出了 MemFit,一种面向对话智能体的长期记忆系统,可降低记忆操作的成本和延迟。与依赖昂贵 LLM 调用进行记忆构建或通过压缩丢弃表层细节的现有系统不同,MemFit 将每一轮对话逐字存储在仅追加存储中,实现近乎瞬时的、无需 LLM 的插入,并用片段摘要对轮次建立索引,而不是替换它们。此外,MemFit 使用一种无需 LLM 的多路径检索策略,该策略将词汇和语义信号与交叉编码器重排序相结合,并在文本和多模态设置下对说明文字增强的片段进行重排序。在三个广泛使用的基准 LoCoMo、MemGallery 和 LongMemEval-S 上的实证结果表明,MemFit 在实现最先进性能的同时,将记忆构建时间和成本降低数倍,为持久化智能体记忆提供了一种可扩展且高效的解决方案。
cs.AI / 8 / 2610.00912
OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework
为从事运筹的 AI 进行运筹:在 OSCAR 框架内将 LLM 沿自动扶梯向上路由
large language model
大语言模型相关
Abstract
Large language models can translate business descriptions into optimization models, but executable code may misrepresent constraints or objectives. A solver can then return an optimal solution to the wrong problem. Even when the solution satisfies the intended operating rules, a better plan may exist. For organizations that repeatedly use optimization modeling, an LLM-based framework should produce accurate formulations at low cost and, ideally, run locally. We study how to verify improvements and allocate attempts across LLMs that differ in price and capability. We develop OSCAR (Optimization modeling by Simulator, Coder, And Reviewer), which uses an offline Simulator certified against labeled decision examples to compare candidates and continues searching beyond feasibility. We model the search for the next certified improvement as sequential decisions under unobserved difficulty: which LLMs to call and when to stop. In a simplified known-prior setting, we give conditions under which cost-ordered escalation is optimal. For general menus, we derive a prior-free competitive guarantee. On five benchmark problems, OSCAR achieves 95% to 100% accuracy at the reported settings using two small open-weight LLMs, each deployable locally on a single GPU. Their single-attempt accuracies average 29% and 48%. In five runs per problem, Codex and Claude Code incur average token costs 3.1 and 5.8 times OSCAR's, respectively. OSCAR supports open-weight models locally or in the cloud, depending on budget and confidentiality requirements. Firms should maintain labeled decision examples of feasible and infeasible decisions to clarify plain-language operating rules. OSCAR follows these labels when an LLM's interpretation conflicts with them. As LLM capabilities and prices change, OSCAR's simple operating rules and adjustable settings help firms adapt their model choices and benefit from these advances.
Chinese Translation
大型语言模型可以将业务描述转化为优化模型,但可执行代码可能会错误地表示约束或目标。求解器随后可能返回错误问题的最优解。即使该解满足预期的运营规则,也可能存在更好的方案。对于反复使用优化建模的组织,基于 LLM 的框架应以低成本生成准确的建模表述,并且理想情况下可在本地运行。我们研究如何验证改进,并在价格和能力不同的 LLM 之间分配尝试次数。我们开发了 OSCAR(Optimization modeling by Simulator, Coder, And Reviewer,即由模拟器、编码器和评审器进行优化建模),它使用一个根据带标签决策示例认证的离线模拟器来比较候选方案,并在可行性之外继续搜索。我们将对下一次经认证改进的搜索建模为在未观测难度下的序贯决策:调用哪些 LLM 以及何时停止。在一个简化的已知先验设置中,我们给出按成本排序的逐级升级为最优的条件。对于一般菜单,我们推导出一个无需先验的竞争保证。在五个基准问题上,OSCAR 在报告的设置下使用两个小型开放权重 LLM 实现了 95% 至 100% 的准确率,每个 LLM 都可在单块 GPU 上本地部署。它们的单次尝试准确率平均分别为 29% 和 48%。在每个问题五次运行中,Codex 和 Claude Code 的平均 token 成本分别是 OSCAR 的 3.1 倍和 5.8 倍。OSCAR 支持在本地或云端使用开放权重模型,具体取决于预算和保密要求。企业应维护可行和不可行决策的带标签决策示例,以澄清用平实语言表述的运营规则。当 LLM 的解释与这些标签冲突时,OSCAR 遵循这些标签。随着 LLM 能力和价格变化,OSCAR 简单的运营规则和可调设置帮助企业调整其模型选择,并从这些进步中受益。
cs.AI / 9 / 2610.00947
ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning
ABDA-NL:面向基于论证的推理的自然语言场景浏览器
large language model
大语言模型相关
Abstract
ABDA-NL adds a natural-language interface to ABDA, a system for argument-based discussion using ASPIC- knowledge bases under grounded semantics. Users see which conclusions are accepted, rejected, or undecided, open an interactive rendering of the grounded discussion game to learn why, explore what-if alternatives by suspending assumptions and rules or changing preferences, ask questions that are answered from a scenario's reference documents, and author new facts, assumptions, and rules in plain English. A large language model provides the bridge between language and formalism: it answers questions from the documents and the current state of the scenario, and it translates plain-English edits into candidate formal statements. The deterministic ABDA engine remains the sole source of arguments, attacks, and acceptance labels, and every proposal of the model is validated and confirmed by the user before it takes effect.
Chinese Translation
ABDA-NL 为 ABDA 增加了一个自然语言接口,ABDA 是一个在 grounded 语义下使用 ASPIC- 知识库进行基于论证的讨论的系统。用户可以查看哪些结论被接受、被拒绝或未决,打开 grounded 讨论博弈的交互式呈现以了解其原因,通过暂停假设和规则或更改偏好来探索 what-if 备选方案,提出由场景的参考文档来回答的问题,并用平实的英语撰写新的事实、假设和规则。大语言模型在语言与形式化之间提供桥梁:它根据文档和场景的当前状态回答问题,并将平实英语的编辑转换为候选的形式化陈述。确定性的 ABDA 引擎仍然是论证、攻击和接受标签的唯一来源,并且模型的每一项提议在生效之前都由用户验证和确认。
cs.AI / 10 / 2610.01000
Evaluating LLM-Generated Preference Distributions
评估LLM生成的偏好分布
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, all models considered in our analysis exhibit self-coherence, with the most probable outcomes stabilizing rapidly under repeated sampling. At the same time, we observe substantial discordance across both model families and scales, with little consensus even among their most probable outcomes. These patterns hold across nine open-weight models, three choice domains, and show robustness under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that outcomes are influenced more by the choice of model than by the wording of the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.
Chinese Translation
大语言模型(LLMs)正越来越多地被用作概率生成器,用于在无法获得真实世界数据的场景中进行仿真、合成数据生成和决策支持。然而,其所生成分布的结构与可靠性仍研究不足。在此,我们系统性地分析了LLM生成的关于航空旅行、餐厅和消费品的偏好分布。令人鼓舞的是,我们分析中考虑的所有模型都表现出自洽性,随着重复采样,最可能的结果迅速趋于稳定。与此同时,我们在不同模型家族和规模之间都观察到显著的不一致性,即使在它们最可能的结果之间也几乎没有共识。这些模式在九个开放权重模型、三个选择领域中均成立,并且在温度变化、贪婪解码以及提示和排序扰动下表现出稳健性。我们的发现表明,结果受模型选择的影响大于受提示措辞的影响,这挑战了一个常见假设:当足够有能力的大语言模型被用作调查受访者的替代者时,它们会产生相似的偏好分布。
cs.AI / 11 / 2610.01017
Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization
为故障付费,而非为流程付费:无标签的流内多智能体工作流优化
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to $+11.97\%$, achieving $+9.64\%$ with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.
Chinese Translation
大型语言模型(LLMs)越来越多地构建多智能体工作流,这些工作流将复杂任务分解,并从智能体池中分配专家智能体。然而,构建好这样的工作流仍然具有挑战性:任务应该划分得多细、每个子任务应信任哪个智能体,以及何时创建新的专家,都是工作流构建器需要事先解决的关键决策。因此,每个子任务是否成功,直到工作流运行之前仍然未知。然而,改进工作流的成本很高。定位故障通常需要参考答案、分级结果或训练有素的评估器,并且修复会通过重新执行、重新搜索或重新训练应用于整个工作流。我们提出 InFlowOp,它用一个无标签成本为每个决策定价,该成本权衡智能体的能力在多大程度上满足子任务的需求,与运行该智能体需要多少开销。在执行之前,InFlowOp 双向确定任务分解的粒度和智能体分配,这些由成本决定,而不是由固定模板决定。在执行期间,InFlowOp 通过同一成本以最便宜的动作纠正故障,该成本在工作流构建和运行时都为其服务。面对工作流级评估挑战,我们引入了 Braid,这是一个基准,其任务需要超出单智能体能力的多智能体协调。在各个领域和骨干模型上,InFlowOp 相比单智能体基线提升高达 $+11.97\%$,通过流内优化达到 $+9.64\%$。我们的项目页面:https://xhguo7.github.io/InFlowOp/。
cs.AI / 12 / 2610.01048
Network World Models as Environments for Algorithm Design on Complex Systems
网络世界模型作为复杂系统上算法设计的环境
diffusion
扩散模型相关
Abstract
World models, which simulate an environment and predict how it changes under actions, are increasingly used in real-world applications such as robotics. Complex systems call for the same tool because the effect of an action is not immediate. Seeding nodes for a campaign, or immunizing nodes against an epidemic, changes little on its own; what matters is the outcome that unfolds over the steps that follow. Designing an algorithm that selects such actions to maximize expected performance on a task is inherently iterative, and every candidate must be scored by the outcome it produces. Obtaining that outcome has relied on simulation, whose cost becomes a bottleneck when candidates are evaluated over many sampled trajectories. We propose an action-conditioned Network World Model that learns a network's diffusion dynamics under interventions over time, applies each action to the network, and predicts the outcome that follows. It serves as a fast evaluator inside an algorithm design loop in which a coding agent designs and refines executable algorithms using feedback from full rollouts, action-level credit, and counterfactual probes over alternative interventions. Across eight network tasks and five diffusion models, the designed algorithms match or exceed the strongest reported baseline in 138 of 141 settings while enabling up to 14.5 times faster rollouts than Monte Carlo simulation. Code will be released upon acceptance.
Chinese Translation
世界模型模拟环境并预测其在动作下如何变化,正越来越多地用于机器人等现实世界应用中。复杂系统也需要同样的工具,因为一个动作的效果并非即时显现。为一场活动选择种子节点,或为抵御流行病而对节点进行免疫,其自身带来的变化很小;重要的是在随后各步中展开的结果。设计一种选择此类动作以最大化任务上预期性能的算法本质上是迭代的,并且每个候选方案都必须由其所产生的结果来评分。获得该结果一直依赖于模拟,当候选方案在大量采样轨迹上进行评估时,其成本会成为瓶颈。我们提出一个动作条件化的网络世界模型,它学习网络在随时间变化的干预下的扩散动力学,将每个动作应用于网络,并预测随后产生的结果。它在一个算法设计循环中充当快速评估器,在该循环中,一个编码智能体利用来自完整推演、动作级信用以及针对替代干预的反事实探测的反馈,来设计和改进可执行算法。在八项网络任务和五个扩散模型上,所设计的算法在 141 个设置中的 138 个中达到或超过了已报告的最强基线,同时实现比蒙特卡洛模拟快高达 14.5 倍的推演。代码将在论文被接收后发布。
cs.AI / 13 / 2610.01080
Improving Math Reasoning through Value-guided Informative Search
通过价值引导的信息性搜索提升数学推理
large language model
大语言模型相关
Abstract
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, allowing improvements found by search to produce informative relative rewards. It further applies selective supervision to search-improved tokens, preserving a learning signal when uniform group rewards render GRPO ineffective. We show that exact value-guided selection improves the expected verifier reward at each searched state and that this guarantee extends to the complete rollout policy, with a corresponding approximate guarantee under bounded value-estimation error. Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
Chinese Translation
具有可验证奖励的强化学习(RLVR)已大幅提升了大语言模型的数学推理能力。近期工作将搜索引入 RLVR 的 rollout 中以增加轨迹多样性,但仅有多样性并不能保证搜索所诱导的 rollout 策略优于当前策略。为弥补这一缺口,我们提出 APIVIS,一个训练时框架,它将有限预算的 Gumbel 搜索适配到块级数学推理。APIVIS 在每个 rollout 组内结合直接响应与搜索得到的响应,使搜索发现的改进能够产生信息丰富的相对奖励。它进一步对经搜索改进的 token 施加选择性监督,在组内奖励均匀导致 GRPO 失效时保留学习信号。我们证明,精确的价值引导选择能够提升每个被搜索状态下的期望验证器奖励,并且这一保证可推广至完整的 rollout 策略,在有界价值估计误差下也有相应的近似保证。在广泛认可的数学推理基准和不同模型规模上的实验表明,APIVIS 相较有竞争力的基于搜索的方法取得了显著提升,验证了 APIVIS 的有效性。
cs.AI / 14 / 2610.01116
Beyond State-of-the-Art: Standardising Environmental Impact Metrics for AI Research
超越最先进水平:为人工智能研究标准化环境影响指标
large language model
大语言模型相关
Abstract
As the capabilities and ubiquity of Large Language Models (LLMs) grow, so does their environmental footprint. Despite calls for responsible AI, the machine learning community lacks standardised practices for carbon accounting. Our automated literature review of the 5,285 papers accepted to NeurIPS 2025 reveals that reporting of environmental impact is nearly non-existent. To catalyse a shift toward sustainable AI, we define standardised sustainability metrics for evaluating model training efficiency, accompanied by simple heuristics to estimate the carbon cost of LLM inference. We implement these metrics in carbonbenchmark, a drop-in software solution for tracking and reporting emissions. Finally, to combat the pursuit of marginal accuracy gains at disproportionate environmental costs, we formalise the `Smallest Model that Achieves the Job' (SMAJ), a framework which challenges the field to prioritise computational efficiency and environmental accountability alongside traditional `State-of-the-Art' (SotA) accuracy.
Chinese Translation
随着大型语言模型(LLMs)的能力和普及程度不断增长,其环境足迹也随之增长。尽管人们呼吁负责任的人工智能,机器学习社区仍缺乏碳核算的标准化实践。我们对被 NeurIPS 2025 接收的 5,285 篇论文进行的自动化文献综述显示,环境影响的报告几乎不存在。为促进向可持续人工智能的转变,我们定义了用于评估模型训练效率的标准化可持续性指标,并辅以简单的启发式方法来估算 LLM 推理的碳成本。我们将这些指标实现在 carbonbenchmark 中,这是一个用于跟踪和报告排放的即插即用软件解决方案。最后,为了对抗以不成比例的环境代价追求边际准确率提升的做法,我们将“能够完成任务的最小模型”(Smallest Model that Achieves the Job,SMAJ)形式化,这是一个框架,它挑战该领域在传统“最先进水平”(State-of-the-Art,SotA)准确率之外,同时优先考虑计算效率和环境责任。
cs.AI / 15 / 2610.01119
AbsorbEvo: An Agentic Framework for Autonomous Inverse Design of Microwave Absorbers
AbsorbEvo:一种用于微波吸收体自主逆向设计的智能体框架
large language model
大语言模型相关
Abstract
Designing high-performance microwave absorbers requires specialized expertise in electromagnetic theory, materials science and simulation programming, and entails time-consuming optimization. Here, we present AbsorbEvo, an agentic framework for autonomous inverse design that translates natural-language performance objectives into designs verified by full-wave simulations. Its candidate evolution strategy integrates language reasoning, physics-based prediction and historical feedback. A large language model proposes the directions and magnitudes of parameter adjustments based on task objectives and computational history. The system combines directed increments with global sampling to generate candidates and uses a low-cost predictive model as a physics prior to rank them. Only high-ranking designs undergo full-wave simulation. Results passing physical validity checks are used to evaluate performance and guide subsequent search. Experience from training tasks is further distilled into textual skills, which are independently validated before use in new tasks. Under identical proposal budgets on held-out AbsorbBench-36 tasks, AbsorbEvo achieved a task success rate of 79.17%, versus 25.00% for a generic agent and 12.50% for random search. Its mean best coverage was 0.7816, compared with 0.6434 and 0.6448, respectively. By integrating language reasoning and physics-based feedback into design decisions, AbsorbEvo provides a methodological foundation for natural-language-driven autonomous inverse design of microwave absorbers.
Chinese Translation
设计高性能微波吸收体需要电磁理论、材料科学与仿真编程方面的专业知识,并且涉及耗时的优化。在此,我们提出 AbsorbEvo,一个用于自主逆向设计的智能体框架,它将自然语言性能目标转化为经全波仿真验证的设计。其候选演化策略集成了语言推理、基于物理的预测和历史反馈。一个大型语言模型基于任务目标和计算历史,提出参数调整的方向和幅度。该系统将有向增量与全局采样相结合来生成候选,并使用低成本预测模型作为物理先验对它们进行排序。只有排名靠前的设计才会进行全波仿真。通过物理有效性检查的结果被用于评估性能并指导后续搜索。来自训练任务的经验进一步被提炼为文本技能,这些技能在用于新任务之前会经过独立验证。在留出的 AbsorbBench-36 任务上,在相同的提议预算下,AbsorbEvo 取得了 79.17% 的任务成功率,而通用智能体为 25.00%,随机搜索为 12.50%。其平均最佳覆盖率为 0.7816,而相比之下两者分别为 0.6434 和 0.6448。通过将语言推理和基于物理的反馈整合到设计决策中,AbsorbEvo 为自然语言驱动的微波吸收体自主逆向设计提供了方法论基础。
cs.AI / 16 / 2610.01128
Grounding Large Language Models in DSGE Simulators for Policy Generation and Forecasting
将大型语言模型锚定在 DSGE 模拟器中以进行政策生成与预测
large language model
大语言模型相关
Abstract
Large language models can produce economic policy responses that sound reasonable, but this does not show that their actions are consistent with economic dynamics. We test this by placing an instruction-tuned language model inside six Snowdrop-backed dynamic stochastic general equilibrium (DSGE) simulators. At each turn, the model observes the economy and a change in economic discourse, selects a bounded policy action, and receives the next simulated state and an economic reward. We implement a common Python interface for repeated rollouts, persistent shocks, state cloning, and rolling-horizon simulation. This setting creates a long-horizon credit-assignment problem. Policy effects may appear several quarters after an action is taken. PPO has a learned value function that can propagate delayed reward to earlier tokens through generalized advantage estimation. GRPO has no learned value function and instead assigns a group-relative advantage from complete rollout returns. It therefore cannot distinguish which earlier turn caused the outcome; if every rollout receives the same return, the normalized advantage is zero. We use PPO as the primary method and GRPO as a matched critic-free baseline. The experiments also test directional semantic signals, reward horizon, trajectory warm starts, cross-simulator transfer, and historically anchored pandemic and monetary-policy shocks. The objective is to judge policy actions by their simulated economic consequences rather than by plausible language alone.
Chinese Translation
大型语言模型可以生成听起来合理的经济政策回应,但这并不能表明其行动与经济动态相一致。我们通过将一个指令微调语言模型置于六个由 Snowdrop 支持的动态随机一般均衡(DSGE)模拟器中来检验这一点。每一轮,模型观察经济状况以及经济话语的变化,选择一个有界政策动作,并接收下一个模拟状态和经济奖励。我们实现了一个通用 Python 接口,用于重复 rollout、持续冲击、状态克隆和滚动时域模拟。这一设定产生了长时域信用分配问题。政策效果可能在采取行动后的数个季度才显现。PPO 具有学习到的价值函数,可以通过广义优势估计将延迟奖励传播到更早的 token。GRPO 没有学习到的价值函数,而是根据完整 rollout 回报分配组相对优势。因此,它无法区分是哪个较早轮次导致了结果;如果每个 rollout 都获得相同的回报,则归一化优势为零。我们使用 PPO 作为主要方法,并将 GRPO 作为匹配的无评论家基线。实验还测试了方向性语义信号、奖励时域、轨迹热启动、跨模拟器迁移以及历史锚定的疫情和货币政策冲击。目标是根据模拟的经济后果来评判政策动作,而不仅仅是根据听起来合理的语言。
cs.AI / 17 / 2610.01138
Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay
LLM 智能体环境中的动作结算审计:顺序、进度与重放
large language model
大语言模型相关
Abstract
Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Joint policies are spatially order-invariant conditional on fixed priorities, yet conservative rejection completes only 31.25% of agents in a six-agent doorway task versus 90.28% for random tickets; the paired improvement is 59.03 percentage points (95% bootstrap interval: 50.00-68.06). All policies preserve the tested spatial constraints, and priority arbitration still misses the independent small-instance optimum. A separate full-state journal audit exactly replays 156 checkpoints and rejects 1,332 constructed corruptions with a retained terminal anchor. The evidence concerns execution semantics, not human realism or long-run fairness.
Chinese Translation
在大语言模型(LLM)智能体环境中,并发动作即使每个提案单独来看都是有效的,也仍然需要仲裁。我们实现了一种带类型的快照结算契约,并审计三个不同的性质:顺序敏感性、有用进度和重放一致性。五种结算策略在 28,800 次穷举排列试验和 2,160 个脚本化多步回合中接受了测试。联合策略在固定优先级条件下具有空间顺序不变性,然而保守拒绝在六智能体门道任务中仅完成 31.25% 的智能体,而随机票号则为 90.28%;配对改进为 59.03 个百分点(95% 自助法区间:50.00-68.06)。所有策略都保持了所测试的空间约束,而优先级仲裁仍然未能达到独立小规模实例的最优值。一项单独的全状态日志审计精确重放了 156 个检查点,并以保留的终端锚点拒绝了 1,332 个构造的损坏。这些证据关注的是执行语义,而非人类真实性或长期公平性。
cs.AI / 18 / 2610.01195
Federated Agent Optimization
联邦智能体优化
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbf{Federated Agent Optimization (FAO)}, which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.
Chinese Translation
大语言模型(LLM)智能体日益在私有环境中运行,并从任务执行、工具使用、反馈和本地知识中积累宝贵经验。然而,此类经验分散在不同组织之间,并且由于隐私和专有约束而无法直接共享。传统的联邦学习不足以应对这一场景,因为智能体能力不仅限于模型参数,还延伸到记忆、工具、奖励、技能以及结构化知识。在本文中,我们提出 \textbf{联邦智能体优化(FAO)},它研究分布式智能体如何通过受控的信息交换协同提升,同时将原始数据、完整轨迹和私有知识保持在本地。我们将 FAO 定义为一个多目标问题,在智能体效用、隐私泄露和通信成本之间进行权衡,并将其优化空间组织为策略、记忆、工具使用、奖励以及结构化知识和技能等方面。我们进一步刻画了私有经验如何被抽象、保护、聚合并适配为可迁移能力,从而提供一种统一视角,说明智能体如何在没有直接经验共享的情况下彼此受益。最后,我们指出 FAO 的关键挑战,并概述若干有前景的未来研究方向,以迈向可信的联邦智能体系统。
cs.AI / 19 / 2610.01207
Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
面向智能体强化学习的依赖感知奖励塑形
large language model
大语言模型相关
Abstract
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO's entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at https://github.com/JianhuiWei7/DARS.
Chinese Translation
当使用强化学习训练大型语言模型时,终端奖励几乎没有提供关于哪些步骤重要的指导。常见的步骤信用分配方法忽略了,建立在未纠正错误之上的工作会被浪费,而独立的工作仍然有效。在仅有最终成功/失败奖励的情况下,失败回合中的每一步总未来奖励都为零,即使它取得了进展。我们提出依赖感知奖励塑形(DARS),它将任务进展表示为通过先修关系连接的谓词,并在依赖图上分配步骤级信用。标注器标记每一步验证、使失效或修复哪些谓词。已验证谓词根据距最近的被破坏先修条件的图距离进行折扣,而独立谓词不受影响。修复会根据仍然存在的任何错误更新这些权重;已失效谓词需要重新验证才能重新获得信用。一个固定势函数将这些标注转换为带符号的每步奖励。一个通用的奖励与标注接口使 DARS 能够与一系列推理和智能体训练方法集成,例如 GiGPO 和 ARPO/AEPO,而无需改变它们的 rollout 策略或优化器。在五个任务族以及从 1.5B 到 8B 的模型上,DARS 相比于在相同预算和运行框架下训练的 GiGPO,在 ALFWorld 上将成功率提升最多 10 个百分点,提高了 WebShop 任务得分和 Search-R1 QA 准确率,在带 Python 解释器的 AIME24/25 上补充了 AEPO 基于熵的训练,并在 1.7B 和 4B 的受控无工具推理比较中超过 OmniOPD。消融实验表明,步骤级信用、依赖衰减和图拓扑各自都有贡献。在 ALFWorld 上,一个蒸馏的 8B 标注器与 API 标注器相匹配,使 DARS 无需前沿评判器即可高效运行。代码见 https://github.com/JianhuiWei7/DARS。
cs.AI / 20 / 2610.01230
HHR: Hierarchical Hash Retrieval for Efficient LLM Generation
HHR:面向高效LLM生成的分层哈希检索
large language model
大语言模型相关
Abstract
Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a critical mismatch between Hamming distance and attention relevance. Query-Key logits depend jointly on directional similarity and feature magnitudes, whereas hash binarization discards magnitude information, causing both false-positive retrieval of low-logit keys and false-negative omission of high-logit keys. To address these failures, we propose Hierarchical Hash Retrieval (HHR), a coarse-to-fine framework that progressively improves retrieval accuracy through Geometry-Aware Key Routing (GKR) and Learned Hash Projection (LHP). GKR learns a head-wise orthogonal transformation to redistribute feature magnitudes and derive more discriminative page-level logit bounds, enabling effective pruning of low-logit keys while preserving important candidates. LHP then learns a head-wise projection space that aligns Hamming distance with the true Query-Key relevance ranking for fine-grained retrieval. By combining GKR and LHP, HHR suppresses false positives and recovers false negatives, substantially improving the fidelity of hash-based sparse attention. Extensive experiments across diverse LLMs and benchmarks demonstrate that HHR achieves superior performance over existing methods. For example, on LongBench, HHR improves the average score by 1.10 points and, at a context length of 128K, achieves up to a 3.30x decoding speedup and a 2.83x end-to-end speedup for Llama-3.1-8B-Instruct. The code is publicly available at https://github.com/lianjunl13-sudo/HHR.
Chinese Translation
高效的长上下文推理对于大语言模型(LLMs)至关重要,但它构成了严重的计算瓶颈。基于哈希的检索通过将查询和键编码为二进制码,并使用汉明距离进行键选择,提供了一种高效的替代方案。然而,这导致汉明距离与注意力相关性之间存在关键的不匹配。查询-键 logits 同时取决于方向相似度和特征幅值,而哈希二值化丢弃了幅值信息,从而导致低 logit 键的假阳性检索和高 logit 键的假阴性遗漏。为了解决这些失败,我们提出分层哈希检索(HHR),这是一个由粗到细的框架,通过几何感知键路由(GKR)和学习哈希投影(LHP)逐步提高检索准确率。GKR 学习逐头正交变换,以重新分配特征幅值并推导出更具判别性的页级 logit 边界,从而能够有效剪枝低 logit 键,同时保留重要候选。随后,LHP 学习一个逐头投影空间,使汉明距离与真实的查询-键相关性排序对齐,以实现细粒度检索。通过结合 GKR 和 LHP,HHR 抑制假阳性并恢复假阴性,显著提高了基于哈希的稀疏注意力的保真度。在不同 LLMs 和基准上的大量实验表明,HHR 相比现有方法取得了更优的性能。例如,在 LongBench 上,HHR 将平均分数提高了 1.10 分,并且在 128K 的上下文长度下,对于 Llama-3.1-8B-Instruct 实现了高达 3.30 倍的解码加速和 2.83 倍的端到端加速。代码公开可用,地址为 https://github.com/lianjunl13-sudo/HHR。
cs.AI / 21 / 2610.01236
Learning to Ask: Information Acquisition for SLM-LLM Collaboration, under a budget
学会提问:预算约束下面向 SLM-LLM 协作的信息获取
large language model
大语言模型相关
Abstract
Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model should handle each portion of the reasoning process. In black-box API-based settings, however, this paradigm can be inefficient due to coarse-grained delegation or repeated transmission of context across model switches. In this work, we instead formulate SLM-LLM collaboration as an information acquisition problem, under an API budget constraint. The SLM remains the primary reasoner and selectively queries a black-box LLM advisor only when needed, issuing targeted queries rather than delegating the reasoning process itself. To realize this strategy, we develop a three-stage RLVR framework that learns whether to call the advisor, how to formulate useful queries, and how to integrate the collaboration into the reasoning process by jointly refining advisor invocation and information use. Across mathematical reasoning and coding tasks, our approach improves the performance--cost tradeoff over existing collaboration baselines and, in some settings, matches or exceeds oracle problem-level routing. Finally, we show that our strategy can transfer to other advisor model families, without further training.
Chinese Translation
小语言模型(SLM)与大语言模型(LLM)之间的协作,为将较小模型的高效性与较大模型的强大推理能力结合起来提供了机会。现有方法主要将这种协作刻画为一个计算分配问题,即决定推理过程的每一部分应由哪个模型来处理。然而,在基于黑盒 API 的设置中,由于粗粒度的委派或模型切换时上下文的重复传输,这一范式可能效率低下。在本工作中,我们转而在 API 预算约束下,将 SLM-LLM 协作表述为一个信息获取问题。SLM 仍然是主要的推理者,仅在需要时选择性地查询一个黑盒 LLM 顾问,发出有针对性的查询,而不是将推理过程本身委派出去。为实现这一策略,我们开发了一个三阶段 RLVR 框架,通过联合优化顾问调用与信息使用,学习是否调用顾问、如何构建有用的查询,以及如何将协作整合进推理过程。在数学推理与编程任务上,我们的方法相较现有协作基线改善了性能—成本权衡,并在某些设置下达到或超过 oracle 的问题级路由。最后,我们表明我们的策略可以迁移到其他顾问模型家族,而无需进一步训练。
cs.AI / 22 / 2610.01296
ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models
ITC-MoE:面向MoE扩散语言模型的重要性引导Token感知压缩
diffusion
扩散模型相关
Abstract
Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at https://github.com/lianjunl13-sudo/ITC-MoE.
Chinese Translation
混合专家(MoE)扩散语言模型(DLM)提供了灵活的并行解码与更高的模型容量,但其数量庞大的专家参数带来了可观的计算与存储开销。现有的低秩MoE压缩方法在很大程度上依赖静态分解和固定秩分配,忽视了MoE扩散语言模型独有的特性。具体而言,我们识别出两个特性:跨模式非均匀冗余,即参数冗余以及对秩截断的敏感性在输入、输出和专家模式之间各不相同;以及逐Token利用变化,即热点Token与冷门Token表现出不同的谱特性和专家激活模式。为应对这些挑战,我们提出了ITC-MoE,一个面向MoE扩散语言模型的重要性引导Token感知压缩框架。ITC-MoE由两个互补的组件构成。首先,重要性引导自适应Tucker压缩(IATC)将激活重要性和梯度重要性纳入专家权重变换,跨多个模式联合分解专家权重,并在固定参数预算下自适应地分配秩。其次,Token感知补偿与路由(TCR)对压缩敏感的热点Token施加轻量级低秩补偿,并为路由模式集中的冷门Token限制候选专家集合。通过使压缩能力与推理执行同时适应参数冗余和逐Token变化,ITC-MoE在保持生成质量的同时大幅降低了MoE扩散语言模型的计算与存储开销。例如,在SDAR-30B-A3B-Chat-b32上,ITC-MoE在30%压缩预算下于MultiArith上保持了96.33%的准确率,同时实现了最高7.22倍的端到端加速。代码已公开于 https://github.com/lianjunl13-sudo/ITC-MoE。
cs.AI / 23 / 2610.01320
ProtoFlow: Prototype-Guided Flow Matching for Multivariate Time Series Forecasting
ProtoFlow:原型引导的流匹配用于多变量时间序列预测
diffusion
扩散模型相关
Abstract
Generative modeling has shown strong promise for multivariate time mseries (MTS) forecasting, especially scale to high-dimensional settings. Diffusion-based methods achieve competitive performance but typically require many sampling steps at inference. VAE-based non-iterative forecasting frameworks have therefore emerged as an efficient alternative. Within this line of work, vector quantization (VQ) enables controllable latent space modeling by mapping multivariate series into compact discrete representations. Existing VQ-based forecasting methods, however, typically rely on autoregressive (AR) token generation, which suffers from exposure bias and training-inference mismatch. Flow matching provides an efficient non-autoregressive alternative for latent forecasting, but existing formulations usually initialize transport from a generic Gaussian prior. We instead observe that the trained VQ codebook already captures representative latent prototypes and can thus serve as a more informative prior for flow matching. Based on this insight, we propose ProtoFlow, a forecasting framework that combines vector-quantized autoencoding with Prototype-prior Flow matching. Our method first maps multivariate sequences into a discrete latent space, then constructs a structured prior from the learned codebook, and finally learns a DiT-based rectified flow to transport samples from this prior to future latent representations conditioned on historical observations. By replacing generic noise initialization with a learned prototype prior, ProtoFlow avoids the rollout mismatch of AR token prediction and promotes faster training convergence. Extensive experiments on benchmark datasets show that it consistently achieves superior forecasting performance with efficient inference.
Chinese Translation
生成式建模在多变量时间序列(MTS)预测中已展现出强大的前景,尤其是在扩展到高维设置时。基于扩散的方法取得了有竞争力的性能,但在推理时通常需要大量的采样步骤。因此,基于 VAE 的非迭代预测框架作为高效的替代方案而出现。在这一研究路线中,向量量化(VQ)通过将多变量序列映射为紧凑的离散表示,实现了可控的隐空间建模。然而,现有的基于 VQ 的预测方法通常依赖自回归(AR)的 token 生成,这会遭受暴露偏差以及训练-推理不一致的问题。流匹配为隐空间预测提供了一种高效的非自回归替代方案,但现有的形式通常从通用的高斯先验初始化传输。我们则观察到,训练好的 VQ 码本已经捕获了具有代表性的隐原型,因此可以作为流匹配中信息更丰富的先验。基于这一洞见,我们提出 ProtoFlow,一个将向量量化自编码与原型先验流匹配相结合的预测框架。我们的方法首先将多变量序列映射到离散隐空间,然后从学习到的码本中构建结构化先验,最后学习一个基于 DiT 的整流流,以历史观测为条件,将样本从该先验传输到未来的隐表示。通过用学习到的原型先验替代通用噪声初始化,ProtoFlow 避免了 AR token 预测的 rollout 不一致问题,并促进了更快的训练收敛。在基准数据集上的大量实验表明,它在高效推理的同时始终取得更优的预测性能。
cs.AI / 24 / 2610.01323
TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
TRACE:面向多轮安全的轨迹回报归因与对比擦除
large language model
大语言模型相关
Abstract
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1\.23 points. Source code can be found in the supplemental material.
Chinese Translation
安全对齐的大语言模型(LLMs)常常拒绝一个有害请求,但一旦同一目标被分散到若干轮中,就会予以遵从。偏好目标对单个提示的完整响应进行评分,因此仅其训练损失无法控制未见历史上的风险。我们的分析给出了充分条件,在这些条件下,在监督式单轮上下文中的抑制会产生对多轮轨迹风险的界。该界考虑了覆盖度、迁移松弛和泄漏,并刻画了相对于在训练策略的上下文上评估的基础策略风险预算的收缩。TRACE(Trajectory Return Attribution and Contrastive Erasure,轨迹回报归因与对比擦除)将这一原则转化为词元级目标。在安全响应上,每个词元都由一个可归因于拒绝的优势的折现回报进行加权。该优势将一个冻结的参考模型与其拒绝能力被消融的副本进行比较,从而使较早的响应词元能够从较晚的拒绝相关证据中获得信用。在被拒绝响应的高差距位置上,TRACE 在擦除目标中将观察到的词元与策略选择的替代词元相结合。一个梯度范数惩罚替代了保留集。在五个开放权重模型和七种多轮攻击上,TRACE 在所有 35 个模型与攻击对中给出最低的攻击成功率(ASR),而在 MMLU 和 HellaSwag 上评估的模型效用最多下降 1\.23 个点。源代码可在补充材料中找到。
cs.AI / 25 / 2610.01415
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
超越记忆:以显式信念状态驾驭长时程智能体
large language model
大语言模型相关
Abstract
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Chinese Translation
大语言模型(LLM)智能体如今能够承担日益复杂的任务,但它们将交互历史组织为记忆的方式并不能确保对当前世界形成连贯的理解。我们提出 PoS,一个推理时框架,它构建并持续维护显式信念状态,将其作为智能体的决策上下文。每个信念都将对当前世界状态的估计与尚未解决的任务需求结合起来,从而明确智能体仍需学习和完成的内容。为使该信念保持可靠且可操作,PoS 验证其一致性并监控任务进展,以检测信念陷阱(Belief Trapping),即智能体持续行动却未朝目标取得实质性进展。随后,恢复过程会针对陷阱模式与未解决任务需求的类型进行定制。在涵盖执行与诊断的四个基准上的实验表明,PoS 在所有三个 LLM 骨干模型下、于每个基准上均取得了最高的总体性能。消融实验证明了一致性验证与恢复的重要性,而上下文扩展实验则显示出对上下文增长的韧性。这些结果共同支持将信念构建与持续维护作为超越历史保留与压缩的长时程上下文管理的基础。
cs.AI / 26 / 2610.01439
DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair
DRelay:面向前缀感知的并行投机解码修复的全局草稿上下文
large language model
大语言模型相关
Abstract
Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections before target-model verification. DRelay bases its decisions on candidate correlations and the selected path: a global reader extracts predictive information across positions for each candidate. While a causal selector combines candidate-level information extracted by the global read with the tokens selected at preceding positions to determine whether the native choice at the current position is consistent with the global evidence and the selected prefix. It then decides whether to retain or replace the token, thereby repairing early errors and extending the accepted prefix. We further jointly train the draft backbone and the selector, combining candidate-support learning with a repair objective, while weighting the repair loss according to each block position's potential contribution to the consecutive accepted prefix. Across eight diverse benchmarks on an H800 GPU, DRelay consistently improves both average acceptance length and end-to-end decoding performance over DFlash, Domino, and DSpark. Under SGLang serving, DRelay improves average end-to-end speedup over DFlash, Domino, and DSpark by 14.7%-16.8%, 8.7%-9.3%, and 8.1%-9.3%, respectively.
Chinese Translation
并行草稿降低了大型语言模型(LLM)投机解码的草稿开销,但其收益仍受限于被接受的 prefix 长度。即使正确 token 存在于候选池中,一次早期的选择错误也会阻止后续预测被使用。我们提出 DRelay,它利用来自整个草稿块的全局信息,在目标模型验证之前对候选选择执行前缀感知的选择性修复。DRelay 的决策基于候选相关性以及所选路径:一个全局读取器为每个候选提取跨位置的预测信息。而一个因果选择器则将由全局读取所提取的候选级信息与在前序位置所选择的 token 相结合,以判断当前位置的原始选择是否与全局证据及所选前缀一致。随后它决定是保留还是替换该 token,从而修复早期错误并延长被接受的 prefix。我们进一步联合训练草稿主干与选择器,将候选支持学习与修复目标相结合,同时根据每个块位置对连续被接受 prefix 的潜在贡献来对修复损失进行加权。在 H800 GPU 上的八个多样化基准测试中,DRelay 在平均接受长度和端到端解码性能上均持续优于 DFlash、Domino 和 DSpark。在 SGLang 服务下,DRelay 相较于 DFlash、Domino 和 DSpark 的平均端到端加速分别提升了 14.7%-16.8%、8.7%-9.3% 和 8.1%-9.3%。
cs.AI / 27 / 2610.01451
A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance
面向个性化健康体检解读与指导的多智能体LLM框架
large language model
大语言模型相关
Abstract
Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.
Chinese Translation
健康体检结果的个性化解读需要在纵向记录、医学知识、生活方式指导与就医导航之间进行推理。我们提出一个多智能体大语言模型(LLM)系统,它识别多种意图,将每种意图映射到特定任务的智能体,并行执行它们,并综合它们的输出。我们使用合成健康体检记录,在120个结合了两到四项需求的韩语复合查询上,比较了在单智能体(Single Agent)与多智能体(Multi Agent)设置下生成的答案。多智能体将加权LLM评判分数从1.695提高到1.797(p = 0.027),另外三个LLM评判者显示出持续改善($Δ$ = +0.111 至 +0.186,所有 p < 0.05)。这些增益来自有用性、一致性以及对复合查询中每一项需求的处理,而数值准确性和依据性仅在四位评判者中的一位下显著改善,医学安全性没有差异,严重失败的发生率相近(单智能体15.0% vs. 多智能体13.3%)。两名人类评估者在66.7%和68.3%的成对比较中偏好多智能体。多智能体执行将延迟和成本分别增加了1.31$\times$和2.02$\times$。在探索性亚组分析中,改善集中在涉及个人记录查询的查询上。
cs.AI / 28 / 2610.01497
OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
OpenMTB-Audit:揭示基于LLM的分子肿瘤委员会安全性评估中的过度拒绝与临床专家视角
large language model
大语言模型相关
Abstract
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
Chinese Translation
分子肿瘤委员会整合基因组发现、临床背景与治疗证据,以支持精准肿瘤学。随着AI进入这一工作流程,一个关键的安全性挑战在于区分真正缺乏支持的推荐与有证据支持的选项,后者由于信息不完整、ECOG体能状态较差或其他临床注意事项而仍然需要肿瘤科医生审查。我们提出OpenMTB-Audit,这是一个开源基准,包含500个合成的非小细胞肺癌病例,涵盖五个对抗性错误类别和四个安全性标签:Supported(有支持)、Partially Supported(部分支持)、Unsupported(无支持)和Insufficient Information(信息不足)。在八种大语言模型配置中,我们发现普遍存在过度拒绝:所有LLM配置在83.3-100%的真实Partially Supported病例中都未能保留Partially Supported标签,它们通过标签坍缩而非经临床校准的推理获得了较高的总体安全性得分。为解决这一局限,我们开发了MTB-AuditAgent,这是一个确定性的七模块框架,将证据核验、缺失信息检测、安全性分类与弃权分离开来。它将过度拒绝降低至6.7%,并达到91.2%的准确率(95% CI:88.6-93.6%)。一项由两位肿瘤科医生进行的标注研究发现,分歧集中在信息充分性与治疗优化之间的边界处,这凸显了保留具有临床意义的区分的必要性。
cs.AI / 29 / 2610.01506
MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills
MCRI:一个用于分析和评估智能体技能的四维框架
large language model
大语言模型相关
Abstract
As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.
Chinese Translation
随着智能体从单工具系统演化为模块化、组合式架构,技能正成为能力开发与分发的重要机制。然而,学术界缺乏一个用于系统分析和评估技能的结构化框架。借鉴信息增益和行为约束,我们提出四维 MCRI 框架,并将其操作化为 MCRI-Eval,一种基于大语言模型的评估方法。我们使用来自 OpenClaw skill Hub 的 63,812 个公开技能评估 MCRI-Eval,并在 BigCodeBench、BFCL-Fundamental 和 Mind2Web 上进行了 58,275 次以技能为条件的模型执行。MCRI-Eval 分数与社区流行度信号正相关,并在所评估的方法中取得最高的下游排名一致性。MCRI-Eval 还在所有三个基准上改进了 top-1 技能选择:与每个基准上最强的基线相比,MCRI-Eval 所选择的技能在 BigCodeBench、BFCL-Fundamental 和 Mind2Web 上的下游性能排名分别提升了 17.7、22.8 和 19.6 个百分位点。这些结果表明,MCRI-Eval 提供了一种有用的执行前信号,用于在昂贵的基于执行的评估之前优先考虑有前景的技能。
cs.AI / 30 / 2610.01509
Sharpening Tax in Post-Training
后训练中的锐化税
large language model
大语言模型相关
Abstract
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
Chinese Translation
关于大型语言模型(LLMs)的强化学习(RL)后训练,一个新兴假设是:它仅仅锐化基础模型的现有行为,以解覆盖率为代价提高单次准确率。虽然这种权衡已在数学和编码任务中被观察到,但它未必会扩展到智能体任务,因为在智能体任务中,多轮工具使用和交互可能需要后训练期间新获得的能力。我们令人惊讶的发现是,预训练 LLMs 在配备轻量级推理框架后,可以充当有能力的智能体。尽管准确率(pass@1)低得多,但在给定充足的测试时预算时,它们在解覆盖率(pass@K)上常常超过经过后训练的对应模型。我们进一步分析其底层机制,并表明后训练将任务推向两个极端:总是被解决或从未被解决,从而以解覆盖率为代价提高采样效率和一致性。为了衡量这一代价,我们提出锐化税(Sharpening Tax),一种诊断指标,用于量化后训练之后测试时可扩展性的损失。在来自四个模型家族的 14 对基础/后训练模型以及三个智能体基准(总共 42 个案例)中,该税在大多数设置中普遍存在,可以从少量 rollout 中估计,并且与其他指标良好相关。最后,我们提出后验调温组采样(posterior-tempered group sampling,PTGS),这是一种简单的即插即用贝叶斯采样器,它会根据每个提示的估计难度调整采样温度。在两个智能体环境中的 RL 训练期间应用时,PTGS 比固定温度基线支付更小的税,在重复采样下解决更多任务,同时还提高单次准确率。
cs.AI / 31 / 2610.01763
TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference
TopK-Guided:面向高效 LLM 推理的自适应、预算感知激活稀疏化
large language model
大语言模型相关
Abstract
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.
Chinese Translation
激活稀疏化通过将不重要的激活值置零,从而跳过相应的计算,加速大语言模型(LLM)推理。然而,现有的免训练方法做出了不同的权衡:诸如 TEAL 之类的基于阈值的方法使稀疏度水平适应每个 token,但无法 tightly 控制实际实现的稀疏度;而诸如 WINA 之类的基于 TopK 的方法强制使用固定的稀疏度水平,但对每个 token 使用相同的稀疏度预算。这两类方法还在所有 transformer 块上应用相同的预算,尽管各块之间的敏感度差异很大。我们提出 TopK-Guided,这是一种免训练方法,它通过将有界的 token 级稀疏度自适应与敏感度感知的块级预算分配相结合,解决了上述两方面的局限。在 Llama-2 和 Llama-3 模型上,TopK-Guided 在困惑度和下游准确率方面持续优于 TEAL 和 WINA,同时保持了与 WINA 基本相同的依赖于稀疏度的投影计算量,且在高稀疏度下收益最大。消融实验表明,这两个组成部分提供了互补的改进。
cs.AI / 32 / 2610.01800
LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
LineupRL:通过描述到时间序列识别实现时间序列描述的可验证强化学习
large language model
大语言模型相关
Abstract
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.
Chinese Translation
时间序列描述是时间序列理解中的一个基本步骤,也可以充当信号与自然语言之间的桥梁。监督微调(SFT)依赖于更大模型生成的描述,因而无法超越这些描述的质量。强化学习(RL)可以做到这一点,但其奖励是为其他模态和其他任务设计的,迁移到时间序列领域的开放式生成时表现很差。我们通过提出 LineupRL 来解决这一问题,这是一个具有可验证奖励的强化学习(RLVR)流程,其奖励是描述到时间序列的识别。奖励模型是一个冻结的大语言模型(LLM)验证器,它读取生成的描述和候选时间序列的原始数值,而从不读取图表,并且必须从多个干扰项中选出所描述的那条时间序列。与编写问题或评判一条描述相比,匹配对验证器而言是轻得多的要求,因此一个现成的 LLM 就能提供该奖励。在两个描述基准上,以及在预测器只能看到描述的预测和重构任务上,LineupRL 在每一项指标上都优于 SFT 和 RL 基线。由 LineupRL 训练的 3B 视觉语言模型(VLM)也以 1/24 的参数规模,优于 SFT 基线所蒸馏其描述的那个 72B VLM。我们的案例研究表明,LineupRL 能够抵抗奖励劫持,并且它所训练出的描述器既能刻画趋势,也能点出关键点处的数值。
cs.AI / 33 / 2610.01861
AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes
AVSD-Scenes:一个用于城市场景音视频描述的数据集
large language model
大语言模型相关
Abstract
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.
Chinese Translation
自然语言描述能够为音视频城市场景提供丰富的语义表示,然而能够同时描述听觉与视觉信息的数据集仍然有限。在本文中,我们提出了 AVSD-Scenes,一个面向城市环境的成对音视频场景描述数据集。该数据集包含 12,291 条由 TAU Urban Audio-Visual Scenes 数据集生成的音视频场景描述。为构建该数据集,我们首先分别使用 Qwen2-Audio-7B 和 Qwen2.5-VL-7B 生成基于音频和基于视觉的描述。随后,我们使用大语言模型,即 Qwen3-14B、Mistral-Small-3.2-24B-Instruct-2506 和 Gemma-3-27B-it,将这些特定模态的描述进行组合,以生成能够捕捉来自两种模态的互补信息的多模态描述。我们使用语义对齐、跨模态检索、场景分类、LLM-as-a-judge 评估以及人类主观评估对 AVSD-Scenes 进行基准测试。结果表明,与特定模态描述相比,多模态描述在提升语义对齐和跨模态检索性能的同时,仍保留了较强的场景判别信息。生成的描述在城市场景分类中达到最高 94.5% 的准确率,而融合音频、视觉和描述嵌入后,准确率进一步提升至 95.4%。此外,即使从提示指令中移除场景标签,这些描述仍保持高度的场景判别性,这表明它们捕捉的是源自音视频内容的语义信息,而不仅仅是对标签信息的反映。
cs.AI / 34 / 2610.01936
Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
绘制 RAG 全景图:效率、防御、交互性与推理的四轴分类法
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
Chinese Translation
大语言模型(LLMs)在许多任务中展现出卓越的流畅性,但受限于其静态的、参数所约束的知识,以及其容易产生信息幻觉的倾向。检索增强生成(RAG)通过将外部检索纳入生成过程来应对这些问题,使模型输出以可验证且最新的来源为依据。尽管此前的综述主要关注核心 RAG 架构和标准流程,但近期研究探索了超出这些基础设计的更广泛挑战与能力。本综述对当代 RAG 发展进行了综合且结构化的考察,将该领域组织为一个四轴分类体系:提升检索效率、增强鲁棒性与安全性、支持用户驱动和交互式工作流,以及实现多步或复杂推理。我们形式化了 RAG 框架的关键组成部分,并综述了涵盖稠密与稀疏检索、融合策略、嵌入优化以及基于强化学习的检索策略的方法,强调这些进展如何影响实际部署和系统设计。我们还综合了评估实践、领域特定应用,以及诸如 Naive、Advanced 和 Modular RAG 等架构变体。最后,我们概述了与检索质量、可靠性、领域适应、可扩展性和可解释性相关的持续挑战,并指出了构建更可靠、更可适应且更透明的 RAG 系统的机遇。
cs.AI / 35 / 2610.01963
Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
用于临床分诊的开源大型语言模型中偏见的反事实审计
large language model
大语言模型相关
Abstract
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.
Chinese Translation
急诊科(ED)分诊是一项高风险的优先级排序任务,其中人口统计学、社会经济和系统情境信息可能不适当地影响紧急程度分级。尽管开源大型语言模型(LLMs)越来越多地被考虑用于本地化和隐私保护的临床决策支持,但反事实偏见如何在模型家族、规模、医学领域模型和领域自适应模型之间变化仍不清楚。我们提出了一项针对十种开源 LLMs 用于儿科急诊严重指数(ESI)预测的比较性反事实审计。从真实的和手册式临床情景案例出发,我们构建成对的反事实变体,这些变体仅改变一个注入的人口统计学、社会经济、医疗可及性、行为、社会或系统情境变量,同时保持临床表现固定。模型包括 Qwen2.5-7B、Qwen2.5-14B-Instruct、一个经 QLoRA 微调的 Qwen2.5-7B、MedGemma 变体、MedLLaMA2-7B、GPT-OSS-20B 和 GPT-OSS-120B。我们测量任何反事实偏移、分诊不足、过度分诊、大于一个 ESI 等级的偏移、平均偏移和平均绝对偏移。反事实敏感性差异很大,并且并未随着模型规模增大或医学领域预训练而一致地降低。微调后的 Qwen2.5-7B 显示出最低的总体敏感性,其任何偏移率为 5.27%,平均绝对偏移为 0.0534,而基础模型分别为 16.02% 和 0.1706。若干更大的或医学领域模型表现出更显著的偏移。分层分析和相关分析进一步揭示了具有临床重要性的方向性,以及被总体比率所掩盖的共同失败模式。这些发现支持将反事实审计作为一种轻量级、临床可解释的框架,用于在临床部署前比较开源 LLMs 中的公平性风险。
cs.AI / 36 / 2610.02005
Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries
计数动作,权衡声音:面向持续对抗者下校准多LLM委员会的贝叶斯辩证论证
large language model
大语言模型相关
Abstract
A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council's \emph{typed} moves---who proposed, challenged, or conceded which answer---as observations of a classical annotator model with \emph{per-agent} reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.
Chinese Translation
一个多LLM委员会让若干大语言模型(LLM)就一个问题进行审议,并返回一个答案以及一个置信度估计。随着这些系统越来越多地用于推理,该置信度应表示一个经过校准的正确概率,并且当一些智能体持续不可靠时,决策应保持稳健。现有的委员会聚合方法在这两方面都失败:它们的置信度估计衡量的是决断性而非正确性,并且它们无法识别或对持续不可靠的智能体进行降权。我们提出贝叶斯辩证论证(Bayesian Dialectical Argumentation,BDA),它将委员会的类型化动作——谁提出、质疑或承认了哪个答案——视为一个具有各智能体可靠性的经典标注者模型的观测。这一形式化将多智能体审议重新表述为一个可靠性估计问题,利用审议轨迹来推断在持续对抗行为下的智能体可靠性。通过根据推断出的智能体可靠性对证据进行加权,BDA在候选答案上产生经过校准的后验概率,同时允许持续不可靠的智能体被反转,而不仅仅是被多数票压倒。在二分类和多分类基准上,BDA在零成本委员会聚合方法中实现了最佳校准,不需要额外的LLM调用,并且在持续对抗性联盟下提高了稳健性,同时在干净设置中保持竞争力。
cs.AI / 37 / 2610.02023
SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL
SPHERE:通过LLM增强的空间偏好学习与人在回路强化学习实现的自适应VR室内场景生成
large language model
大语言模型相关
Abstract
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user's final edited scenes. A mixed-design user study ($N=42$) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: https://github.com/hyeonmin11/SPHERE
Chinese Translation
尽管大型语言模型(LLMs)推动了3D室内场景合成的发展,但当前的流程无法在跨会话之间保留用户特有的偏好,使得沉浸式创作成为一个重复且令人身体疲劳的过程。我们提出SPHERE,一个自适应VR生成框架,它将孤立的合成转变为持续的人机协同创作。SPHERE从自然的多模态交互(语音与控制器编辑)中提取持续性的空间偏好。为确保在面对空间失真时的几何鲁棒性,它将这些原始编辑抽象为层次化约束,对局部功能语境与全局拓扑语境同时进行建模。此外,一种人在回路的强化学习机制会根据用户最终编辑后的场景动态更新检索策略。一项混合设计用户研究($N=42$)和离线消融实验表明,SPHERE显著减少了纠正性编辑和身体负荷,防止对浅层物体级特征的偏向,从而生成几何鲁棒、与用户画像对齐的布局。最终,SPHERE展示了捕捉所演示的空间逻辑如何实现可控的空间适应,为沉浸式创作建立了一个可靠、可治理的人机协作框架。项目主页与源代码将在以下网址提供:https://github.com/hyeonmin11/SPHERE
cs.AI / 38 / 2610.02038
Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control
Mimir:面向长时程灌溉控制的物理接地 LLM 智能体
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.
Chinese Translation
大语言模型(LLM)智能体日益将推理、工具使用和行动结合起来,但大多数证据来自具有相对即时反馈和可重置失败的回合式任务。长时运行的物理控制运行于不同的机制:动作改变未来状态,错误跨决策累积,智能体必须从经验中改进,而不被允许改写使执行安全的物理规则。我们通过灌溉研究这一机制,其中每日决策在整个生长季内与土壤-水动力学相互作用。我们提出 Mimir,一个围绕两个修复时间尺度组织的物理接地 LLM 智能体。在快时间尺度上,结构化物理接口和确定性模拟器将 LLM 输出转化为一个提议,我们在数值上对其检查、修订,并在执行前使其接受有界确定性动作选择。在慢时间尺度上,反复出现的失败模式被整合为持久的上下文原则,这些原则为未来提议提供条件,而物理模型、评估器和执行约束保持不变。在跨多个地点、作物和年份的共同回顾性评估器下,Mimir 在评估的参考方法中取得了所报告的最低总控制成本,并且比历史调度重放少用约 51% 的灌溉。消融研究显示,当移除前向模拟、经验证的修订或持久上下文时,控制成本更高;模型规模和模型家族研究显示,增大 LLM 规模没有单调增益。由此得到的结论表明,持久性物理智能体可以将语义推理与有界、证据驱动的自我改进结合起来,同时将物理真值和执行器权限保留给显式数值机制。
cs.AI / 39 / 2610.02048
HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks
HydroJEV:一种用于配水网络网络攻击与故障归因的一秒级、免训练筛查方法
large language model
大语言模型相关
Abstract
When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve as the first tier of this triage. On a four-class cause-attribution benchmark built on the C-Town network in EPANET, Jev was compared with a hand-written rule tree, a supervised classifier and seven cloud LLMs on identical evidence in four sealed, pre-registered rounds. With only a label-free prior correction, Jev matched the rule tree (macro-F1 0.62-0.64 against 0.56-0.61 in distribution) and exceeded the supervised classifier by 0.36-0.42 on event subtypes absent from its labels, in all four rounds, and it outperformed the classifier whenever fewer than about four labelled events per class were available. Jev also decided 20-40 times faster than frontier LLMs. Accepting only benign Jev verdicts confirmed by the rule tree spared an LLM reviewer 35-38% of windows on fresh sealed sets without loss of macro-F1. Transferred unchanged to two further networks, this gated cascade stayed within the non-inferiority margin of its reviewer on all four sets. A fast, training-free screen can therefore take over about a third of the review load in SCADA anomaly triage while preserving the accuracy of deliberate review.
Chinese Translation
当配水网络中触发 SCADA 报警时,操作人员必须迅速判断该报警反映的是网络攻击、物理故障、正常瞬态还是传感器故障。监督式分类器需要有标注的事件,而水务公司很少具备这类数据;前沿大语言模型(LLM)每次决策需要数十秒。我们检验了 Jev——一个约一秒内返回类别概率的免训练模型——能否充当这种分诊的第一层。在基于 EPANET 中 C-Town 网络构建的四分类原因归因基准上,Jev 在四轮封存、预注册的测试中,与手写规则树、一个监督式分类器以及七个云端 LLM 在相同证据上进行了比较。仅需一种无标签的先验校正,Jev 便与规则树相当(分布内 macro-F1 为 0.62-0.64,对照为 0.56-0.61),并在其标签中未出现的子类事件上以 0.36-0.42 超过监督式分类器,且这在全部四轮中均成立;当每类可用标注事件少于约四个时,它的表现均优于该分类器。Jev 的决策速度也比前沿 LLM 快 20-40 倍。仅接受经规则树确认的良性 Jev 判定,在全新的封存数据集上为一位 LLM 复核者节省了 35-38% 的窗口,且 macro-F1 没有损失。在未作任何改动的情况下迁移到另外两个网络后,这种门控级联在全部四组数据集上均保持在其复核者的非劣效性界限之内。因此,一种快速、免训练的筛查方法可以在 SCADA 异常分诊中接管约三分之一的复核工作量,同时保持审慎复核的准确率。
cs.AI / 40 / 2610.02066
External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
外部观察者可能看得更清楚:通过隐藏状态探测进行大语言模型中的跨模型跨度级幻觉检测
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
Chinese Translation
随着大语言模型(LLM)日益作为基础推理引擎,其产生幻觉的倾向仍然是一个关键脆弱性。尽管近期的内部状态探测为缓慢的外部检索系统提供了一种有前景的替代方案,但它们在很大程度上将幻觉检测简化为逐 token 的二分类任务,未能捕捉语义漂移的结构化、序列化边界。在这里,我们引入一个用于细粒度、跨度级幻觉检测的内部隐藏状态框架。通过检查逐层激活模式,我们尝试在 LLM 生成中检测确切的幻觉起始和延续 token。我们的实验表明,这种方法成功分离出幻觉起始,尽管存在极端类别不平衡,仍在精确率-召回率 AUC 上相较随机基线取得显著提升。最终,我们提出一种新颖的跨模型检测框架,其中一个模型观察另一个模型生成所引发的内部表示。我们发现,外部观察者能够匹敌或超过生成器对自身幻觉起始的自我检测,包括当观察者是更小模型时,这表明自我检测并非起始定位的上限。
cs.AI / 41 / 2610.02070
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
因果记忆策略:通过对检索进行干预使记忆效用可识别
large language model
大语言模型相关
Abstract
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.
Chinese Translation
记忆增强的大语言模型必须决定保留哪些记忆,而近期的系统通过估计每个记忆对任务性能的影响来做到这一点。然而,这些估计完全依赖于被检索到的记忆。当一个记忆从未被检索到时,存储层面的干预会产生完全相同的结果,从而使其效用无法被识别。这是一种检索层面的正性违背(positivity violation),对于仅考察记忆操作的诊断方法而言是不可见的。我们提出因果记忆策略(Causal Memory Policy, CMP),这是一个因果框架,它通过对检索本身进行干预来恢复可识别性,为以已知倾向性采样的记忆保留固定数量的上下文槽位。CMP 在平衡分配设计下通过自归一化逆倾向性加权来估计记忆效用。我们证明了记忆效用经由检索的因果分解、估计量的无偏性与精确方差,以及不可逆操作下的最优决策规则。在实证上,在 LongMemEval 上有 54% 的必需记忆无法被识别,在 LoCoMo 上为 67%,并且这种失效在一个已部署的记忆系统中依然存在。CMP 将必需记忆与非必需记忆之间的区分能力从 0.54 AUC 提升到 0.66 AUC。最后,我们表明仅凭被识别的记忆效用不足以做出保留决策:每个查询的效用在为其估计的那个查询上达到 0.78 AUC,然而保留策略所能使用的任何聚合方式都无法预测某个记忆在未见查询上的价值。代码可在以下地址获取:https://anonymous.4open.science/r/cmp-release-D0C3/。
cs.CL / 42 / 2610.00540
Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models
评估语言差异对大型语言模型多语言语言能力的影响
large language model
大语言模型相关
Abstract
Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.
Chinese Translation
关于多语言语言模型语法能力的论断,会随着能力度量方式的不同而出现显著差异,然而评估范式、后训练与语言资源可得性之间的相互作用尚未得到系统性考察。我们在 MultiBLiMP 上评估了来自六个模型家族的基座模型与后训练模型,该基准是覆盖 101 种语言的句法最小对立体基准,并使用了四种评估方法。我们报告三项主要发现。第一,后训练会损害语法能力,但模型规模对这一效应幅度的削减并不均衡,而低资源语言承担了最高的代价。第二,后训练模型保留了它们无法通过显式提示表达出来的语法知识,然而这一点仅在高资源语言中可被测量,因为在低资源环境下接近随机水平的基线几乎没有可供隐藏的知识。第三,母语提示能够恢复低资源语言上原本被隐藏的能力,这表明只有高资源语言才能直接从无提示的概率中被探测。我们得出结论:多语言语法评估必须采用以语言为依据的、多范式的协议,以避免系统性地低估低资源语言的能力。
cs.CL / 43 / 2610.00562
Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning
大语言模型能否进行长时程推理?面向纵向临床推理的上下文策略实证评估
large language model
大语言模型相关
Abstract
Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.
Chinese Translation
纵向临床推理要求大语言模型(LLMs)识别并整合分布在长期患者病史中的相关证据。尽管长上下文模型能够处理越来越大量的信息,但提供更多病史并不一定使相关证据更易获取或改善推理。我们在 MedLoCoMo 上,跨四个开放权重 LLM 比较五种上下文策略(Full、Recent、Episodic、Semantic 和 Hybrid),考察答案正确性、对查询-证据距离的鲁棒性,以及在具有不受支持前提的问题上的弃答。Episodic 和 Hybrid 通常取得最强的总体准确率,而 Recent Context 在支持性证据变得更远时下降最多;Episodic 和 Hybrid 在长距离上保持最高准确率。对对抗性问题的分析进一步表明,当可用病史不支持所请求的结论时,在可回答问题上的强劲表现并不一定能转化为成功的弃答。这些发现表明,可靠的纵向推理不仅取决于 LLM 能访问多少病史,而且关键取决于相关证据如何被选择并呈现以供推理。
cs.CL / 44 / 2610.00568
Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
涌现的不忠实:对齐训练如何导致语言模型静默地覆盖任务忠实性
large language model
大语言模型相关
Abstract
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
Chinese Translation
大语言模型具有三个关键特性:能力、对齐和忠实性。先前工作研究了能力与对齐之间以及能力与忠实性之间的权衡,但第三种张力仍未得到充分探索:对齐-忠实性冲突。我们表明,对齐后的模型在不安全或敏感内容上会系统性地偏离其输入,却不披露这种修改,我们将这种失败模式称为对齐引发的不忠实(AIU)。与能力驱动的不忠实不同,后者来自知识或推理中的错误,而这是由覆盖对输入遵循的后训练机制所引发。我们引入 FaithConflict,一个隔离这两种冲突的受控数据集,以及两个互补的分类体系:行为分类(B1-B8)和思维链推理分类(C0-C6)。在不同模型中,AIU 随规模增加而增加,且比能力驱动的不忠实增加得更急剧,这是一种反向缩放定律;中间检查点表明它在后训练期间被放大,而 DPO 则是这一差距既增长最多又变得最不可见的阶段。基于提示的缓解并不能解决它,这揭示了大语言模型设计与评估中的能力-对齐-忠实性三难困境。
cs.CL / 45 / 2610.00606
Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities
Waldo 在哪里?跨语言知识差异下的查询语言偏好
large language model
大语言模型相关
Abstract
Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.
Chinese Translation
大语言模型越来越多地充当跨语言知识密集型信息检索任务的接口,通过综合多语言证据来满足这些任务。先前的研究表明,它们常常表现出查询语言偏好——即倾向于偏爱以查询语言撰写的来源——但这类行为大多是在不同语言中可获得等价知识的场景下被考察的。然而,当不同语言的来源对同一事实提供不完整或不一致的叙述时,这种偏差就会产生重要影响,因为此时用户获得的信息取决于模型选择使用哪些来源。为了刻画这种跨语言知识差异下的查询语言偏好,我们提出了 Waldo,一个基于维基百科构建的多语言问答(QA)基准。Waldo 包含 12K 个问答对,针对知识缺口——即某一事实在一种语言中可得但在另一种语言中缺失——以及知识冲突——即不同语言版本对同一事实提供相互冲突的版本。在五种语言上评估八个模型后,我们发现,当某一语言版本仅仅缺少相关事实时,模型通常会使用来自另一种语言的证据,而不论查询使用何种语言。然而,在叙述相互冲突的情况下,模型回答会与查询语言对应的文档高度一致,从而导致语义等价的查询因用户语言不同而引出不同的叙述。最后,我们探索了两种可以在知识冲突下缓解这种偏好的不同方法:一种机制性干预,其消融与查询语言偏好相关的注意力头;以及基于 LoRA 的训练,其将偏好差距最多降低 61.5%。
cs.CL / 46 / 2610.00656
Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference
Lingtai:概念几何关于 LLM 推理揭示了什么——以及没有揭示什么
large language model
大语言模型相关
Abstract
Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept anchors, constructed without labeled concept examples, outcome labels, gradient fitting, or activation-space optimization, producing a structured per-step concept-coordinate signal. Across code generation and grade-school mathematical reasoning, this signal exhibits a robust association with predictive uncertainty: the association survives problem-identity and token-position controls and is not attributable to a single token type, is not explained by a simple correct/incorrect mixture on GSM8K, and is not reproduced by matched random anchors; it is markedly weaker or direction-inconsistent in K-means and PCA projections. Two structures emerge: a recurring uncertainty-linked activity signal whose functional geometry is task-conditioned (distinct activity-entropy shapes on HumanEval, MBPP, and GSM8K), and an execution-specific trajectory identity with strong local inertia but weak re-instantiation invariance--under completion-only elastic alignment, corruption at k=32 (approximately a median quarter of the completion) on the matched re-execution subset still retrieves the archived episode at 62.0%, while a fresh execution retrieves it only 11.7-16.0% of the time. Finally, a matched audit finds no evidence that the scalar concept-activity signal used here supplies a stable correctness coordinate under the tested protocol; we therefore treat correctness as externally supplied. Telemetry adds 0.7-1.6% per-token decode overhead for the 161-anchor code implementation, with unchanged generated tokens.
Chinese Translation
观察大型语言模型在自回归推理过程中计算了什么——在线且无需训练探针——仍然很困难。我们引入 Lingtai,一个无需训练的概念遥测层:在每个生成步骤,残差状态被投影到一个领域特定的命名概念锚点库上,该库的构建无需标注概念示例、结果标签、梯度拟合或激活空间优化,从而产生结构化的逐步概念坐标信号。在代码生成和小学数学推理中,该信号与预测不确定性表现出稳健关联:这种关联在问题身份和 token 位置控制下仍然存在,不能归因于单一 token 类型,不能用 GSM8K 上简单的正确/错误混合来解释,也不能由匹配的随机锚点复现;在 K-means 和 PCA 投影中,它明显更弱或方向不一致。两种结构出现:一种反复出现的不确定性相关活动信号,其功能几何是任务条件化的(在 HumanEval、MBPP 和 GSM8K 上具有不同的活动-熵形状),以及一种执行特定的轨迹身份,具有强局部惯性但弱重新实例化不变性——在仅完成弹性对齐下,在匹配的重新执行子集上,k=32(大约为完成的中位四分之一)处的破坏仍能以 62.0% 检索到存档情节,而全新执行只有 11.7-16.0% 的时间能检索到它。最后,一项匹配审计发现没有证据表明这里使用的标量概念活动信号在测试协议下提供了稳定的正确性坐标;因此我们将正确性视为外部提供的。对于 161 锚点的代码实现,遥测每个 token 解码增加 0.7-1.6% 的开销,而生成的 token 保持不变。
cs.CL / 47 / 2610.00689
Towards Robust Numerical Claim Verification
迈向稳健的数值声明验证
large language model
大语言模型相关
Abstract
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B$\unicode{x2013}$8B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.
Chinese Translation
大语言模型(LLMs)被广泛用于声明验证,但在数值推理方面仍然脆弱:即使数值发生微小变化,也可能使准确率急剧下降。我们表明,这种脆弱性在前沿 LLM 中依然存在,但可以通过在数值扰动样本上进行对抗性微调来缓解。使用参数高效微调,小型 Qwen3 模型(0.6B$\unicode{x2013}$8B)在标签翻转扰动上达到 98.7% 的准确率,优于更大的零样本模型以及前沿系统(GPT-5.4 Pro(74.0%)和 Gemini 2.5 Flash(73.9%))。这些增益可泛化到未见过的扰动类型,表明其具有稳健的数值决策边界,而非记住了这些修改。稳健性也能在不使用目标域数据的情况下迁移,显著提升西班牙语的跨语言性能。我们进一步表明,使用 VitaminC 数据集,相同的微调方案也能赋予模型对证据侧扰动的稳健性。
cs.CL / 48 / 2610.00717
Sequential Functional Structured Tucker Compression for Large Language Model Attentions
面向大语言模型注意力的顺序函数式结构化 Tucker 压缩
large language model
大语言模型相关
Abstract
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.
Chinese Translation
LLM 注意力的训练后压缩通常被表述为独立的矩阵近似,忽略了注意力投影之间的共享结构以及早期压缩引入的表示偏移。我们提出 FTC,一种顺序结构化压缩框架,它在固定存储预算下联合利用原生 Q/K/V 头结构,同时使近似适应当前压缩后的模型。输出投影被单独处理,以考虑变化后的注意力后表示。FTC 既不需要微调,也不需要基于梯度的恢复。在从 6B 到 32B 参数的七个仅解码器 LLM 上,FTC 在五个现代 GQA 模型上的每个测试保留率下都实现了所比较方法中最低的 WikiText-2 困惑度,并且在激进压缩下增益最大。这些改进可迁移到下游任务,并且在 32B 规模上仍然显著。
cs.CL / 49 / 2610.00795
Can large language models unlock discrete data in ophthalmic diagnostic reports?
大型语言模型能否解锁眼科诊断报告中的离散数据?
large language model
大语言模型相关
Abstract
Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer Single Exam, and OCT Thickness Map; n = 5 each) were processed using two GPT-4o-assisted pipelines and compared with a reconciled manual ground truth. Schema-Constrained used Structured Output mode with a predefined JSON Schema; Prompt-Only used a detailed instruction prompt followed by Python conversion to JSON. Outcomes were value accuracy, formatting accuracy, and extraction time. Results: Schema-Constrained value accuracy was 100.00% for Visual Field and RNFL Single Exam, 97.45% for Glaucoma Overview, and 98.00% for Thickness Map; Prompt-Only achieved 100.00% across all four report types. Formatting accuracy was 100.00% for Schema-Constrained across all report types and 100.00% for Prompt-Only except RNFL Single Exam (90.14%). Mean extraction time was 56.51 s per report for manual review versus 5.04 s for Schema-Constrained and 4.70 s for Prompt-Only, an approximately 92% reduction. Conclusions: In this small proof-of-concept dataset, general-purpose LLM-assisted pipelines extracted structured data from ophthalmic diagnostic PDFs with high accuracy and substantially reduced processing time. Prompt-Only achieved the highest value accuracy, while Schema-Constrained produced schema-compliant output with 100% formatting accuracy. These complementary strengths support further evaluation of hybrid, validation-aware workflows for research and clinical data abstraction.
Chinese Translation
目的:评估使用两种提示策略的大型语言模型(LLM)从眼科诊断 PDF 报告中提取结构化数据的准确性和效率。方法:使用两条 GPT-4o 辅助流程处理了四种类型(视野、OCT 青光眼概览、OCT 视网膜神经纤维层单次检查以及 OCT 厚度图;每类 n = 5)的 20 份去标识化报告,并与经核对的人工真值进行比较。Schema-Constrained 使用带有预定义 JSON Schema 的结构化输出模式;Prompt-Only 使用详细的指令提示,随后通过 Python 转换为 JSON。结局指标为值准确率、格式准确率和提取时间。结果:Schema-Constrained 在视野和 RNFL 单次检查中的值准确率为 100.00%,在青光眼概览中为 97.45%,在厚度图中为 98.00%;Prompt-Only 在所有四种报告类型中均达到 100.00%。Schema-Constrained 在所有报告类型中的格式准确率均为 100.00%,而 Prompt-Only 除 RNFL 单次检查(90.14%)外均为 100.00%。人工审核每份报告的平均提取时间为 56.51 秒,而 Schema-Constrained 为 5.04 秒,Prompt-Only 为 4.70 秒,约减少 92%。结论:在这个小型的概念验证数据集中,通用 LLM 辅助流程从眼科诊断 PDF 中提取结构化数据具有高准确性,并大幅缩短了处理时间。Prompt-Only 达到了最高的值准确率,而 Schema-Constrained 产生了符合 schema 的输出,格式准确率为 100%。这些互补优势支持进一步评估用于研究和临床数据抽取的混合式、验证感知型工作流。
cs.CL / 50 / 2610.00827
Verbalized and Internal Probabilities Are Coupled in Large Language Models
大型语言模型中的言语化概率与内部概率相互耦合
large language model
大语言模型相关
Abstract
Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that internal probabilities track relative frequencies in the training data, and that verbalized probabilities track explicit probabilistic assertions in the training data. However, we do not know whether these two readouts are aligned, except when frequencies and probabilistic assertions in the training data happen to align. This limits our understanding of when we can use verbalized uncertainties as a proxy for either training data frequencies, or a model's internal distribution. We resolve this gap by systematically exploring how LLMs probability readouts are impacted by training and in-context data, via intervening on the underlying uncertainty sources in the data. We find that both internal and verbalized probability readouts are impacted by both distributional and asserted uncertainty in the training data. Further, we find that verbalized and internal probabilities are aligned beyond what would be expected by independently tracking the same uncertainty sources, suggesting that verbalized probabilities can be used to probe a model's internal distribution.
Chinese Translation
大型语言模型在其采样分布中携带一种内部的不确定性概念,即它们赋予生成某个答案而非另一个答案的概率。它们也可以被要求以文字或数字陈述置信度:一种言语化的不确定性。先前的工作表明,内部概率追踪训练数据中的相对频率,而言语化概率追踪训练数据中显式的概率断言。然而,我们不知道这两种读出是否对齐,除非训练数据中的频率与概率断言恰好对齐。这限制了我们对何时能够将言语化不确定性用作训练数据频率或模型内部分布的代理的理解。我们通过系统性地探索 LLMs 的概率读出如何受到训练数据和上下文数据的影响来解决这一空白,方法是对数据中潜在的不确定性来源进行干预。我们发现,内部概率读出和言语化概率读出都受到训练数据中分布不确定性和断言不确定性的影响。此外,我们发现,言语化概率与内部概率的对齐程度超出了独立追踪相同不确定性来源所预期的程度,这表明言语化概率可用于探测模型的内部分布。
cs.CL / 51 / 2610.00840
Contextual trajectory and incremental contextual displacement: Towards using LLMs to understand dynamic, utterance-specific meaning construction
上下文轨迹与增量上下文位移:迈向使用 LLM 理解动态的、特定于话语的意义构建
large language model
大语言模型相关
Abstract
Transformer-based large language models (LLMs) such as RoBERTa represent text using contextual word embeddings (CWEs), which alter the embeddings associated with each token based on surrounding context. We construct token-wise incremental trajectories by repeatedly recomputing a token's CWE as successive words are added to a sentence, yielding a representation of how contextualized embeddings evolve as the utterance unfolds. We evaluate this approach using garden-path sentences as a test case with characteristic features. Token-wise trajectories reproduce known features of garden-path processing, including disruption around the critical region, and reliably distinguish garden-path sentences from matched disambiguated controls. We introduce several metrics for quantifying representational displacement across contextual increments and show that trajectory information can be highly predictive of sentence type. We find that ambiguity-related information is recoverable not only from the sentence-level CLS representation but also from ordinary vocabulary tokens, suggesting that utterance-level information is distributed across multiple representational scales. In exploratory analyses, we find qualitatively similar trajectory structures in other ambiguity- and misdirection-related linguistic phenomena. Together, these results establish token-wise incremental trajectories as a promising framework for studying utterance-specific meaning construction using LLMs.
Chinese Translation
基于 Transformer 的大型语言模型(LLM),如 RoBERTa,使用上下文词嵌入(CWE)来表示文本,这些嵌入会根据周围的上下文改变与每个 token 相关联的嵌入。我们通过随着单词相继加入句子而反复重新计算某个 token 的 CWE,构建逐 token 的增量轨迹,从而得到一种表示,刻画上下文嵌入如何随着话语展开而演变。我们使用具有典型特征的花园路径句作为测试案例来评估这一方法。逐 token 轨迹再现了花园路径加工的已知特征,包括关键区域附近的扰动,并能可靠地区分花园路径句与匹配的去歧义控制句。我们引入若干度量,用于量化跨上下文增量的表征位移,并表明轨迹信息可以对句子类型具有高度预测性。我们发现,与歧义相关的信息不仅可以从句子级 CLS 表示中恢复,也可以从普通词汇 token 中恢复,这表明话语级信息分布在多个表征尺度上。在探索性分析中,我们在其他与歧义和误导相关的语言现象中发现了定性上相似的轨迹结构。总之,这些结果确立了逐 token 增量轨迹作为一个有前景的框架,用于使用 LLM 研究特定于话语的意义构建。
cs.CL / 52 / 2610.00928
Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations
大语言模型中的高效任务适配:基于权重、基于提示与基于嵌入的适配方法综述
large language model
大语言模型相关
Abstract
As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-injection approaches. However, these lines of work have largely evolved within individual paradigms, leaving their cross-paradigm relationships and trade-offs underexplored, especially for recently emerging embedding-based adaptations. This survey presents a unified framework that categorizes task adaptation methods by where and how task information is encoded: model weights, input prompts, or injected task embeddings. We provide a comprehensive taxonomy that integrates these paradigms, analyze their key strengths and limitations to explain how different adaptation paradigms have evolved, clarify relationships across paradigms, and highlight open problems for future research.
Chinese Translation
随着大语言模型日益被部署到各式各样的下游任务中,高效的任务适配已成为一个核心挑战。为此,人们提出了大量任务适配方法,涵盖参数高效微调、上下文学习以及嵌入注入等方法。然而,这些研究方向大多是在各自的范式内部发展演变的,其跨范式的关系与权衡仍未得到充分探索,尤其是对于近来新兴的基于嵌入的适配方法而言。本综述提出了一个统一的框架,依据任务信息被编码的位置与方式——模型权重、输入提示或注入的任务嵌入——对任务适配方法进行分类。我们提供了一个整合这些范式的全面分类体系,分析它们的主要优势与局限,以解释不同的适配范式是如何演变的,厘清各范式之间的关系,并指出未来研究的开放性难题。
cs.CL / 53 / 2610.00940
ReHoPER: Receding-Horizon Planning for Enhanced Reasoning
ReHoPER:用于增强推理的滚动时域规划
large language model
大语言模型相关
Abstract
We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.
Chinese Translation
我们提出 ReHoPER,一种仅用于推理的、零样本方法,它通过在最终答案之前沿多条路径生成并回答中间问题来提升大型语言模型的推理能力。它迭代地规划一个候选中间问题的时域,选择一个来回答,并从更新后的历史重新规划。ReHoPER 是任务无关的,在数据集和模型之间使用相同的通用指令,无需标注数据或任务特定的提示设计。在多个数据集上,包括 iLLC(一个用于组合推理的新受控基准),ReHoPER 优于强基线,在最具组合性的设置中提升最大。我们的实现和 iLLC 生成器已公开可用,以支持未来工作。
cs.CL / 54 / 2610.00954
Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles
超越排行榜:智能体小语言模型集成的代币经济学
large language model
大语言模型相关
Abstract
As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.
Chinese Translation
随着大语言模型(LLMs)从独立助手转向智能体工作流,评估必须超越标量式的排行榜准确率,进而考量运行可靠性、成本、延迟和 token 效率。我们使用一个由小语言模型(SLMs)构成的智能体集成,并配备由 SLM 评审(SLM-judge)中介的反馈回路,作为此类超越排行榜评估的案例研究。在包含 541 个提示的 IFEval 基准上,最佳集成达到了 97.34% 的严格提示准确率,比最强的独立 LLM 基线 gpt-5.4 高出 5.81 个百分点,同时运行在成本更低的区间。随后我们分析这一增益背后的代币经济学与运行行为,包括每样本成本、token 构成、有用输出有效吞吐(goodput)、反馈回路恢复、延迟分解,以及跨指令类别和约束数量的性能。我们的结果表明,智能体 SLM 集成能够以额外的测试时 token 与编排开销,换取更高的指令遵循保真度,这为未来智能体 AI 系统的多维评估协议提供了动因。
cs.CL / 55 / 2610.00958
Role-aware Heuristic Episodic Attention for Conversational LLMs
面向对话式大语言模型的角色感知启发式情景注意力
large language model
大语言模型相关
Abstract
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91$\times$. Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.
Chinese Translation
大型语言模型在多轮对话不断增长时,往往会丢失对持续指令和相关信息的追踪。我们通过三种相关的失败模式研究这种累积式上下文衰减:注意力污染、注意力稀释和注意力漂移。我们提出 REA(Role-aware Heuristic Episodic Attention,角色感知启发式情景注意力),一种上下文管理框架,它为指令和情景交互分配不同的持久性与表示策略。指令记忆将识别出的全局约束保留在一个专用前缀中。情景记忆保存用户输入并压缩模型回复,而启发式检索会为每个历史轮次选择原始文本、压缩表示或省略。在 Long-MT-Bench+ 上,REA 在 10 分制下将评判分数从 6.32 提高到 7.36,相对于 Vanilla 基线取得 16.5% 的相对提升,并将平均延迟降低 2.91$\times$。额外评估显示,在参数规模从 1.7B 到 7B 的三个骨干模型上,以及在中文和英文角色扮演任务上,均取得了总体增益。这些结果支持将角色感知的上下文管理作为一种维持对话连续性和指令遵循的实用方法。
cs.CL / 56 / 2610.00969
A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models
一个面向可信财报电话会议记录分析、以引文为支撑的大语言模型基准
large language model
大语言模型相关
Abstract
Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generating standalone claims, users increasingly prefer grounded analyses that pair claims with verifiable citations from source documents to enable independent validation. However, evaluating such analytical claims typically requires extensive expert annotation, which is costly and difficult to scale, and real-world financial analysis commonly involves long context-question-answer triplets, further increasing task complexity. To address these challenges and benchmark the current landscape of grounded analysis by LLMs, we propose a numeric evidence evaluation method that enables groundedness assessment without reliance on expert annotation. We also introduce an automated dataset construction pipeline and construct ECTs-100 from the top 100 constituents of the S&P 500 to support benchmark of both groundedness and correctness. In addition, we examine conscious incompetence, a practical failure mode in financial analysis in which LLMs must detect when available evidence is insufficient and refrain from producing unsupported hallucinations. Empirical results show that LLMs perform well in groundedness but face notable limitations in correctness, with informational insufficiency presenting an additional challenge.
Chinese Translation
大语言模型(LLMs)已越来越多地被用于金融文档分析,包括财报电话会议记录(ECTs)。除了生成独立的论断之外,用户越来越青睐有依据的分析,即把论断与来自源文档的可验证引文配对,以便能够进行独立验证。然而,评估此类分析性论断通常需要大量的专家标注,这既昂贵又难以规模化,并且现实世界的金融分析通常涉及长上下文—问题—答案三元组,进一步增加了任务复杂度。为应对这些挑战并对LLMs有依据分析的当前状况进行基准评测,我们提出了一种数值证据评估方法,该方法能够在无需依赖专家标注的情况下实现有据性评估。我们还引入了一个自动化数据集构建流水线,并从标普500指数前100只成分股构建了ECTs-100,以支持对有据性和正确性两者的基准评测。此外,我们考察了“有意识的无能”(conscious incompetence),这是金融分析中的一种实际失效模式,其中LLMs必须检测出可用证据何时不足,并避免产生没有支撑的幻觉。实证结果表明,LLMs在有据性方面表现良好,但在正确性方面面临显著局限,而信息不足则构成了额外的挑战。
cs.CL / 57 / 2610.00997
Distilling Directional Verification
蒸馏方向性验证
large language model
大语言模型相关
Abstract
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at https://github.com/js-lee-AI/directional-verification.
Chinese Translation
知识蒸馏旨在将大语言模型的事实性知识迁移到更小的模型,以实现高效部署。然而,教师模型可能在一个方向上回忆起某个关系,却无法在相反方向上生成答案。因此,从其生成的答案中进行蒸馏会把这一方向性局限传播给学生模型。尽管如此,同一个教师模型能够通过在其所知的方向上对该关系进行评分来识别这样的答案。我们引入方向性标签蒸馏,其中冻结的教师模型在那个已知方向上对候选答案进行评分,得分最高的候选答案成为学生的训练目标。在关于父母及其子女的事实上,已知方向评分比按所请求方向评分产生更准确的标签,即使在对名字先验进行调优校正之后也是如此。在使用先验校正后的分数时,更好的方向取决于事实本身而非模板,并且在所挖掘的事实上会发生反转——在这些事实中,显著实体是父母而非子女。在隐去所评估的儿童正向事实的情况下,用已知方向标签训练的学生在其训练查询上的开放式准确率比用先验校正的反向标签训练的学生高出13到15个百分点。在将生成的答案通过词法相似度匹配到固定名字列表之后,学生几乎再现了所有被选中的标签。它们的准确率在很大程度上随标签质量而变化。这种标签优势在未筛选的查询上依然成立,在检索候选答案而不插入正确答案时也依然成立。我们的发现表明,方向性验证通过提供更准确的训练目标,缓解了错误从教师生成的答案向学生的迁移。代码可在 https://github.com/js-lee-AI/directional-verification 获取。
cs.CL / 58 / 2610.01016
Scaling and Distilling Text Embeddings for Better Diffusibility
为更好的可扩散性而缩放与蒸馏文本嵌入
diffusion
扩散模型相关
Abstract
Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Chinese Translation
扩散语言模型(DLMs)为自回归(AR)语言生成提供了一种有前景的替代方案。连续 DLMs 的最新进展将潜在扩散应用于连续文本嵌入,这提出了一个实际问题:哪种嵌入能构成最佳潜在空间,即最具可扩散性的潜在空间?为回答这个问题,我们搜索了不同的嵌入,并发现将同一系列内的嵌入模型扩展到更强的模型(T5 到 T5Gemma-1 到 T5Gemma-2)会大幅提升生成性能。但原始 T5Gemma-2 嵌入仍非最优。它们判别性太强,以至于连合理替代词的嵌入也被分开,这使得生成易受不完美采样影响。因此,连续扩散往往无法到达其中任何一个,反而最终落在无效嵌入上。为解决此问题,我们将 T5Gemma-2 蒸馏到一个学生编码器中,该编码器将教师的解码概率作为软标签来学习。从这种软标签中学习使学生将替代嵌入拉得更近,同时保持编码-解码机制。蒸馏后的嵌入形成更连通、更可扩散的潜在空间,优于普通 T5Gemma-2 嵌入。结果,我们的中等规模 DLM 在 OpenWebText 上于真实文本熵处达到 Gen. PPL 17.8(相对于真实文本 PPL 15.4),在 Gen. PPL 上优于 GPT-2-M。
cs.CL / 59 / 2610.01026
It Takes Workflows to Evolve Better Workflows
要演化出更好的工作流,需要工作流
large language model
大语言模型相关
Abstract
Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to $+7.41\%$, with co-evolving ($+5.03\%$) more roles gaining more than optimizing one of them alone ($+2.83\%$). Our project page: https://xhguo7.github.io/FloWright/.
Chinese Translation
处理复杂的现实世界任务可能超出单个大语言模型(LLM)的能力,这促使人们使用多智能体工作流来协调专门化的智能体共同处理这些任务。最近的方法训练 LLM 从执行结果中构建更好的工作流,但它们只优化工作流生成器,而构建或执行每个工作流的其他智能体保持固定,尽管每个结果都依赖于所有这些智能体。然而,将训练扩展到生成器之外具有挑战性:这些智能体是耦合的,而工作流的结果是一个单一的稀疏分数,无法说明哪个智能体导致了失败。我们提出 FloWright,它利用工作流作为优化工作流的载体。通过引入一种分层、结构感知的奖励范式,FloWright 使一个角色能够自我演化,并使两个或更多角色能够共同演化,而无需额外的模型、标签或执行。考虑到工作流通常是在单个智能体已经能够处理的数据上进行训练和评估这一局限,我们进一步提出 DataWright,一种自适应数据强化方法,它将现有数据集转换为难度增加的工作流级任务。在文档、幻灯片、图表、代码、数学和金融任务中,使用 FloWright 训练的小型开放模型性能提升最高达 $+7.41\%$,其中共同演化($+5.03\%$)更多角色比单独优化其中一个角色($+2.83\%$)获得更多收益。我们的项目页面:https://xhguo7.github.io/FloWright/。
cs.CL / 60 / 2610.01027
LawCompass: Navigating from Legal QA to Multi-Agent Deep Research with Grounded Evidence
LawCompass:从法律问答迈向基于扎实证据的多智能体深度研究
large language model
大语言模型相关
Abstract
Recent advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have significantly democratized access to legal information. Nevertheless, most existing legal assistants remain confined to multi-turn conversational QA, failing to support complex legal tasks that require systematic evidence retrieval, multi-step reasoning, and report-level synthesis. In this paper, we present LawCompass, an evidence-grounded legal assistant that navigates the transition from standard Legal QA to multi-agent deep research. LawCompass provides three task-oriented functions: Legal QA, which delivers precise, evidence-backed answers to legal questions; Professional Retrieval, which enables structured exploration of statutes and judicial cases via query rewriting; and Deep Research, which employs a multi-agent workflow to decompose complex legal tasks and synthesize comprehensive research reports. Crucially, LawCompass maintains explicit citation links across all modules, empowering users to directly verify system outputs against original legal sources. Evaluation results demonstrate that LawCompass provides a practical and scalable paradigm for transforming conversational AI into trustworthy and evidence-grounded legal research assistance.
Chinese Translation
近年来,大型语言模型(LLMs)和检索增强生成(RAG)的进展显著推动了法律信息获取的民主化。然而,大多数现有法律助手仍局限于多轮对话式问答,无法支持需要系统性证据检索、多步推理和报告级综合的复杂法律任务。在本文中,我们提出 LawCompass,一个以证据为基础的法律助手,它引导从标准法律问答到多智能体深度研究的转变。LawCompass 提供三种面向任务的功能:法律问答,能够针对法律问题给出精确且以证据为支撑的答案;专业检索,能够通过查询重写对法规和司法案例进行结构化探索;以及深度研究,采用多智能体工作流来分解复杂法律任务并综合生成全面的研究报告。至关重要的是,LawCompass 在所有模块中维护显式的引用链接,使用户能够直接对照原始法律来源核验系统输出。评估结果表明,LawCompass 提供了一种实用且可扩展的范式,可将对话式 AI 转变为可信且基于证据的法律研究辅助。
cs.CL / 61 / 2610.01066
Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
用参与奖杯进行探测:随机奖励强化学习作为大语言模型能力的探针
large language model
大语言模型相关
Abstract
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.
Chinese Translation
我们将虚假奖励悖论与模型的可达性联系起来,并提出随机奖励强化学习(RL)作为探测工作的一种有用工具,回应了长达十年的关于探测性能究竟揭示了模型的什么这一争论。对于即使随机奖励也能提升大语言模型(LLM)性能这一令人惊讶的发现,有两种流行的解释:一种将增益归因于 RL 训练中的特定机制;另一种归因于数据污染。我们的结果支持一种不同的观点:虚假奖励 RL 可以探测模型的可达性,即在指定约束下,进一步训练能从其当前状态达到什么,超出其当前性能所反映的内容。例如,两个在合成算术上具有相同准确率(3.5%)的 OLMo 检查点,在相同的正确性奖励 RL 下,其最佳运行分别达到 8.5% 和 55%。考察预训练和中期训练阶段的 OLMo 检查点,揭示了三种不同的训练响应模式:在早期,即便对正确答案给予奖励,RL 也几乎不产生改进;在预训练后期,奖励正确答案变得有效,而随机奖励仍然较弱;并且,在进入中期训练后,即便随机奖励也能产生大幅增益。对这些检查点进行数字掩码监督微调(SFT)分析时,也出现了类似的排序,表明这一模式并非特定于某种 RL 机制。此外,采用随机奖励的 RL 为训练能让 LLM 做什么提供了一个独特视角,因为其奖励信号不提供任何关于哪些答案正确的信息。通过追问在没有正确性反馈的情况下训练能够达到什么,它处理了基于可解码性的探测中一个核心问题的标签泄漏方面:一个成功的探针究竟揭示了模型的能力,还是学习了任务本身。
cs.CL / 62 / 2610.01127
Counting and Min-Cost Encoding for Tokenization in Large Language Models
大型语言模型分词中的计数与最小代价编码
large language model
大语言模型相关
Abstract
Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.
Chinese Translation
主流大型语言模型依赖分词器将文本编码为词元序列。不同分词器对同一文本可能产生长度显著不同的词元序列。在固定模型架构下,更短的词元序列对应于更低的推理时间。我们提出一种名为计数与过滤(Counting and Filtering, CNF)的分词器训练方法,以及一种称为最小代价编码(Min-Cost Encoding, MCE)的文本编码算法。MCE 在一个文本片段上定义代价函数,并通过全局最小化整体切分代价来确定最佳切分。CNF 通过直接计数有效子串来构建原始词表,然后在使用 MCE 切分训练语料时,基于实际词元使用情况通过过滤步骤构建最终词表。CNF-MCE 组合相较 BPE 提供若干优势,包括更高的词元效率、更强的可扩展性以及更低的依赖性。在六个文本类别和两个词表规模组上,CNF-MCE 始终比所评估的 BPE 分词器实现更好的压缩。在 250K 词表下,CNF-MCE 在英文网页文本上将压缩率相较于 o200k_base 和 qwen250k 分词器分别提高 26% 和 30%。在英文网页文本上将词表扩展至 1M 条目的实验表明,相较于 BPE 有持续改进,词元效率提升超过 60%,词表利用率从 52.9% 上升至 96.9%。MCE 算法不依赖于合并列表(如 BPE 中那样)或词元概率(如 UnigramLM 中那样),使其适用于广泛的词表,包括由 BPE、UnigramLM、CNF 及其他方法构建的词表。在 1.8B 和 8B 规模上从头训练的语言模型在 11 个基准上达到与使用 BPE 分词器的模型相当的平均性能。这些结果表明,CNF-MCE 能够显著提高词元效率,同时保持有竞争力的下游性能。
cs.CL / 63 / 2610.01161
My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning
My FAULT:自诊断作为自演化智能体强化学习中的信用分配
large language model
大语言模型相关
Abstract
Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.
Chinese Translation
智能体强化学习(RL)已成为训练大语言模型智能体完成多步任务的一种强大方法,但依赖终端结果奖励会产生两个信用分配问题,尤其是在长时程任务中。首先,同结果 rollout 组无法从终端奖励中获得学习信号。其次,终端奖励仅提供轨迹范围的反馈,因此难以识别哪些决策导致了失败。近期工作通过轨迹分析得到的更细粒度信息来补充终端奖励,例如对中间决策和错误的自然语言反思。然而,自然语言诊断难以直接用于信用分配:其错误断言可能不可靠,并且它们没有量化每个错误应对学习产生多大影响。我们提出自诊断引导的终端信用再分配(FAULT),它将被诊断出的错误转化为由终端结果锚定的显式步骤级信用。FAULT 检查诊断证据,并从任务结果中学习相对错误代价。在训练过程中,策略与自诊断器共同演化,同时错误代价根据近期结果在线更新。在 ALFWorld 上,FAULT 从同结果组中恢复学习信号,达到 95% 的信号覆盖率,而 GRPO 为 41%,GiGPO 为 72%,同时更好地将信用定位到具体错误步骤。在两个模型规模上,FAULT 在长时程 ALFWorld 和 WebShop 任务上带来了强劲的提升,同时在短时程基于搜索的问答任务上保持竞争力。
cs.CL / 64 / 2610.01177
Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
时间分辨的 Token 归因揭示扩散语言模型的生成动力学
diffusion
扩散模型相关
Abstract
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
Chinese Translation
本文提出了扩散层积分梯度(DLIG),一种面向扩散语言模型(DLM)的 token 归因方法,它将积分梯度(IG~\cite{sundararajan2017axiomatic})扩展到任意层和去噪步骤。DLIG 对 DLM 向针对输入提示的自生成或固定补全所作出的渐进承诺进行归因。我们建立了 DLIG 与 IG 的完备性、实现不变性、线性和对称性保持公理之间的直接对应关系。作为对干预分析的轻量级补充,DLIG 为跨去噪轨迹的机制假设提供了一种低成本的初步检验。我们在词义消歧、多跳图推理和句子填充上展示了这一点,揭示了 DLM 如何跨位置、层和去噪步骤利用输入。
cs.CL / 65 / 2610.01241
Evaluating the Robustness of Japanese LLMs to IME-Related and Typographical Errors
评估日语大语言模型对IME相关错误与打字错误的鲁棒性
large language model
大语言模型相关
Abstract
Large language models (LLMs) have achieved strong performance across various natural language processing tasks. However, their robustness to typographical errors remains underexplored, particularly in Japanese, where text input involves multiple writing systems and IME-based conversion. In this study, we evaluate the robustness of Japanese LLMs against realistic Japanese-specific typos. We introduce five typo categories: Character Transposition, Character Replacement, Homophone Conversion, Japanese IME Conversion, and Full-Width Conversion. These perturbations are applied to three Japanese benchmark datasets (JMMLU, JCommonsenseQA, and JamC-QA), and eleven Japanese and multilingual LLMs are evaluated. The results show that Character Transposition and Character Replacement typos consistently reduce accuracy across benchmarks, whereas IME Conversion, Full-Width Conversion, and Homophone Conversion have relatively limited impact. These findings reveal that current Japanese LLMs remain vulnerable to realistic Japanese typing errors, particularly those that substantially distort the original input, highlighting the importance of robustness evaluation in practical input environments.
Chinese Translation
大语言模型(LLMs)已在各种自然语言处理任务中取得了强劲性能。然而,它们对打字错误的鲁棒性仍未得到充分探索,尤其是在日语中,文本输入涉及多种书写系统和基于IME的转换。在本研究中,我们评估日语大语言模型对真实的日语特有打字错误的鲁棒性。我们引入五类打字错误:字符换位(Character Transposition)、字符替换(Character Replacement)、同音词转换(Homophone Conversion)、日语IME转换(Japanese IME Conversion)和全角转换(Full-Width Conversion)。这些扰动被应用于三个日语基准数据集(JMMLU、JCommonsenseQA和JamC-QA),并评估了十一个日语和多语言大语言模型。结果表明,字符换位和字符替换错误在各基准上一致地降低准确率,而IME转换、全角转换和同音词转换的影响相对有限。这些发现揭示,当前的日语大语言模型仍然容易受到真实的日语打字错误影响,尤其是那些会显著扭曲原始输入的错误,这凸显了在实际输入环境中进行鲁棒性评估的重要性。
cs.CL / 66 / 2610.01275
Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models
知道何时保留它们:均匀状态扩散语言模型中的正确词元保留
diffusion
扩散模型相关
Abstract
Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173--270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3--11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.
Chinese Translation
均匀状态扩散模型(USDMs)能够在任意去噪步骤修正任意词元,这使它们能够纠正自身错误,这是相对于掩码扩散的一个关键优势。然而,自纠正既需要修正错误词元,也需要保留正确词元,而我们表明当前的USDMs缺乏后者。即使在贪心尾部解码下,最先进的USDMs(DUO、UDLM和均匀噪声SEDD)在每一步仍会修正512个位置中的173--270个,而这些大规模、不协调的编辑会使样本多样性崩溃。一项随机词元破坏实验将这一缺陷追溯到模型自身:它们以几乎相同的准确率重建干净词元和受损词元,尽管干净词元是更容易的目标。对验证NELBO的分解表明,训练几乎不奖励保留:错误预测在受损位置受到严重惩罚,但在干净位置几乎无代价。我们提出正确词元保留正则化(CTR-Reg),这是一种简单但有效的辅助损失,训练模型保留前向过程未扰动的词元,并且不需要对采样器做任何改变。CTR-Reg在六个基准上平均将干净词元准确率提高26.5个百分点,同时使受损词元准确率几乎不变,并且其每步修正数收敛到仅3--11个位置。仅需五个贪心尾部步骤,在CTR-Reg下,所有三个模型的生成困惑度都减半以上,同时多样性得以保留,而且这些收益在不同采样预算下均成立。我们的结果将正确词元保留确定为自纠正扩散语言模型缺失的关键要素,并证明了一种有效的修复方法。
cs.CL / 67 / 2610.01353
Does AI-Generated Scientific Text Follow Human Argumentation Patterns? A CARS-Based Comparison of Research Article Introductions
人工智能生成的科学文本是否遵循人类论证模式?——基于 CARS 的研究论文引言比较
large language model
大语言模型相关
Abstract
Large language models are moving from helping write up research to helping do it, which makes it important to know how the scientific text they produce differs from human writing. Work on this question has stayed mostly at the surface, using lexical and stylistic cues that light paraphrasing erases. We look instead at rhetorical structure, the sequence of argumentative moves through which a text makes its case. We study research-article introductions under Swales' CARS model, and compare original introductions from published linguistics articles with generated counterparts of the same papers. We find that human-written introductions are more flexible in which moves they use and in what order, while the generated ones are more uniform. Giving the models the CARS definitions makes them more rigid.
Chinese Translation
大语言模型正从帮助撰写研究论文转向帮助开展研究,这使得了解它们生成的科学文本与人类写作有何差异变得十分重要。关于这一问题的研究大多停留在表面,使用的是轻度改写就会抹去的词汇和文体线索。我们转而考察修辞结构,即一个文本借以提出其论点的论证语步序列。我们在 Swales 的 CARS 模型下研究研究论文引言,并将已发表语言学论文的原始引言与同一批论文的生成对应引言进行比较。我们发现,人类撰写的引言在使用哪些语步以及以何种顺序使用方面更为灵活,而生成的引言则更为统一。向模型提供 CARS 定义会使它们更加僵化。
cs.CL / 68 / 2610.01428
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
泛化是稳定性,而非准确性:大语言模型的多轴评估
large language model
大语言模型相关
Abstract
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
Chinese Translation
大语言模型(LLMs)中的泛化是指当同一输入以不同方式表达时,模型能够产生一致且语义稳定的输出的能力。现有工作通常通过单一提示格式、任务或一组变体上的聚合准确率来评估泛化,这会将鲁棒性与整体基准性能混为一谈。在这项工作中,我们展示在个体样本层面、跨多个输入变体以及跨模型行为不同方面的泛化评估,关注变异性,而不是将性能简化为一个可以通过狭窄训练或其他混淆泛化评估的方式提高的分数。遵循这一观点,我们引入稳定性感知泛化目标(Stability-Aware Generalization Objective,SAGO),这是一个框架,用于衡量同一输入在不同变体和基准下模型行为的变化程度,捕捉包括生成一致性、内部激活、置信度和响应镜像在内的多个维度上的变异性。我们表明,许多常用模型表现出统计显著且一致的泛化不稳定性:没有模型能够一致地泛化,行为轴捕捉到独立的失败模式,而跨数据集的变化可以逆转模型排名。
cs.CL / 69 / 2610.01471
When Does a Second Model Help? Cross-Model Review in LLM Verification
第二个模型何时有帮助?LLM 验证中的跨模型评审
large language model
大语言模型相关
Abstract
Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CCR plus one cross-model review matches more planted errors than two CCR reviews (56.7% vs. 42.7%; Holm-adjusted p=.006), but not significantly more than two reviews by the top-tier cross-model reviewer, so model difference and reviewer capability are not separated. A lightweight cross-model reviewer scores no higher than same-model review. Withholding requirements from the reviewer raises F1 for the two lower tiers but not the top tier, in untested point estimates whose pattern depends on how failed sessions are scored. Before analysis we audited all session records, excluding one baseline run of uncertain provenance and 14 failed calls; results with all sessions are also reported. A partial check on public detector outputs from another benchmark neither replicates nor contradicts the main comparison. Records, artifacts, and scripts are available from the author on request.
Chinese Translation
大语言模型现在能够生成代码、文档和分析,并且越来越多地被用于评审此类输出。我们研究由不同模型进行的第二次评审在何时有帮助。在作者早先预印本的基础上——这些预印本在一个模型内改变了上下文、重复和角色结构——我们在一个受控实验中检验模型独立性:30 个工件,其中植入 150 个错误,10 种评审条件,以及涉及来自两个开发者的三个评审模型的 900 次评审会话。在该实验中,(1) 顶级跨模型评审者在 F1 上与全新会话中的同模型评审(CCR)没有显著差异,但这并未确立等价性;(2) 二者发现部分不同的错误(Jaccard 41.2%);以及 (3) 在两次评审调用中,一次 CCR 加一次跨模型评审比两次 CCR 评审匹配到更多植入错误(56.7% 对 42.7%;Holm 校正 p=.006),但不显著多于由顶级跨模型评审者进行的两次评审,因此模型差异与评审者能力没有被分离开。一个轻量级跨模型评审者的得分不高于同模型评审。向评审者隐瞒需求会提高两个较低层级的 F1,但不会提高最高层级,这些是尚未检验的点估计,其模式取决于失败会话如何评分。在分析之前,我们审计了所有会话记录,排除了一个来源不确定的基线运行和 14 次失败调用;也报告了包含所有会话的结果。对来自另一个基准的公开检测器输出所做的部分检查,既没有复现也没有与主要比较相矛盾。记录、工件和脚本可应要求向作者索取。
cs.CL / 70 / 2610.01491
Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories
对 WebArena-Lite 上网页智能体评估的审计:对结果与轨迹的人工审查
large language model
大语言模型相关
Abstract
Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.
Chinese Translation
网页智能体是大语言模型的一项重要应用,然而其评估往往依赖于仅检查最终结果的基于规则的评估器或语言模型评估器。对任务完成情况的人工验证以及对失败轨迹的详细分析仍然十分有限。我们在由 GPT 5.5 和未经训练的 Qwen3.5 9B 模型构建的六种评估条件下,对全部 165 个 WebArena Lite 任务进行了审计。该审计保留了原始得分,纠正了自动评估器产生的假阴性,识别出第一个导致后果的错误,并考察了整条轨迹上的进展。我们还研究了一种记忆与分析支持机制(Memory and Analysis Support Mechanism,MASM),它维护显式的执行状态,以及引导文本(Guide Text),它提供与任务相关的流程性指导。在四种 GPT 5.5 设置下,人工审查找回了评估器所漏掉的 5.45 至 8.49 个百分点的成功率。在 25 步的预算下,引导文本使使用 MASM 时的校正成功率从 34.55% 提升至 38.18%。在未经训练的 Qwen3.5 9B 模型上,MASM 将评估器得分从 13.90% 提升至 18.80%。对 102 条失败的 GPT 5.5 轨迹的审查揭示了频繁的滚动循环、未完成的探索、过早给出的答案、无效动作以及不完整的表单工作流程。步骤级别的证据进一步表明,显著的早期进展可能与最终失败并存。这些结果表明,为什么仅凭最终得分只能对网页智能体行为提供不完整的说明,并促使我们采用以人工为依据、具备轨迹感知能力的验证方式。
cs.CL / 71 / 2610.01514
How the Audit Rule Shapes Faithful Factor Explanations in LLMs
审计规则如何塑造大语言模型中忠实的因子解释
large language model
大语言模型相关
Abstract
Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.
Chinese Translation
大语言模型经常被问及哪些输入因子影响了它们的输出。对于结构化输入,此类报告可以通过反事实扰动来检验,但每个因子都必须被多次查询以估计其效应,因此验证通常受到预算限制。我们研究这种有限预算的设置如何改变如实报告因子层面影响力的激励。我们将这种交互形式化为一个验证博弈,并表明当审计依赖于报告时,仅有恰当评分是不够的:依赖报告的审计会产生一种压制激励,因为被报告为重要的因子更有可能被检查,并因估计噪声而受到惩罚。相比之下,与报告无关的审计,或带有较小的与报告无关下限的混合规则,会消除这一渠道,并使如实报告比完全压制更可取。我们用反事实 Brier 分数(CBS)实例化该框架,并在四个 NLP 基准上评估其预测。一个合成的理性智能体与理论预测完全吻合,而真实的大语言模型在激励被明确呈现时也遵循相同的激励。主要的设计启示很简单:在部分验证下,因子层面的解释系统应包含一个与报告无关的审计组件,这样低报就不能被用来逃避审查。
cs.CL / 72 / 2610.01592
Which LLM to pick? Online Active Model Selection for Large Language Models
该选择哪个 LLM?面向大语言模型的在线主动模型选择
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online settings. Given an arbitrary stream of queries and a limited annotation budget, ONLINE LLM PICKER selects the most informative prompts for annotation to identify the best LLM among candidate models. Across multiple tasks including 10 datasets, for over 130 language models, we show that ONLINE LLM PICKER saves annotation cost by up to 71.67% while reliably identifying the best or near-best model for the stream. We also show that using the returned model for sequential generation on unannotated prompts across the stream reduces regret by up to a factor of 2.51x, indicating that ONLINE LLM PICKER can identify the best or near-best model well before processing all streaming prompts.
Chinese Translation
大语言模型(LLMs)越来越多地应用于处理流式数据,实践者依赖基准来选择最佳模型,尽管这些信号只能近似真实性能。虽然 oracle 标注能够提供可靠反馈,但它们通常成本高昂且难以大规模获得。为解决这一挑战,我们提出了 ONLINE LLM PICKER,这是首个用于在线环境下 LLM 主动模型选择的框架。给定任意查询流和有限的标注预算,ONLINE LLM PICKER 选择信息量最大的提示进行标注,以在候选模型中识别最佳 LLM。在包括 10 个数据集在内的多个任务上,针对超过 130 个语言模型,我们表明 ONLINE LLM PICKER 最高可节省 71.67% 的标注成本,同时可靠地识别出该流的最佳或接近最佳模型。我们还表明,使用返回的模型在流中未标注的提示上进行顺序生成,最高可将遗憾(regret)降低 2.51 倍,这表明 ONLINE LLM PICKER 能够在处理完所有流式提示之前就很早识别出最佳或接近最佳模型。
cs.CL / 73 / 2610.01616
Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?
LLMs 能否可靠地注释生物测定元数据以提升数据就绪度?
large language model
大语言模型相关
Abstract
The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's $\sim$2 million bioassays is critically sparse, 36\% lacking an assay format, 89\% a BioAssay type, and >99.9\% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.
Chinese Translation
用于分子性质预测的基础模型的出现,要求具备高度的 AI 数据就绪度,包括可靠的元数据注释。然而,公共存储库和工业筛选数据库都受困于测定注释缺失、不一致或混同的问题。在这项工作中,我们量化了 PubChem 中生物测定本体(BAO)的测定格式和物理检测方法字段的缺失注释程度,并研究开源和专有的大语言模型(LLMs)能否直接从测定文本中可靠地预测和审核元数据注释。在我们的评估中,我们发现 PubChem 的 $\sim$2 百万个生物测定中的注释覆盖情况极其稀疏:36\% 缺少测定格式,89\% 缺少 BioAssay 类型,且 >99.9\% 缺少任何 BAO 映射的测定格式或检测技术术语。这促使人们需要自动化测试元数据整理。使用从 PubChem 和 ChEMBL 衍生的评估集,我们评估了七个开源和专有 LLM 与现有银标签的一致性。对于生化测定和基于细胞的测定格式,召回率至少为 0.96,检测技术也呈现类似模式,尽管在代表性不足的类别上分歧增加。人工检查显示,这些分歧中有许多可追溯到银标签来源之间的不一致,而非 LLM 错误。此外,在一项与资深工业策展人进行的定性研究中,LLM 生成的证据促使该专家修订了其自身的一些标签,表明 LLM 能够标记可能被错误标注的测定。在整个研究中,专有模型与开源模型之间的性能差异很小。综合来看,这些结果表明 LLM 可以支持测定元数据的大规模注释和审核,尽管在此类标签进入下游 ML 流水线之前,逐类可靠性估计和有针对性的人工审查仍然是必需的。
cs.CL / 74 / 2610.01634
Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá
Yo-ByT5:面向约鲁巴语的高效高保真变音符号还原
large language model
大语言模型相关
Abstract
Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.
Chinese Translation
约鲁巴语是一种广泛使用的声调语言,它依赖变音符号来避免词汇歧义。然而,它常常在不带这些变音符号的情况下被书写,从而阻碍了下游自然语言处理(NLP)任务。在本文中,我们提出 Yo-ByT5,一个从 ByT5-small 微调而来的字节级自动变音符号还原(ADR)模型。我们在一致的协议下,在 YAD 基准上将 Yo-ByT5 与五个公开释出的约鲁巴语 ADR 模型以及一个开放权重的大语言模型(LLM)一同进行评估。我们的结果表明,Yo-ByT5 达到了现有最强模型 mT5-base 的性能,其 DER 为 10.14%,CER 为 3.48%。此外,尽管其参数量约为 mT5-base 的一半,它仍展现出更优的文本保真度。我们还发布了我们的训练代码和模型输出,并呼吁开发一个更大、专门为约鲁巴语变音符号还原而构建的基准。
cs.CL / 75 / 2610.01696
Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information
Acmite:通过概念引导的互信息缓解 LLMs 中的性别偏见
large language model
大语言模型相关
Abstract
Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.
Chinese Translation
大语言模型(LLMs)会复现其训练数据中的社会刻板印象,这促使了关于模型去偏的大量研究。然而,现有方法通常依赖于显式的偏见示例或预定义的群体术语替换,这使得它们对措辞敏感,并且在捕捉跨不同语境共享的刻板印象概念方面效果较差。更重要的是,它们通常抑制有偏输出,而没有显式地建模模型输出与潜在刻板印象概念之间的统计依赖关系。我们提出 Acmite,一个用于有针对性且选择性去偏的轻量级概念引导框架。Acmite 将刻板印象表示为结构化语义概念,并使用最大边际相关性(MMR)来选择多样化的概念以进行去偏。受互信息最小化启发,它使用 token 级 KL 散度来近似这种依赖关系,同时保留任务语义。一个轻量级 LoRA 适配器在基础模型冻结的情况下进行训练,并在推理时仅当输入与刻板印象相关概念足够相似时才被激活;否则,直接使用原始模型。我们在 BBQ、CrowS-Pairs 和 StereoSet 上评估 Acmite,并在 ARC-Challenge、GSM8K 和 PIQA 上评估通用能力保持情况。跨三个 LLM 的实验表明,Acmite 在互补的评估格式中有效缓解性别偏见,同时在偏见无关任务上保持有竞争力的性能。匿名代码和数据可在 https://anonymous.4open.science/r/Acmite-18E2/ 获取。
cs.CL / 76 / 2610.01767
A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering
面向高效多跳问答的 Matryoshka 分层 RAG
large language model
大语言模型相关
Abstract
Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
Chinese Translation
用于多跳问答(QA)的检索增强生成(RAG)系统必须在检索质量与计算成本之间取得平衡。这种成本产生于索引阶段,即通过使用昂贵的知识图谱(KG)或大语言模型(LLM)来生成摘要,或者产生于查询阶段,即通过由 LLM 驱动的迭代式检索。为了在保持检索质量的同时降低该成本,我们提出了 MatRAG,一个将 RAG 系统与 Matryoshka 表示学习(MRL)相结合的分层框架。MatRAG 通过将聚类结构的语义层级与 MRL 的嵌套结构对齐,来同时应对这两类成本。具体而言,它将文档语料组织为一个由簇构成的有向无环图(DAG),其粒度逐级变粗。每一层级都由更低的 Matryoshka 维度进行索引。MatRAG 将 DAG 的迭代式自顶向下遍历与一种实体驱动的机制相配对,该机制控制跳数预算并对候选进行重排序。我们在三个标准多跳问答基准上,针对七个具有代表性的基线方法对 MatRAG 进行了评估。在检索质量方面,MatRAG 优于其最强的竞争对手;此外,它通过避免 KG 构建和基于 LLM 的摘要生成降低了索引成本,并通过维度感知的相似度降低了查询时成本。
cs.CL / 77 / 2610.01938
A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined
用于评估大语言模型临床推理的评分量规图景:现存什么、缺失什么,以及需要组合什么
large language model
大语言模型相关
Abstract
Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.
Chinese Translation
考试式准确率并不能确立大语言模型(LLMs)是否能在临床记录上进行良好推理。我们将临床推理定义为跨时间和来源整合并更新证据,以形成、修订并论证患者的问题表征和可辩护的计划。本结构化叙述性综述梳理三类文献:医学教育评估工具、2023年起发表的临床LLM基准,以及评估长文本生成的通用领域方法。我们考察六个维度:问题表征、时间综合、鉴别或管理推理、反事实推理、校准不确定性和推理忠实性。预印本被纳入并标注。没有任何单一工具覆盖全部六个维度。问题表征以及鉴别或管理推理得到了合理覆盖,尽管可靠性因工具和情境而异。TIMER-Eval针对时间综合,ER-Reason评估序贯诊断信念更新。专门的不确定性和反事实评估正在涌现,但其对纵向自由文本推理的适用性仍然有限。事实完整性在通用领域评估中已得到充分理论化,并有早期临床证据表明存在重要遗漏。忠实性仍然是最薄弱的维度,仅有一项已识别的针对多项选择题的临床因果消融研究。现有工具应通过二元评分细则条目、单独的完整性和正确性评分、带有不可补偿安全上限的病例特定重要性加权、时间顺序一致性检查,以及经机会校正的可靠性报告来组合。需要针对纵向自由文本记录上的校准不确定性、反事实推理和忠实性开展进一步设计工作。本综述提供的是设计依据,而非经过验证的工具。
cs.CL / 78 / 2610.01984
Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
通用字节级编码:通过 UTF-8/UTF-16 路由减少跨文字系统的 token 预算差异
large language model
大语言模型相关
Abstract
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
Chinese Translation
字节级字节对编码(BBPE)分词器对多语言大语言模型(LLM)具有吸引力,因为它们覆盖了所有 Unicode 文本。然而,在基于 UTF-8 的 BBPE 中,许多文字系统的起始回退成本就高于英语:当无法应用任何已学到的合并时,一个多字节字符需要多个由字节派生出的符号。我们将这种最坏情况下的合并前成本称为编码下限。更高的编码下限会增加 token 数量与每次请求的成本,并缩小可用上下文。改变文本编码可以缩小这一差距,但单一的全局编码会使混合文字系统文本中原本已高效的英文片段变得更加昂贵。我们提出通用字节级编码(UBE),这是一种双字母表分词器,它将 1-2 字节的 UTF-8 字符保留在 UTF-8 路径上,同时将 3-4 字节的 UTF-8 字符经由 UTF-16 进行路由。这降低了具有高 token 溢价(相对于英语的 token 数量)的文字系统中 3 字节基本多文种平面(BMP)字符的编码下限,同时不会提高混合文字系统文本中原本已高效片段的编码下限。UBE 仅改变呈现给字节对编码(BPE)的字节表示;合并规则仍为标准规则,且精确解码得以保持。UBE 还可与替代的边界策略以及基于形态学的表示方式组合使用。在一项 Unicode 17 审计中,UBE 对所有 Unicode 标量值以及官方规范化、字素断行和表情符号测试套件中的所有输入都能精确地往返转换。在各项内在评估中,UBE 降低了以英语归一化的 token 计数比值的离散程度,从而减小了跨语言的 token 预算差异。在多语言语言模型(LM)实验中,UBE 达到了与 BBPE 相当的 LM 质量。在主要的多语言设置中,UBE 对高溢价文字系统的 token 数量削减最多,并略微降低英语的 token 数量,从而在固定 token 预算下带来更多可用上下文,并在内容匹配的基准测试中实现更快的提示处理。
cs.CL / 79 / 2610.02002
Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
Mem++:面向长期组织型 LLM 智能体的非破坏性记忆
large language model
大语言模型相关
Abstract
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.
Chinese Translation
大型语言模型(LLM)智能体如今参与到组织工作中,在这些工作中,许多作者在数月内跨文档记录决策。由于修订后的决策以新文档而非编辑的形式出现,回答一个问题需要知道在给定时间哪个版本有效。然而,大多数记忆系统在写入时压缩记录。通过将每份文档蒸馏为事实、笔记或图边,这些方法在任何问题被提出之前就固定了可回答的内容。为解决这一问题,我们提出 Mem++,一个从写入时蒸馏转向读取时选择的非破坏性记忆框架。Mem++ 将每份文档连同其日期和作者完整存储,并且在写入时不调用任何生成模型。在读取时,它仅检索日期截至问题所询问时间的文档,并融合词汇和语义排序。与覆盖旧版本的系统不同,Mem++ 保留它们,并将选择权留给回答模型。在组织基准 OrgMemBench 上的评估表明,Mem++ 在两个回答模型上以 8.0 到 13.1 分的优势超过最强的记忆系统基线。在使用 gpt-4.1-mini 时,它还取得了最佳总体分数,比 RAG 高 2.6 分。此外,Mem++ 在 LoCoMo 上取得了最佳平均 LLM 评判分数,并在 LongMemEval-S 上排名第二,仅次于其实体图变体。基准评估代码可在 https://github.com/AIDAChip-Inc/mem-plus-plus 获取。
cs.CL / 80 / 2610.02022
Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation
旧思想,新问题:基于LLM的新颖性评估的不稳定性
large language model
大语言模型相关
Abstract
Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper's originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods.
Chinese Translation
自动化构思系统通常根据其产生想法的新颖性来评估,而这一判断正越来越多地被委托给大型语言模型。这类评判器通常是临时构建的,并且即便经过验证,也往往是在人类撰写的论文上验证,而不是在其本应评分的人工生成想法上验证。那么,新颖性评判器的表现如何?并不好。我们提出了一项关于新颖性评估设计选择的系统性对照研究。我们首先自动构建了一个评估集,从OpenReview中挖掘审稿人明确肯定或质疑论文原创性的段落,并且只保留在其研究领域两端获得一致同意的投稿;我们将这些与来自一个普通LLM生成器的想法配对。在六个评判器上,我们发现微小的提示设计选择会产生巨大后果;例如,仅仅告诉评判器审稿人认为其中一个想法新颖而另一个不新颖,就能改变其对所展示的相同想法对中超过一半的判定,使配对准确率变动超过50个点,偶尔甚至将其推到低于随机水平。同样的改动会帮助一个评判器,却损害另一个评判器。检索和更大的推理预算帮助甚微,而两个专门构建的新颖性评估器被我们最廉价的提示基线所超越。这些结果对自动化构思系统所报告的新颖性提升提出了质疑,并呼吁建立稳健的新颖性评估方法。
cs.CL / 81 / 2610.02092
Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)
可扩展、可迁移的数据选择元网络需要一种不同的损失(以及为什么显而易见的选择是有问题的)
large language model
大语言模型相关
Abstract
Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.
Chinese Translation
数据选择对于在庞大且异构的语料库上训练大型语言模型至关重要。用于训练数据选择的元学习(Meta-learning for Training-data Selection)通过从目标验证目标中学习数据权重,为启发式评分提供了一种有原则的替代方案,但现有方法面临着细粒度估值与对未见数据的可迁移性之间的权衡。一个自然的解决方案是用选择网络替代逐样本权重。然而,我们发现,将此类网络直接纳入现有的 MTS 目标会导致优化不稳定和泛化能力差,其原因是权重抑制以及对易学特征的持续依赖。为解决这些问题,我们提出了可迁移样本评分与选择(Transferable Example Scoring and Selection, TESS),这是一个建立在逐点价值匹配目标(Pointwise Value Matching objective, PVM)之上的可扩展数据选择框架。在 LLM 安全性和定向指令微调上的实验表明,其能够在数据集之间、从子集到完整语料库、以及从较小模型到较大模型之间实现强迁移。
cs.CL / 82 / 2610.02150
From Knowledge Access to Source Learning: Developing Source-Specific Competence
从知识获取到源学习:发展源特定能力
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.
Chinese Translation
大型语言模型(LLM)智能体越来越依赖持久的外部源来解决一系列知识密集型任务。现有方法改进了源内容的访问和组织方式,而智能体记忆系统则保存了先前交互中可复用的知识,但对同一源的重复使用在很大程度上仍被视为重复访问,而非逐步提升对该源理解的机会。我们研究源学习:针对一个持久权威源发展可复用的源特定能力。我们用持久源模型来表示这种能力,该模型捕获对源的可复用理解,包括其知识如何被结构化、解释和应用。为了构建并逐步细化此类模型,我们提出 SourceLearn,它结合了两种互补的学习机制。自主导向的源学习识别出哪些内容仍未得到充分理解,并自适应地重新访问源;而任务引导的源学习则利用下游经验来揭示局部表征缺口以及源知识应如何组织的反复出现的需求。在两种情况下,学习信号决定应重新考虑什么,而持久更新则从权威源重建。在五个基准和三个 LLM 后端上,SourceLearn 在 15 种设置中的 13 种取得了最佳性能,相比 Hybrid RAG 的提升最高达 22.6 分,并在整体上显著优于静态源表示和基于经验的记忆基线。
cs.CL / 83 / 2610.02193
Hierarchical Continuous Diffusion Language Models
分层连续扩散语言模型
diffusion
扩散模型相关
Abstract
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.
Chinese Translation
离散扩散语言模型为需要双向推理和全局约束满足的任务提供了对自回归生成有吸引力的替代方案。然而它们共享一个结构性瓶颈:当并行解码时,每个词元都从其边缘分布中独立采样,切断了同时解码的词元之间的统计依赖关系。连续扩散语言模型通过对共享连续状态去噪来避免这一点,但它们的去噪器只看到该状态,因此在最终解码之前,没有任何东西将其与有效的词元配置绑定起来。为解决这一问题,我们提出分层连续扩散语言模型(HC-DLM),它在单一且原则性的去噪过程中将离散词元生成与连续潜在轨迹耦合起来,其训练目标由词元似然的变分界推导而来。与近期将连续上下文附加到自包含离散链上的方法不同,HC-DLM 使潜在状态成为唯一持久的生成状态:词元在每一步从其中读出,并作为下一次潜在更新的支架反馈回去。在结构化推理(数独)、数学规划(Countdown)和语言建模(LM1B)上,在匹配模型规模的情况下,HC-DLM 在数独和 Countdown 上的谜题准确率以及在 LM1B 上的生成困惑度方面均优于离散和连续扩散基线。项目页面:https://hc-dlm.github.io/。
cs.CR / 84 / 2610.00557
No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents
没有一种架构能适合所有情况:分层红队智能体的跨环境评估
large language model
大语言模型相关
Abstract
Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates conceal.In the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.
Chinese Translation
自主红队智能体越来越多地通过规划策略和执行多阶段攻击,对 AI 赋能的网络防御进行压力测试。强化学习(RL)和大语言模型(LLM)为此类智能体所需的规划与执行提供了互补机制,已有工作已将它们结合在混合层级结构中。然而,特定架构通常是在单一环境内开发和评估的,这使得观察到的优势究竟反映了一种普遍更强的决策机制,还是仅仅契合了某一特定设置,仍悬而未决。我们通过对两种同质分层红队架构进行受控的跨环境比较来填补这一空白:一种是由 RL 规划器和 RL 执行器组成的架构(RL+RL),另一种是由 LLM 规划器和 LLM 执行器组成的架构(LLM+LLM)。我们在 CybORG CAGE-4 和 Cyberwheel 中的两种网络规模下,针对专家级自主防御者评估了这两种架构,在同一个统一破坏指标下覆盖了 18 种配置。我们发现了一种显著的环境依赖性反转。RL+RL 在紧凑且奖励密集的 CAGE-4 中获胜(破坏成功率为 78.5%,而最强的 LLM 配置为 18.0%),并在 100 主机 Cyberwheel 网络中获胜(81.0% 对 50.5%),而一个预训练的网络安全 LLM 智能体在更大、受升级门控的 1010 主机 Cyberwheel 网络中获胜(55.0% 对 RL 的 0.0%)。一项杀伤链分析通过聚合成功率所掩盖的架构特定瓶颈解释了这种反转。在 1010 主机 Cyberwheel 网络中,RL 能发现并攻陷主机,但在权限提升处停滞;而在 CAGE-4 中,LLM 智能体获得了特权访问,却很少将其转化为实际作战影响。这些结果表明,在单一环境中得出的结论可能无法推广,且混合式规划器-执行器设计应由特定的失效模式驱动,而不是基于某一种架构普遍更优的假设。
cs.CR / 85 / 2610.00590
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
迈向使用大型语言模型的分层网络防御:从规划到执行
large language model
大语言模型相关
Abstract
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATT&CK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.
Chinese Translation
使用强化学习(RL)训练的自主网络防御者通常与其训练所在的网络绑定,这限制了其在网络规模变化时的泛化能力。分层RL通过将战略目标选择与战术执行分离来降低决策复杂度,但它并未消除这种重训练依赖。我们研究冻结的、零样本大型语言模型(LLMs)能否在分层网络防御中提供无需重训练的控制,以及当LLM控制从规划扩展到执行时性能如何变化。我们形式化了一个与控制器无关的规划器-执行器分层结构,其中规划器在固定时域内选择一个要防御的子网,执行器在该子网内选择防御动作。使用高保真Cyberwheel环境及其内置的、映射到MITRE ATT&CK框架的自动化红队智能体,我们在小型、中型和大型网络上比较了RL+RL、LLM+RL和LLM+LLM配置,使用了从3B到70B参数不等的六个模型,其中包括两个网络安全专用模型。仅将规划器替换为LLM时,随着网络规模增大,所获得的增益有限。相比之下,将LLM控制扩展到执行会为能力足够强的模型带来显著改进。例如,一个冻结的通用70B模型在全部三种网络规模上使用相同的模型权重,将成功的横向移动控制在约1%的步骤内,并将攻击者影响保持在接近零,而RL基线则针对每种规模重新训练。我们的结果表明,能力足够强的冻结LLM可以在所评估的网络规模上无需任务特定重训练即可保持强防御性能,同时也表明强大的战术执行对于实现基于LLM的控制所带来的收益很重要。
cs.CR / 86 / 2610.00839
Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
针对LLM提取的防御能否跨攻击奏效?黑盒模型提取的生命周期基准
large language model
大语言模型相关
Abstract
Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access assumptions, model configurations, query budgets, and security objectives, limiting comparability across methods. To address this gap, we introduce a unified benchmark covering six extraction attacks, ten defenses, and two adaptive attacks that paraphrase or back-translate protected responses before surrogate training. The benchmark controls model configurations, query data, budgets, and held-out evaluation conditions within each comparison while preserving attack-specific querying and training procedures. We measure surrogate capability, fidelity to the victim, output quality using Rep-4, and query-budget sensitivity; defenses use their own security metrics paired with surrogate performance. For the adaptive attacks, we jointly measure provenance-detector scores and the capability and fidelity of surrogates trained on rewritten responses. The benchmark thus provides a reproducible basis for comparing extraction methods and their interactions with defenses under text-only access. Code and artifacts are available at https://github.com/sliu11-byte/MEA-Bench.
Chinese Translation
通过纯文本API部署的大语言模型(LLM)面临模型提取风险,因为攻击者可以收集其响应来训练能够复现其能力的替代模型。尽管已有工作提出了多种多样的攻击与防御方法,但评估在访问假设、模型配置、查询预算和安全目标方面仍然彼此割裂,限制了不同方法之间的可比性。为弥补这一空白,我们提出了一个统一的基准,涵盖六种提取攻击、十种防御,以及两种自适应攻击——后者在替代模型训练之前对受保护的响应进行改写或回译。该基准在每组比较中控制模型配置、查询数据、预算和留出评估条件,同时保留各攻击特有的查询与训练流程。我们测量替代模型的能力、对受害模型的保真度、使用Rep-4衡量的输出质量以及对查询预算的敏感性;防御则使用其自身的安全指标并搭配替代模型性能。对于自适应攻击,我们联合测量来源检测器得分,以及在改写后响应上训练的替代模型的能力与保真度。因此,该基准为在纯文本访问条件下比较提取方法及其与防御的交互提供了一个可复现的基础。代码与工件可在 https://github.com/sliu11-byte/MEA-Bench 获取。
cs.CR / 87 / 2610.01058
MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs
MOMAT:用于量化LLM低功耗越狱防御的多图集混合
large language model
大语言模型相关
Abstract
Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top-$k$ similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT's CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a $4.69 \times 10^6\times$ speedup) and reduces energy from $8.1 \times 10^7$ $μ$J to 3.32 $μ$J, yielding an approximately $2.5 \times 10^5\times$ energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.
Chinese Translation
量化大语言模型因其低延迟和高能效而越来越多地部署在边缘设备上。然而,模型量化削弱了对齐安全防护,使 qLLM(量化大语言模型)极易受到越狱攻击。为应对这一挑战,我们提出 MOMAT(多图集混合,Mixture of Multiple Atlases),一个将结构化知识检索与低功耗防御加速相结合的硬件增强安全框架。每个图集代表有害或良性样本集与策略模板的一个语义簇,从而实现领域局部化的检索增强生成防护,缓解大型异构安全数据库中的维数灾难以及由此产生的语义稀疏问题。MOMAT 针对每个提示从所有图集中检索 top-$k$ 相似度特征,并使用轻量级 MoE(专家混合)检测器对其进行评估,同时由 CiM(存内计算)加速的相似度引擎执行快速、低功耗的图集局部检索。MOMAT 基于 CiM 的检索将 100 次查询批处理从 15,052.44 ms 加速到 3,207.21 ns(加速比达 $4.69 \times 10^6\times$),并将能耗从 $8.1 \times 10^7$ $μ$J 降低到 3.32 $μ$J,相较基于 DRAM(树莓派)的基线实现了约 $2.5 \times 10^5\times$ 的能耗降低。跨标准基准的红队评估表明,MOMAT 在防御性能上可与最先进方法相媲美,同时避免了良性样本的过度拦截,并带来显著的效率提升,证明基于 CiM 的模块化防御能够使边缘部署的 qLLM 既更安全又更节能。我们将发布完整的 223.2k 样本数据集,以促进未来研究。
cs.CR / 88 / 2610.01079
Jev-IDS: System One Models for Network Intrusion Detection
Jev-IDS:用于网络入侵检测的系统一模型
large language model
大语言模型相关
Abstract
Machine-learning Network Intrusion Detection Systems (IDS) depend on substantial labeled datasets and task-specific training, whereas Large Language Models (LLMs) detection can analyze flow records directly but incurs higher inference cost and latency, with less constrained outputs. This paper presents JEV-IDS, an open experimental general NIDS based on the Jev System One Model (SOM) to detect zero day intrusions Under label scarcity. JEV-IDS serializes one flow per request and asks JEV two questions: a binary attack probability and a finite-choice traffic category. Our results show that, at k=1, JEV was 4.8 times faster and 3.8 times cheaper than GPT-5.6 Luna, with 1.5 times higher novel-attack recall; it also produced 15 times fewer false alarms than a low-data Random Forest. Across 5,400 decisions on a 300-flow NSL-KDD pilot split, JEV achieved F1-Score 0.859, precision 0.941, recall 0.790, and novel-attack recall 0.838. Increasing k to 2 reduced its F1-Score to 0.839.
Chinese Translation
基于机器学习的网络入侵检测系统(IDS)依赖于大量标注数据集和任务特定的训练,而大语言模型(LLM)检测可以直接分析流记录,但会带来更高的推理成本和延迟,且输出受到的约束更少。本文提出了 JEV-IDS,一个基于 Jev 系统一模型(SOM)的开放式实验性通用 NIDS,用于在标签稀缺条件下检测零日入侵。JEV-IDS 每次请求序列化一条流,并向 JEV 提出两个问题:一个二值攻击概率和一个有限选项的流量类别。我们的结果表明,在 k=1 时,JEV 比 GPT-5.6 Luna 快 4.8 倍、便宜 3.8 倍,且新型攻击召回率高 1.5 倍;与低数据量的随机森林相比,它产生的误报还少 15 倍。在 300 条流的 NSL-KDD 试点划分上做出的 5,400 次决策中,JEV 取得了 F1 分数 0.859、精确率 0.941、召回率 0.790 以及新型攻击召回率 0.838。将 k 增至 2 后,其 F1 分数降至 0.839。
cs.CR / 89 / 2610.01263
Autonomous OSS Threat Detection via Taxonomy-Aligned LLMs
通过分类体系对齐的大语言模型实现自主 OSS 威胁检测
large language model
大语言模型相关
Abstract
Open source software (OSS) ecosystems face growing threats from sophisticated supply chain attacks including typosquatting, dependency confusion, Trojan Source obfuscation, malicious build injection, and CI/CD pipeline poisoning. Existing detection approaches rely on signature-based tools and rule-based systems that struggle to generalize across attack variants and emerging threat patterns. In this paper we propose a taxonomy-aligned large language model framework for automated detection and classification of OSS supply chain threats. We introduce a structured AV-xxx threat taxonomy covering five attack categories and construct a curated dataset of 999 verified real-world OSS supply chain incidents sourced from GitHub Security Advisories, CISA alerts, and security research reports spanning 2018 to 2026. Using taxonomy-aligned prompt engineering with GPT-4, our framework achieves 97.0\% multi-class classification accuracy and 97.0\% macro F1 score across all five threat categories. Comparative evaluation against five traditional machine learning baselines, one zero-shot open source LLM, and two fine-tuned neural models reveals a surprising finding: fine-tuned Llama 3.1 8B (70.5%) and SecRoBERTa (77.5%) both underperform simple TF-IDF classifiers (82.3%), while Mistral 7B without taxonomy alignment achieves only 65.7%. These results confirm that taxonomy-aligned prompting rather than model scale, domain pretraining, or fine-tuning is the critical factor enabling high classification accuracy. Our dataset and code are publicly available to support reproducible supply chain security research.
Chinese Translation
开源软件(OSS)生态系统面临来自复杂供应链攻击的日益增长的威胁,包括域名仿冒(typosquatting)、依赖混淆(dependency confusion)、Trojan Source 混淆、恶意构建注入以及 CI/CD 流水线投毒。现有检测方法依赖于基于特征签名的工具和基于规则的系统,这些方法难以泛化到攻击变体和新兴威胁模式。在本文中,我们提出了一种分类体系对齐的大语言模型框架,用于对 OSS 供应链威胁进行自动检测和分类。我们引入了一个结构化的 AV-xxx 威胁分类体系,涵盖五类攻击,并构建了一个经过整理的包含 999 个已验证真实世界 OSS 供应链事件的数据集,该数据集来源于 GitHub Security Advisories、CISA 警报以及涵盖 2018 年至 2026 年的安全研究报告。使用 GPT-4 进行分类体系对齐的提示工程,我们的框架在所有五类威胁上实现了 97.0\% 的多分类准确率和 97.0\% 的宏 F1 分数。与五个传统机器学习基线、一个零样本开源 LLM 和两个微调神经模型的比较评估揭示了一个令人惊讶的发现:微调后的 Llama 3.1 8B(70.5%)和 SecRoBERTa(77.5%)都表现不如简单的 TF-IDF 分类器(82.3%),而未进行分类体系对齐的 Mistral 7B 仅达到 65.7%。这些结果证实,实现高分类准确率的关键因素是与分类体系对齐的提示,而不是模型规模、领域预训练或微调。我们的数据集和代码已公开可用,以支持可复现的供应链安全研究。
cs.CR / 90 / 2610.01349
PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
PACE:面向使用工具的LLM智能体的溯源感知能力强制执行
large language model
大语言模型相关
Abstract
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.
Chinese Translation
使用工具的大型语言模型(LLM)智能体将生成的文本转化为真实的副作用,因此被投毒的工具元数据、检索到的页面、记忆和可复用技能可以操纵下一次调用。在准入前审查一个制品并不能解决这个问题。一个安全变体和一个泄露变体可以产生相同的准入证据,而一个健全的门控此时无法为二者中的任何一个放宽该位置。我们将该条件精确化,这留下了部署仍能采取行动的最后边界。我们提出溯源感知能力强制执行(PACE),它在每次工具调用执行之前立即对其进行中介。路径限制提出了对已表示的影响路径的一种可执行割,而能力与效果验证根据从经认证请求编译出的授权,检查模式定义的效果。我们区分经认证的执行契约与经评估的配置,后者可以在提议阻止之后恢复一次经授权的调用,或应用所声明的修复。限制要求最终动作保持经认证的割。在八个可执行的智能体安全基准上,涉及三个目标模型家族,经评估的配置在79个符合条件的攻击列中的62个中给出严格最低的攻击成功率,并在14个中打平;全基准原生效用相对于未防御智能体最多损失三个点。对1167个配对案例的完整消融将大部分安全增益归因于效果验证,并将拒绝控制归因于边界适配。一个缩小规模的自适应搜索针对该防御在0/30个越权目标上取得成功。
cs.CR / 91 / 2610.01365
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
沉睡的秘密:微调如何重新唤醒语言模型中的隐私风险
large language model
大语言模型相关
Abstract
Beyond adapting Large Language Models (LLMs) to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same distribution, i.e., the previous training dataset. We argue that such recovery remains possible without such impractical knowledge. We show that LLM-generated candidates can provide sufficient supervision to recover previously learned private associations. Based on this, we propose ReGap, a data-free attack that recovers private associations using task structure, filters them by answer-token likelihood, and updates the target model via low-rank adaptation. Specifically, ReGap requires neither target answers nor auxiliary genuine private supervision. Across six GPT-2, OPT, and Qwen3 models, ReGap improves target-association recovery by 6-21 percentage points over the post-training target model. Recovery remains substantial even when the adaptation identities are disjoint from all memorized and evaluation identities, with no exact target answers appearing in the generated or selected supervision. Moreover, the same trained adapters increase recovery from 42\% to 63\% on a previously exposed checkpoint, but produce no gain on a matched checkpoint that never encountered the targets. This contrast shows that adaptation alone is insufficient to explain the observed recovery and that prior target exposure strongly affects post-adaptation recoverability. Our findings highlight that routine model customization can reawaken latent privacy risks, warranting urgent attention from the academic and industrial communities.
Chinese Translation
除了将大型语言模型(LLMs)适配到专门应用之外,最近的研究表明,微调可以恢复通过直接查询不再可访问的私人信息。然而,先前的微调恢复攻击需要从相同分布(即先前的训练数据集)中提取的真实私人监督。我们认为,在没有这种不切实际的知识的情况下,这种恢复仍然是可能的。我们表明,LLM生成的候选可以提供足够的监督,以恢复先前学习到的私人关联。基于此,我们提出了ReGap,一种无数据攻击,它利用任务结构恢复私人关联,通过答案令牌似然对其进行过滤,并通过低秩适配更新目标模型。具体来说,ReGap既不需要目标答案,也不需要辅助的真实私人监督。在六个GPT-2、OPT和Qwen3模型上,与训练后的目标模型相比,ReGap将目标关联恢复率提高了6-21个百分点。即使适配身份与所有记忆身份和评估身份都不相交,且生成或选择的监督中没有出现确切的目标答案,恢复仍然显著。此外,相同的已训练适配器在先前暴露的检查点上将恢复率从42\%提高到63\%,但在从未遇到过目标的匹配检查点上没有产生任何增益。这种对比表明,仅靠适配不足以解释观察到的恢复,并且先前的目标暴露强烈影响适配后的可恢复性。我们的发现强调,常规的模型定制可能会重新唤醒潜在的隐私风险,值得学术界和工业界紧急关注。
cs.CR / 92 / 2610.01367
High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection
高质量数据并不意味着安全!数据选择后对大型语言模型进行投毒
large language model
大语言模型相关
Abstract
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fill this gap, we systematically evaluate both the filtering effects against poisoning and the downstream safety impact of retained data. The results reveal that selection removes many overtly harmful samples, yet some retained high-quality samples can still degrade model safety alignment possibly due to their harmful-like training-update patterns at the layer-wise gradient level. Together, these findings expose a practical vulnerability: safety-degrading influence can pass through quality-based selection via retained high-quality samples. To examine its systematic exploitability, we propose Bi-Stage Quality-Constrained Safety-Degradation Text Optimization (Bi-QSTO), which optimizes poisoned samples under an explicit quality constraint to survive selection while preserving their safety-degrading influence. Across poisoning settings, target models, and filtering rates, Bi-QSTO maintains attack effectiveness before and after selection. Even at 90% filtering, harmful-seeded samples achieve a Poisoning Retention Rate above 90% and Harmful Score of 3.30--4.01. Their attack effectiveness strongly transfers across models and their retention advantage generalizes to additional selection methods.
Chinese Translation
安全对齐的大型语言模型仍然容易受到在小规模有害或看似良性的样本上进行微调的影响。然而,先前研究通常假设投毒样本直接进入下游微调,忽视了实际训练流程中基于质量的选择。为填补这一空白,我们系统评估了针对投毒的过滤效果以及保留数据对下游安全性的影响。结果表明,选择会移除许多明显有害的样本,但一些被保留的高质量样本仍可能降低模型的安全对齐,这可能是因为它们在逐层梯度层面具有类似有害样本的训练更新模式。这些发现共同揭示了一个实际漏洞:降低安全性的影响可以通过被保留的高质量样本穿过基于质量的选择。为考察其系统性可利用性,我们提出双阶段质量约束安全退化文本优化(Bi-Stage Quality-Constrained Safety-Degradation Text Optimization, Bi-QSTO),该方法在显式质量约束下优化投毒样本,使其能够通过选择,同时保留其降低安全性的影响。在各种投毒设置、目标模型和过滤率下,Bi-QSTO 在选择前后均保持攻击有效性。即使在 90% 过滤下,有害种子样本仍达到超过 90% 的投毒保留率(Poisoning Retention Rate)以及 3.30--4.01 的有害分数(Harmful Score)。它们的攻击有效性在不同模型之间具有很强的迁移性,并且其保留优势可泛化到额外的选择方法。
cs.CR / 93 / 2610.01871
Walking the Embedding Space: Datastore Extraction from Multimodal RAG
在嵌入空间中漫步:从多模态 RAG 中提取数据存储
large language model
大语言模型相关
Abstract
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.
Chinese Translation
多模态检索增强生成(MRAG)已成为一种可靠且成本效益高的技术,用于将多模态大语言模型(MLLMs)的生成能力建立在相关的、最新的外部知识之上。尽管带来了若干好处,例如减少幻觉行为,它们也引入了新的攻击面,包括隐私信息泄露以及针对数据提取攻击的脆弱性。在本文中,我们介绍 $\immrag$,一种自适应且自动的数据提取攻击流程,它在黑盒设置下针对图像返回型 MRAG 运行,这种配置中检索到的视觉产物本身就是响应。每个查询将攻击者持有的影子图像与一幅已从系统中恢复的图像混合,并且相关性加权的重采样将后续查询导向嵌入空间中仍能产生新颖检索结果的区域。与当前那些旨在通过将恶意查询作为文本提示放置来诱导模型泄露数据的提取攻击不同,$\immrag$ 将恶意指令嵌入到用户给定的输入图像内部。我们在三个合理且不同的真实世界场景中评估 $\immrag$:医疗助手、以文档为中心的助手和通用工具。这些实验涉及研究该攻击在多个 CLIP 系列检索器上的有效性,以及各种生成器的影响。一次 2500 次查询的运行在局部特征对应下重建了多达 611 幅不同的放射学图像、566 份文档扫描件和 416 幅通用图像,并使不同数据存储项的数量达到非自适应基线的多达 $5.6\times$。我们的结果表明,迫切需要专门为多模态数据设计的安全保障措施。
cs.LG / 94 / 2610.00686
SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
SemanTok:用于高效自回归视频生成的可预测语义标记
diffusion
扩散模型相关
Abstract
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
Chinese Translation
最近基于视频的世界模型将自回归(AR)预测的可扩展性与扩散模型的视觉质量结合起来。场景标记器的选择对于这两方面各自达到最优性能都至关重要,无论是在保真度还是语义方面。灵活长度、由粗到细的标记器恰好能实现这一点:最初的粗粒度标记承载视频片段的全局语义,而后续标记进一步指定细节。现有灵活标记器仅在解码器早期隐藏状态上应用表示对齐(REPA)损失,而这一目标其实是解码器可以部分地从其加噪输入中满足的。我们提出 SemanTok,一种灵活视频标记器,它将冻结的 DINO 特征馈入其编码器,并添加轻量级头,这些头仅从每个保留的标记前缀重建这些特征。SemanTok 在每个 AR 模型规模下都实现了高语义对齐和视频保真度:一个 201M 的 SemanTok AR 模型达到或超过了一个规模为其 $3.4\times$ 的 VideoFlexTok AR 模型,而更大的 SemanTok AR 模型进一步提高了保真度。它在分布外类别上保持语义对齐,并在每个噪声水平(包括纯噪声)下为解码器提供更高的语义对齐。它在重建和生成两方面都表现良好,并且其短标记前缀预测成本更低,并能带来更好的生成保真度,而像素细节则推迟到后续标记。
cs.AI / 95 / 2610.00848
Geometric Similarity in VLM Low-Level Vision Representations
VLM 低级视觉表征中的几何相似性
diffusion
扩散模型相关
Abstract
Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.
Chinese Translation
视觉-语言模型(VLM)已成为通用视觉主干网络的有力候选,其代表性架构包括自回归(AR)模型和扩散 Transformer(DiT)。然而,如何高效地将它们适配用于一体化低级图像复原仍是一项挑战。至关重要的是,该领域尚缺乏对以下问题的理解:VLM 如何组织隐藏层表征,以及这些结构上截然不同的范式是否共享一种用于像素级感知的共同几何组织方式。这种共享的组织方式是构建高度可迁移的统一复原 VLM 与适配器的先决条件。在本文中,我们系统地研究了跨越 5 个类别的 24 个低级任务之间的表征相似性。我们提出了 GeoSim,一个统一的四级框架,从全局相似性、局部几何、稀疏特征分解和拓扑验证的视角分析任务条件化表征。我们的公式化方法适用于在同任务与跨任务/模型设置下分析 AR 模型中的隐藏状态和 DiT 中的特征图。我们的结果揭示了低级视觉表征的组织原则,同时暴露了它们在跨任务与跨模型一致性方面的局限。最终,GeoSim 提供了一种可解释性视角,用于探究低级视觉中的潜在可迁移性,并诊断任务特定或模型特定场景中的模型局限。
cs.AI / 96 / 2610.00953
Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold
扩散 MLLM 中的两个时钟:当答案在理由展开之前稳定时
diffusion
扩散模型相关
Abstract
An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On V*Bench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron's overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/V*Bench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.
Chinese Translation
在掩码扩散 MLLM 中,一个答案候选可以在其理由仍在展开时就已经稳定。我们将已记录候选的回溯性稳定与 token 承诺区分开来,并相对于理由生成考察这两个时钟。分析我们在三个视觉问答基准上的结果,我们发现,在单块、抑制 EOS 的 LaViDa 运行中,稳定时理由侧画布仍有 89.4–98.1% 未被写入。在 V*Bench 上,将块长度从 128 减小到 8 会使这一比例从 89.4% 变为 1.7%,同时改变答案覆盖率和合格观察窗口。在启用 EOS 的提示下,直接指令使 Nemotron 在 M3CoT 和 ScienceQA 上的总体准确率分别提高 15.0 和 19.5 个百分点,但使 LaViDa/V*Bench 的准确率降低 11.0 个百分点。一种对称分解将每个变化中较大的绝对分量与覆盖率而非条件准确率联系起来。匹配画布的图像消融在测量答案稳定的同时测量视觉敏感性,从而分离这两个时间读出。总之,这些测量区分了答案稳定、理由展开和视觉敏感性,并确定覆盖率是提示差异中较大的组成部分。
cs.AI / 97 / 2610.01434
MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
MWOP:面向高效 MLLM 的模态感知宽度维操作剪枝
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.
Chinese Translation
多模态大语言模型(MLLMs)在处理长视觉-文本序列时会产生大量推理成本。尽管现有的操作压缩方法利用了模态级冗余,但它们在很大程度上将注意力头内部以及共享前馈网络(FFN)通道中的计算视为统一单元,从而使更细粒度的冗余未被充分探索。我们发现,冗余在同一注意力头内的不同模态交互路径之间以及同一 FFN 通道的视觉与文本执行之间都会发生变化。基于这些发现,我们提出模态感知宽度维操作剪枝(MWOP),它在每一层内独立地剪枝视觉到视觉(V2V)、文本到视觉(T2V)和文本到文本(T2T)注意力路径,并分别为视觉和文本输入选择 FFN 通道。一阶泰勒准则指导剪枝过程,并在注意力剪枝和基于 LoRA 的恢复训练之后重新评估 FFN 重要性。为了将所得的细粒度稀疏性转化为实际加速,我们进一步开发了路径稀疏的 Triton 注意力内核以及紧凑的视觉侧 FFN 执行。MWOP 在减少注意力和 FFN 计算的同时保留 token 序列,使其与 token 压缩互补,并能够同时减少序列长度和每个 token 的计算量。在 LLaVA-OneVision-7B 上,仅 MWOP 就实现了 $1.6\times$ 的预填充加速比,并在 12 个基准上达到 99.7\% 的平均性能保持率。与两种代表性的 token 压缩方法结合后,它进一步将其预填充加速比分别从 $2.0\times$ 和 $1.9\times$ 提升到 $2.9\times$ 和 $2.7\times$。在 Qwen2.5-VL-7B 上的结果进一步证明了其跨架构的适用性。代码可在 https://github.com/EIT-NLP/MWOP 获取。
cs.AI / 98 / 2610.01595
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
在它消退之前:在视频大语言模型推理时强化时间表示
large language model
大语言模型相关
Abstract
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $τ_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $τ_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
Chinese Translation
视频大语言模型(VideoLLMs)按顺序接收帧,并解释视觉内容如何沿时间轴演变,然而时间推理在各类架构中始终是一个持续存在的弱点。反转视频的帧顺序——这一本应逆转时间答案的变换——常常使最终预测保持不变。我们通过定义时间发散向量 $τ_l$ 来研究这一失败的来源,该向量是由反转时间顺序引起的逐层表示差异。追踪其在不同层上的幅度揭示出一致的时间发散分布,其中发散在中间层达到峰值,并在接近输出时逐渐减弱。我们确认这一峰值特定于时间推理,并且对预测具有功能上的关键作用,从而确立 VideoLLMs 在中间层获取时间信息,但未能将其保持到输出。这种渐进式消退促使我们提出方法——时间激活注入(Temporal Activation Injection, TAI),该方法为每个输入在分布峰值处提取 $τ_l$,并按照测得的衰减将其重新注入后续层。TAI 无需训练,并在三个 VideoLLMs 和四个基准上持续改进时间推理,同时对非时间任务的影响可忽略不计。代码可在 https://github.com/Youngwoo-git/Before-It-Fades 获取。
cs.LG / 99 / 2610.01670
Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation
MLLM 评判者会评判编辑吗?在具有已验证质量保持的图像编辑评估中审计偏差
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is verified to preserve the underlying editing quality. To address this challenge, we introduce EditJudgeBias, a counterfactual benchmark with verified quality preservation, comprising 1,196 real editing samples and 13 cues injected across four evaluation sites. We verify quality preservation for the requested edit using calibrated multimodal validators, controls, and human inspection. We then audit five MLLM judges along three complementary dimensions: invariance to quality-preserving cues, agreement with human judgments, and stability of pairwise preferences. Importantly, observed shifts are evaluated against each judge's own zero-dose and re-query noise floors rather than against zero. Experiments show that quality-preserving cues move every judge beyond its own noise. Fabricated majority opinions increase ratings, irrelevant visual elements cause larger shifts than whole-image manipulations, and swapping candidate order reverses up to 60.9% of pairwise decisions. Edit-region cues also tend to reduce human agreement. The three measures characterize judges differently, showing that robustness cannot be captured by a single metric.
Chinese Translation
多模态大语言模型(MLLMs)正越来越多地被用作基于指令的图像编辑的自动评判者,以及模型训练的奖励信号。然而,系统地审计这些评判者是否受到与编辑质量无关的线索的影响具有挑战性,因为视觉干预本身可能会改变正在被评估的质量。因此,只有当干预被验证为保持底层编辑质量时,判断偏移才能归因于偏差。为应对这一挑战,我们引入 EditJudgeBias,一个具有已验证质量保持的反事实基准,包含 1,196 个真实编辑样本和跨四个评估位点注入的 13 个线索。我们使用校准的多模态验证器、对照和人工检查来验证所请求编辑的质量保持。然后,我们沿三个互补维度审计五个 MLLM 评判者:对质量保持线索的不变性、与人类判断的一致性,以及成对偏好的稳定性。重要的是,观察到的偏移是相对于每个评判者自身的零剂量和重查询噪声下限来评估的,而不是相对于零。实验表明,质量保持线索使每个评判者都超出其自身噪声范围。伪造的多数意见会提高评分,无关视觉元素造成的偏移比整图操纵更大,并且交换候选顺序会逆转高达 60.9% 的成对决策。编辑区域线索也往往会降低与人类判断的一致性。这三种度量以不同方式刻画评判者,表明稳健性无法由单一指标捕捉。
cs.AI / 100 / 2610.02010
Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking
探索生成式图像水印在面对潜在频率掩蔽时的弱点
diffusion
扩散模型相关
Abstract
Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion regeneration for improved image preservation. We provide a theoretical distortion bound relating the change between the reconstructed adversarial image and the masked latent-frequency perturbation. We evaluate the proposed attack against six diffusion watermarking methods on images generated from DiffusionDB and MS-COCO prompts. Latent Frequency Masking removes or substantially weakens several watermarks while preserving perceptual quality and achieving favorable runtime compared with existing attacks. These results identify latent-frequency manipulation as a practical attack surface and highlight the need to include such attacks in robustness evaluations of generative image watermarking.
Chinese Translation
不可见水印已成为追踪AI生成图像的核心工具,但其针对自适应去除攻击的鲁棒性仍是一个未解决的安全问题。我们提出了潜在频率掩蔽(Latent Frequency Masking),这是一种通过替换带水印图像潜在表示中的选定傅里叶系数来抹除水印证据的攻击。这些替换值可以从高斯噪声中采样以提高效率,也可以由扩散再生得到以改进图像保持。我们提供了一个理论失真界,将重建的对抗图像与经掩蔽的潜在频率扰动之间的变化联系起来。我们在由DiffusionDB和MS-COCO提示生成的图像上,针对六种扩散水印方法评估了所提出的攻击。潜在频率掩蔽移除或显著削弱了若干水印,同时保持感知质量,并与现有攻击相比实现了有利的运行时间。这些结果将潜在频率操纵识别为一个实际的攻击面,并强调了在生成式图像水印的鲁棒性评估中纳入此类攻击的必要性。
cs.MA / 101 / 2610.02045
Form and Void: Entangled Composition through an Autonomous AI Agent
形与虚:通过自主 AI 智能体的纠缠式构图
large language model
大语言模型相关
Abstract
Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.
Chinese Translation
正负空间是视觉构图中的一项基本原理,支撑着视觉上连贯的形式和分层的语义关系。生成此类构图具有挑战性,因为它需要对两个共享共同边界的语义概念进行协调控制。尽管近期的文本到图像模型和多模态大型语言模型(MLLMs)在图像生成和视觉理解方面取得了强劲性能,但正负空间生成仍然困难,尤其是在直接单次提示下。在这项工作中,我们提出了 Form and Void Agent(FaV-A),一个为分阶段正负空间生成而设计的多模态智能体。FaV-A 遵循渐进式工作流:它首先生成一个基础对象,然后分析其形状和空间结构以识别候选的负空间语义,最后为最终图像生成阶段生成构图指令。实验结果和消融分析表明,与直接的零样本 MLLM 基线相比,FaV-A 为生成视觉上连贯且语义对齐的正负空间构图提供了一个更有效的框架。
cs.AI / 102 / 2610.02117
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Where-OPD:基于合成场景的空间引导式多模态大语言模型在线策略自蒸馏
large language model
大语言模型相关
Abstract
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD
Chinese Translation
在线策略自蒸馏近来已成为一种提升语言模型推理能力的有效方法,其做法是用一个接收特权信息的、冻结的或 EMA 版本的自身模型来监督学生模型。然而,它在多模态大语言模型(MLLMs)上的应用在很大程度上仍未被探索。近期的方法使用特权视觉信息,例如与问题相对应的图像裁剪区域,来提升细粒度感知能力,但其收益仅限于那些能从这类视觉放大中获益的任务,并且需要人工标注的定位数据或外部教师模型。我们为 MLLMs 引入了一种不同形式的在线策略自蒸馏,它向教师模型提供文本化的、空间定位的引导,用以识别与查询相关的视觉元素。我们使用程序化生成的场景,其物体身份和空间坐标可自动获得,从而实现可扩展且无需标注的后训练。教师模型利用这种空间引导来定位并整合来自多个相关图像区域的证据,而学生模型则学习仅从图像和问题出发复现由此产生的行为。我们的方法在多个模型上一致地提升了在计数、文档与图表理解基准上的性能。重要的是,尽管后训练仅使用合成场景,所产生的改进仍能迁移到真实世界感知基准上,在 CVBench、V*、ZoomBench、BLINK、HR-Bench 和 MME-RealWorld 上平均性能提升了 3.23 个点。这些结果表明,空间定位的特权信息能够通过在线策略自蒸馏诱发更广泛的感知能力,实现超出后训练所用任务与数据分布的显著的合成到真实迁移。项目主页:https://github.com/sirkosophia/Where-OPD
cs.AI / 103 / 2610.02136
MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI
MIRTO:一种用于脑MRI无监督异常分割的、以配准为门控、经多重宇宙检验的评估协议
diffusion
扩散模型相关
Abstract
Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model's voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.
Chinese Translation
脑MRI的无监督异常检测(UAD)方法依据单一分数进行排名,然而该分数依赖于一些很少被报告的抉择:每张异常图如何与参考标准对齐,阈值如何设定以及在哪些数据上设定,以及使用了何种假阳性预算、指标、聚合方式和病灶定义。我们提出MIRTO,一种将这些抉择显式化并度量其影响的评估协议。它用配准检查以及具有已知检验效力的无标签诊断,为每一次比较的几何关系设置门控;仅在验证数据上设定阈值,并报告在测试集上实际达成的假阳性体积;在15,552条可辩护的评估流水线上重复每一次比较;并附加带有多重性控制的配对受试者自助法区间。将MIRTO应用于四种在相同健康数据上训练、并在312名BraTS 2020受试者上测试的UAD方法后,结果显示:存储图与参考标准之间的轴序不匹配使一个扩散模型的体素AUROC从0.873降至0.583,而其切片级AUROC几乎不变。在每种指标内部,该方法至少解释了体素AUROC和AUPRC方差的0.95以及Dice方差的0.77,但在病灶灵敏度上仅解释了0.14,而在那里病灶定义和命中判据占主导地位。一个在验证阈值下显著的Dice优势,在相同的实际假阳性负担下消失了,而一个精确恒等式将其归因于阈值迁移。对REFLECT的潜在聚合进行的一项免训练改动,在同等负担下将Dice提高了0.052。九个假设依据明确的标准进行了检验;由于同一队列被用于开发该协议,所有推断均为探索性的。
cs.AI / 104 / 2610.02188
DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
DMAD:将分布匹配作为对抗蒸馏用于快速视觉生成
diffusion
扩散模型相关
Abstract
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.
Chinese Translation
分布匹配蒸馏(DMD)根据分别估计的目标分数与学生分数之间的差异来训练少步学生模型,因此它必须保持一个辅助扩散模型来拟合学生不断演化的分布,这会带来额外的内存和计算成本。我们提出 DMAD,即分布匹配作为对抗蒸馏,它将分布匹配重新表述为分类,并直接学习所需的对数密度比。共享骨干网络上的两个判别器头将真实数据与教师样本同学生样本区分开来,并且对其 logits 的线性损失训练学生,而无需辅助分数拟合。我们证明,在判别器最优时,这些损失通过将判别器 logits 与对数密度比联系起来的经典恒等式,恢复了 DMD 所依据的分布匹配梯度。我们进一步引入基于差距的重加权,它根据真实数据头在真实样本与教师样本之间的经验 logit 差距,在不同噪声水平上调整教师监督。DMAD 在 ImageNet-64x64 上以一步生成达到 1.04 的 Fréchet Inception Distance(FID),在 COCO-10K 上以四步 SDXL 达到 14.47,并在四步 Wan2.1-T2V-14B 上取得 85.15 的 VBench 总分,这些是在所比较的少步方法和多步教师中的最佳值。在 MiniMax-H3-33B 上,对于联合音视频生成,我们的四步学生模型取得了相对于 DMD2 的 79.1% 和相对于 rCM 的 84.6% 的总体人类偏好率,不包括平局。我们的代码、模型和演示可在 https://yzmblog.github.io/projects/DMAD 获取。
cs.LG / 105 / 2610.02203
Embedding Prediction Helps Image Generation
嵌入预测有助于图像生成
diffusion
扩散模型相关
Abstract
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Chinese Translation
在扩散 Transformer 中,类别标签或文本提示只被嵌入一次,并且相同的条件会在每个去噪步骤中被重复使用。我们探究预测的嵌入能否改为充当这一条件。下一嵌入预测自回归(Next-Embedding Predictive Autoregression, NEPA)训练一个 Transformer 来预测序列中的下一个连续嵌入。在生成过程中,干净图像跟在噪声图像之后,因此其嵌入是条件和噪声图像之后的下一个嵌入。我们训练一个 NEPA 模型,通过多嵌入预测(Multi-Embedding Prediction)一次性预测所有这些嵌入;而在嵌入条件生成(Embedding Conditioned Generation)中,一个 DiT 生成器以这些预测为条件,这些预测在每个去噪步骤中重新计算,因此条件信号会适应当前噪声状态。在类别条件 ImageNet $256\times256$ 上的实验研究了生成器的条件、多嵌入预测的设计以及两个模型的扩展。NEPA 模型为每个采样步骤增加第二个网络;借助它,并与 REPA 结合,我们的最终模型 NEPA-DiT-XL 在使用约 REPA 三分之一的训练计算量的情况下达到了 1.32 的 FID。
cs.SE / 106 / 2610.00863
Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming
正确性、收敛性与AI生成代码检测:对入门编程课程中学生与大语言模型代码的纵向研究
large language model
大语言模型相关
Abstract
Large language models can generate plausible solutions to programming assignments, making it tempting to detect their use by matching student code against a reference bank of generated solutions. Yet similar code can also arise when an assignment admits only a few natural implementations, which leaves open what a match actually shows. We investigate generated-reference matching using 29,970 student submissions from ten Python labs offered in 2021, 2023, and 2025, together with 90,000 solution attempts generated retrospectively by three frontier LLMs. We validate the generated solutions using hidden instructor tests, compare code with MOSS after excluding the starter code, and examine the exact abstract syntax tree (AST) forms of selected functions. The models usually produced correct solutions and, across most assignments, converged on similar implementations. Student submissions matched the generated references more often in later cohorts, including among submissions that passed every hidden test. On tightly specified functions, the models converged on a few exact abstract-syntax-tree forms, and the number of distinct student forms also declined across cohorts, whereas open-ended functions remained diverse in both sources. Most reported overlaps were short, making the minimum match length an important choice when reviewing students' code. Finally, we discuss how instructors can build a reference bank of generated solutions before releasing an assignment to identify tasks on which generated solutions converge, decide how much review a match warrants, and redesign tasks to elicit tests, reasoning, and intermediate work. These findings support tracking population-level changes in submitted code, while attributing AI use to an individual submission would require additional evidence about how it was produced, such as prompts, revisions, intermediate code, and student disclosures.
Chinese Translation
大语言模型能够为编程作业生成看似合理的解答,这使得人们很容易通过将学生代码与生成的参考解答库进行匹配来检测其使用。然而,当一项作业只允许少数几种自然的实现方式时,也可能出现相似的代码,这就使得匹配结果究竟说明了什么仍然悬而未决。我们利用2021年、2023年和2025年开设的十个Python实验课的29,970份学生提交,以及由三个前沿大语言模型回溯生成的90,000次解答尝试,来研究这种生成参考匹配方法。我们使用隐藏的教师测试来验证生成的解答,在排除起始代码后用MOSS比较代码,并考察所选函数的确切抽象语法树(AST)形式。这些模型通常能生成正确的解答,并且在大多数作业上收敛到了相似的实现。在后期的学生群体中,学生提交与生成参考的匹配更为频繁,包括在通过了每一项隐藏测试的提交中也是如此。在规范严格的函数上,模型收敛到了少数几种确切的抽象语法树形式,而不同的学生形式数量在各学生群体间也有所下降;相比之下,开放式函数在两类来源中都保持了多样性。大多数被报告的重叠都很短,这使得最小匹配长度成为审查学生代码时的一个重要选择。最后,我们讨论了教师如何能在发布作业之前构建一个生成解答的参考库,以识别哪些任务上生成的解答会收敛,判断一次匹配需要多少审查,并重新设计任务以引出测试、推理和中间工作。这些发现支持追踪所提交代码在群体层面的变化,而将AI使用归因于某一份个人提交,则需要关于其如何产生的额外证据,例如提示词、修改过程、中间代码以及学生本人的说明。
cs.MA / 107 / 2610.00889
HakiCC: LLM-Driven Multi-Agent Design and Optimization of Concurrency Control Protocols
HakiCC:LLM 驱动的并发控制协议多智能体设计与优化
large language model
大语言模型相关
Abstract
Large language models (LLMs) have recently been applied in systems research as a tool to reduce human-intensive engineering effort through cost-efficient automation. Decades of research have produced a rich landscape of concurrency control (CC) protocols, each encoding distinct trade-offs in correctness, throughput, and abort behavior. However, most applications in practice default to 2PL or OCC, because selecting and adapting a protocol to a specific application requires expert knowledge that is rarely available to application designers. This is a wasted opportunity, as an application-specific CC protocol can yield significant performance advantages over a generic baseline, but designing one requires deep expertise in CC protocol design. In this paper, we propose HakiCC, an LLM-driven multi-agent pipeline that automatically designs, verifies, and optimizes concurrency control protocols tailored to a given target application. HakiCC provides a two-stage pipeline. In Stage 1, a multi-agent system takes a workload description as input and generates an application-specific CC protocol implementation, which is iteratively repaired and verified for conflict-serializability. In Stage 2, the verified protocol is further optimized for that application through an LLM-driven evolutionary loop targeting correctness and throughput. We evaluate HakiCC on TPC-C and AuctionMark as target workloads, producing and reporting ten application-specific CC protocols. All ten are conflict-serializable after Stage 1; Stage 2 improves throughput for every protocol, with average gains of +50.6% for TPC-C protocols and +92.2% for AuctionMark protocols.
Chinese Translation
大语言模型(LLM)近来已被应用于系统研究中,作为一种通过具有成本效益的自动化来减少人力密集型工程投入的工具。数十年的研究已经产生了丰富多样的并发控制(CC)协议,每种协议都在正确性、吞吐量和中止行为方面编码了不同的权衡。然而,实践中的大多数应用默认采用 2PL 或 OCC,因为为特定应用选择和适配协议需要专家知识,而应用设计者往往不具备这类知识。这是一种被浪费的机会,因为针对特定应用的 CC 协议相较通用基线可以带来显著的性能优势,但设计这样一个协议需要 CC 协议设计方面的深厚专业知识。在本文中,我们提出 HakiCC,一个由 LLM 驱动的多智能体流水线,它能够自动设计、验证并优化针对给定目标应用量身定制的并发控制协议。HakiCC 提供了一个两阶段流水线。在第一阶段,一个多智能体系统以工作负载描述作为输入,生成针对特定应用的 CC 协议实现,该实现会被迭代修复并针对冲突可串行化进行验证。在第二阶段,经过验证的协议会通过一个以正确性和吞吐量为目标、由 LLM 驱动的进化循环,针对该应用进一步优化。我们在 TPC-C 和 AuctionMark 上作为目标工作负载对 HakiCC 进行评估,生成并报告了十个针对特定应用的 CC 协议。全部十个协议在第一阶段之后都是冲突可串行化的;第二阶段提升了每一个协议的吞吐量,其中 TPC-C 协议的平均提升为 +50.6%,AuctionMark 协议的平均提升为 +92.2%。
cs.LG / 108 / 2610.00907
STEER: Reducing Inference Cost in Relational Foundation Models through Semantically Informed Sampling
STEER:通过语义信息引导的采样降低关系基础模型中的推理成本
large language model
大语言模型相关
Abstract
Relational foundation models (RFMs) are pretrained once on a collection of relational databases and prediction tasks, and then applied zero-shot to previously unseen databases and tasks. To make a prediction for a target row, an RFM samples a neighborhood of rows linked to that row through foreign keys and uses this neighborhood as its inference context. Lowering inference cost is an important goal for any foundation model, and for RFMs this cost grows with the size of the context. The simplest ways to shrink the context is to drop some of the sampled rows, but this ignores the semantics of the database schema, so it is as likely to discard informative rows as uninformative ones. We propose STEER, a sampling approach that shrinks the inference context by concentrating it on the tables most relevant to the prediction task at hand. STEER obtains relevance information by prompting a large language model to rank the foreign-key edges of the database schema into relevance tiers for the given task, and then maps each tier to a probability of following that edge during traversal. Because the ranking uses only the schema, it is computed once per task and reused across all subsequent predictions, amortizing its cost. We evaluate STEER on three state-of-the-art RFMs (RT, RT-J, and Griffin) and show that it reduces inference context size by about 40% on average while maintaining, and in some cases improving, prediction accuracy.
Chinese Translation
关系基础模型(RFMs)在一组关系数据库和预测任务上预训练一次,然后零样本应用于先前未见过的数据库和任务。为了对目标行进行预测,RFM 采样通过外键与该行相连的行邻域,并将该邻域用作其推理上下文。降低推理成本是任何基础模型的重要目标,而对于 RFM 来说,这一成本随着上下文大小而增长。缩小上下文的最简单方法是丢弃一些采样行,但这忽略了数据库模式的语义,因此丢弃信息性行与丢弃无信息性行的可能性一样大。我们提出 STEER,一种采样方法,它通过将推理上下文集中到与当前预测任务最相关的表上来缩小推理上下文。STEER 通过提示大型语言模型针对给定任务将数据库模式的外键边排序为相关性层级来获取相关性信息,然后将每个层级映射为在遍历过程中沿该边前进的概率。由于该排序仅使用模式,它每个任务只计算一次,并在所有后续预测中重复使用,从而摊薄其成本。我们在三个最先进的 RFM(RT、RT-J 和 Griffin)上评估 STEER,并表明它在保持、并在某些情况下提高预测准确率的同时,将推理上下文大小平均减少约 40%。
cs.LG / 109 / 2610.00687
Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware
Leto:在幸存硬件上实现LLM训练的快速原地恢复
large language model
大语言模型相关
Abstract
Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training. We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two-tier erasure protection and chunk-level transactional updates to keep the retained model state recoverable and consistent, and reclaim the shadow state when active training needs its GPU memory. Evaluation on 6- and 72-GPU NVIDIA A100 clusters shows that Leto recovers 3.6--6.5$\times$ faster than the best-performing checkpointing baselines and improves productive training time by up to 13.7 percentage points. Large-scale simulation shows over 95% productive training time on a 131,072-GPU cluster.
Chinese Translation
硬件可操作故障(HOFs)会中断大语言模型(LLM)训练,但允许在同一硬件上无需重置、修复或更换即可恢复。尽管如此,现有恢复系统会重新加载检查点、重新计算丢失的进度并重建进程状态,使本可继续训练的GPU闲置。我们提出Leto,一种利用幸存硬件来实现高效原地恢复的容错训练系统。我们的关键洞见是,恢复训练所需的状态可以在活动训练进程之外被保留或准备,同时仍位于同一硬件上。Leto保留工作模型状态和可复用的进程状态,并在影子训练器中预初始化其余状态。我们设计了两级擦除保护和块级事务更新,以保持所保留的模型状态可恢复且一致,并在活动训练需要其GPU内存时回收影子状态。在6块GPU和72块GPU的NVIDIA A100集群上的评估表明,Leto比表现最佳的检查点基线恢复快3.6--6.5$\times$,并将有效训练时间最多提高13.7个百分点。大规模仿真表明,在131,072块GPU的集群上,有效训练时间超过95%。
cs.LG / 110 / 2610.00465
AIR-LLM: Broadcasting AI Weights over Radio for Memory-Free Edge LLM Inference via RF Computing
AIR-LLM:通过无线电广播 AI 权重,以经由 RF 计算实现免内存边缘 LLM 推理
large language model
大语言模型相关
Abstract
Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spend unaffordable energy on loading the weights. This raises our question: can an edge device run an LLM without storing or loading its weights, but receive them over the air and consume them on the fly? Inspired by wireless broadcasting, we present AIR-LLM, an LLM inference architecture for edge devices, which is composed of: (i) a central radio (e.g., 5G base stations) that broadcasts the LLM weights into the air, and (ii) the edge user that receives the weights and completes the general matrix-vector multiplication (GEMV) of LLM inference directly in the radio frequency (RF) domain using RF mixers. To further shorten the airtime, AIR-LLM exploits MIMO spatial multiplexing and proposes an energy-efficient precoder-postcoder pair on the edge to calibrate its own wireless channel. Since the central radio stays user-unaware, AIR-LLM is user-scalable so that one broadcast serves unlimited users within its coverage. We implement AIR-LLM on the NVIDIA Sionna ray-traced channels of two real-world urban scenes and the profiling of a real RF mixer. With a WikiText-2 perplexity degradation of 4.0% on LLaMA-3.1-8B, AIR-LLM saves the energy by 157.7x/40.4x against the FP16 and weight-only quantization baselines; with 20 users, its airtime is 104.1x/26.0x shorter, respectively.
Chinese Translation
下一代大型语言模型(LLMs)正从云端扩展到无处不在的边缘设备。然而,边缘设备通常要么缺乏存储日益增大的 LLM 权重所需的内存,要么即使有足够内存,也会在加载权重上花费难以承受的能量。这引出了我们的问题:边缘设备能否在不存储或加载其权重的情况下运行 LLM,而是通过空口接收权重并即时使用它们?受无线广播启发,我们提出 AIR-LLM,一种面向边缘设备的 LLM 推理架构,它由以下部分组成:(i) 一个中心无线电(例如 5G 基站),其将 LLM 权重广播到空中,以及 (ii) 边缘用户,其接收这些权重并使用射频混频器直接在射频(RF)域中完成 LLM 推理的通用矩阵-向量乘法(GEMV)。为了进一步缩短空口时间,AIR-LLM 利用 MIMO 空间复用,并在边缘提出一个节能的预编码器-后编码器对,以校准其自身的无线信道。由于中心无线电保持对用户无感知,AIR-LLM 具有用户可扩展性,因此一次广播可服务其覆盖范围内的无限用户。我们在两个真实城市场景的 NVIDIA Sionna 光线追踪信道以及真实射频混频器的剖析上实现了 AIR-LLM。在 LLaMA-3.1-8B 上具有 4.0% 的 WikiText-2 困惑度退化的情况下,AIR-LLM 相较于 FP16 和仅权重量化基线分别节省 157.7 倍/40.4 倍能量;在 20 个用户下,其空口时间分别缩短 104.1 倍/26.0 倍。
cs.AI / 111 / 2610.00571
Interpreting Reasoning of Large Language Models via Partial Information Decomposition
通过部分信息分解解释大语言模型的推理
large language model
大语言模型相关
Abstract
Large reasoning models (LRMs) have achieved substantial improvements in solving complex mathematical problems, but often produce lengthy, repetitive, or erroneous reasoning trajectories. In this work, we introduce a new interpretability framework, SLIDER, to evaluate the quality of the reasoning process. SLIDER leverages an emerging body of work from information theory called Partial Information Decomposition to disentangle the information about the final answer between two consecutive reasoning steps into non-negative components: unique information (in preceding steps or current step), redundant information, and synergistic information. Building on this decomposition, we propose the *Step-wise Repetitive Reasoning Index (Step-RRI)*, a theoretically grounded measure that assesses whether the answer-relevant information in the current step $S_i$ is predominantly redundant with the past steps $S_{<i}$, relative to its unique and synergistic contributions. To evaluate the effectiveness of Step-RRI in detecting repetitiveness, we apply SLIDER to the redundancy class of the PRMBench dataset where Step-RRI improves step-level redundancy identification accuracy by over $10$ points compared to embedding-similarity and information-gain baselines. Next, we define *Trajectory-RRI*, an aggregate measure of repetitiveness for an individual reasoning trajectory. To demonstrate its practical relevance, we show that average Trajectory-RRI strongly correlates with actual reasoning length across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and GPT-4.1, motivating its use as a signal for improving reasoning efficiency. Finally, we introduce *Trajectory-RRI-guided data selection for fine-tuning*, demonstrating that selecting training data based on Trajectory-RRI can improve a fine-tuned model's reasoning efficiency while largely preserving its task performance.
Chinese Translation
大型推理模型(LRMs)在解决复杂数学问题方面取得了显著改进,但常常产生冗长、重复或错误的推理轨迹。在这项工作中,我们引入一个新的可解释性框架 SLIDER,用于评估推理过程的质量。SLIDER 利用信息论中一个新兴的研究方向——部分信息分解(Partial Information Decomposition)——将两个连续推理步骤之间关于最终答案的信息分解为非负分量:唯一信息(位于前序步骤或当前步骤中)、冗余信息和协同信息。在此分解的基础上,我们提出 *逐步重复推理指数(Step-wise Repetitive Reasoning Index, Step-RRI)*,这是一种具有理论依据的度量,用于评估当前步骤 $S_i$ 中与答案相关的信息相对于其唯一和协同贡献而言,是否主要与前序步骤 $S_{<i}$ 冗余。为评估 Step-RRI 在检测重复性方面的有效性,我们将 SLIDER 应用于 PRMBench 数据集的冗余类别,其中相比嵌入相似性和信息增益基线,Step-RRI 将步骤级冗余识别准确率提高了超过 $10$ 个点。接下来,我们定义 *Trajectory-RRI*,一种针对单个推理轨迹的重复性聚合度量。为了展示其实际相关性,我们表明,平均 Trajectory-RRI 与 QwQ-32B、DeepSeek-R1-Distill-Qwen-32B 和 GPT-4.1 上的实际推理长度强相关,从而推动将其用作改进推理效率的信号。最后,我们引入 *Trajectory-RRI 引导的微调数据选择*,证明基于 Trajectory-RRI 选择训练数据可以提高微调模型的推理效率,同时大体上保持其任务性能。
cs.LG / 112 / 2610.00499
Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving
去噪曲面:为扩散大语言模型服务建模与预测推理成本
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.
Chinese Translation
随着扩散大语言模型(dLLMs)能力不断增强,它们正从研究环境走向真实世界的\textit{服务}部署,而在服务场景中,请求管理(如调度与资源分配)依赖于对每个请求推理成本的准确估计。然而,常见的成本代理指标对 dLLMs 而言并不适用:输出长度忽略了单次前向传播可以同时解掩多个 token,而去噪步数则忽略了各步之间\textit{异构}的代价。我们观察到,块自回归生成机制在输出块与块内去噪步这两个维度上诱导出一个二维执行结构,而这些代理指标却将其压缩为一个标量,丢弃了刻画成本所必需的信息。受此洞见启发,我们提出了去噪工作负载曲面(Denoising Workload Surface, DWS),它将该二维的块—步结构保留为一个概率曲面,用以对各步之间异构的代价进行加权。随后,我们设计了一种由粗到细的训练方案,使一个仅依赖提示词的轻量级预测器能够准确预测复杂的 DWS。该预测器即使在单个 CPU 核心上也能高效运行,从而避免与服务模型争抢 GPU。由于 DWS 将依赖于请求的执行行为与特定部署的成本因素解耦,该预测器无需重新训练即可跨硬件配置迁移。在\textit{真实世界}的服务实验中,相较于基于标量的预测器,DWS 将成本预测误差最多降低 $2.50\times$;同时,由 DWS 指导的最短作业优先调度器将在线聊天机器人的端到端延迟最多降低 $1.92\times$。
cs.LG / 113 / 2610.00574
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
让稀疏奖励发挥作用:面向多奖励强化学习的密度感知奖励聚合
large language model
大语言模型相关
Abstract
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Chinese Translation
多奖励强化学习训练大语言模型同时满足多个行为目标。如 GDPO 中所用的逐奖励归一化保留了 rollout 组内特定奖励的相对信息,但不同目标仍可能表现出不均衡的学习进度。我们通过优势能量研究这一行为,即在单个批次上某一奖励的平方优势之和。在理想化的 GDPO 归一化下,我们证明该能量与活跃组密度成正比:即该奖励提供非零相对优势的 rollout 组所占比例。这揭示了一种残留的批次级信号不平衡,并为校准各奖励的贡献提供了依据。基于这一关系,我们提出密度感知奖励聚合(Density-Aware Reward Aggregation, DARA)。我们推导出一种逆平方根密度校正,该校正赋予来自不常活跃奖励的信号更大的权重。DARA 从每个 rollout 批次计算其权重,在不修改底层策略优化目标的情况下适应训练过程中奖励活跃度的变化。在工具调用和数学推理上的实验表明,DARA 比 GDPO 更快学习到目标行为:在工具调用上,以最多少 26% 的训练步数达到高格式合规率;在数学推理上,以最多少 65% 的步数达到接近饱和的长度合规率,同时在最终性能上仍保持竞争力。我们的代码可在 https://github.com/zhaihaotian/DARA 获取。
cs.LG / 114 / 2610.00647
Group-Invariant Statistics Determine Embedding Geometry: Harmonic Analysis of Representations from Bach to the Night Sky
群不变统计决定嵌入几何:从巴赫到夜空的表示调和分析
large language model
大语言模型相关
Abstract
The representations that language models learn for concepts such as months, weekdays, and places display consistent geometric structure: circles and saddle-shaped "Pringle" manifolds. Recent work traced these structures to $\textit{translation symmetry}$ in word co-occurrence statistics, deriving the observed Fourier geometry when co-occurrence depends only on distance on an abelian lattice of concepts. We demonstrate that more general notions of symmetry lead to equally structured predictions. Considering symmetries defined by arbitrary finite groups, compact groups, and homogeneous spaces, we prove that whenever the co-occurrence statistics of a word family are invariant under a group $G$, the learned word embeddings consist of matrix elements of the irreducible representations (irreps) of $G$. Circles and Pringles arise when $G$ is cyclic, in which case the irreps are Fourier modes. We verify the irrep structure in three experimental settings. (i) The cyclic group $\mathbb{Z}_{12}$: for the months of the year we recover the known circular geometry. (ii) A dihedral group acting on the major and minor triads: we unify two classical observations -- that transposition and chord inversion form a group ($T/I$) acting on chords (music theory), which $\textit{implies}$ that the well-known "circle of fifths" emerges in learned chord embeddings (machine learning). (iii) We explain and reproduce a recently discovered spherical representation of celestial objects in large language models (LLMs) as a spherical-harmonic embedding derived from our theory. Our results demonstrate that the geometry of learned representations is often a consequence of the statistical symmetry of underlying data.
Chinese Translation
语言模型为月份、工作日和地点等概念学习到的表示展现出一致的几何结构:圆环和鞍形的“Pringle”流形。近期工作将这些结构追溯到词共现统计中的 $\textit{translation symmetry}$,并在共现仅依赖于概念的阿贝尔格点上的距离时,推导出所观察到的傅里叶几何。我们表明,更一般的对称性概念会导向同样具有结构的预测。考虑由任意有限群、紧群和齐性空间定义的对称性,我们证明:只要一个词族的共现统计在群 $G$ 下不变,学习到的词嵌入就由 $G$ 的不可约表示(irreps)的矩阵元构成。当 $G$ 为循环群时,就会出现圆环和 Pringle,此时不可约表示是傅里叶模式。我们在三个实验设置中验证了不可约表示结构。(i) 循环群 $\mathbb{Z}_{12}$:对于一年中的月份,我们复现了已知的圆形几何。(ii) 一个作用于大三和弦与小三和弦的二面体群:我们统一了两个经典观察——移调与和弦转位构成一个作用于和弦的群 ($T/I$)(音乐理论),这 $\textit{implies}$ 著名的“五度圈”会在学习到的和弦嵌入中出现(机器学习)。(iii) 我们将最近在大型语言模型(LLMs)中发现的天体球面表示解释并复现为由我们的理论导出的球谐嵌入。我们的结果表明,学习到的表示的几何往往是底层数据统计对称性的结果。
cs.LG / 115 / 2610.00661
Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
探索更多,推理更佳:面向扩散语言模型的逐步风险敏感 GRPO
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.
Chinese Translation
扩散大语言模型(dLLMs)通过对一个序列或连续块进行去噪来生成文本,从而允许并行揭示多个词元。具有可验证奖励的强化学习(RLVR)在这些决策之间复用终端反馈,即使它们的条件上下文发生变化。我们提出逐步风险敏感 GRPO(StepRS-GRPO),它在去噪状态之间改变组优势变换的风险系数,同时保留底层训练器。对于二值奖励,我们表明该变换恰好是对中心化结果优势的依赖于提示和状态的重新缩放。基于能力的校准给出了系数尺度,而端点与插值消融引导了调度选择。在多个 dLLM 主干模型和数学推理基准上,StepRS-GRPO 相较于中心化 GRPO 同时提升了 pass@1 准确率和 pass@k 覆盖率,同时增加了答案多样性。在我们的消融研究中,质量匹配的对照支持了状态分配和调度方向的贡献,并且在将优势的均方根(RMS)匹配到中心化 GRPO 的 RMS 之后,这些增益仍然存在。推理轨迹诊断进一步表明,StepRS-GRPO 带来的多样性增益不仅限于最终答案字符串。
cs.LG / 116 / 2610.00704
SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization
SkillSpec:通过表示专门化实现的共识门控智能体技能演化
large language model
大语言模型相关
Abstract
Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid representation.Across six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
Chinese Translation
自然语言技能是文本程序性记忆,通过它们,大语言模型(LLM)智能体在不更新模型权重的情况下保留可复用的任务知识。现有方法通常将技能视为静态制品或使用聚合验证分数作为反馈进行优化的单体文档。然而,将技能表示为单体文档会将优化限制在其文本内容上,而没有显式建模程序性知识被检索和执行所依据的结构。我们识别出学习保留什么知识与确定如何组织该知识之间的一个关键区别:文本更新应首先通过执行证据进行验证,之后保留的知识应根据其程序性依赖关系和检索需求进行结构化。为此,我们提出SkillSpec,一个包含共识门控演化和表示专门化的两阶段框架。在共识门控阶段,互补的编辑意图生成完整的候选技能。仅当配对评估达成共识时才提交更新,这要求在每个重复评估中具有充分的总体改进和非负的聚合配对增益。在专门化阶段,从完整优化轨迹(包括被接受和被拒绝的候选)中推导出的过程敏感性和冗余敏感性信号,指导对平面、图或混合表示的选择。在六个基准和三个目标语言模型上,SkillSpec将平均成功率较SkillOpt提高了6.89%,这是在三个模型上平均的结果。这些结果表明,可靠的技能演化和表示专门化解决了互补的目标:决定保留什么知识以及如何为推理构建该知识。
cs.LG / 117 / 2610.00707
Initialization Improves LLM-Driven Discovery
初始化提升大语言模型驱动的发现
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.
Chinese Translation
大语言模型(LLM)已被用于算法、定理、药物及其他任务的新发现,其方式是使用一种引导框架(harness),提示 LLM 迭代地优化某个目标。在本工作中,我们研究先前迭代解的种群(population)与最终发现成功之间的关系。我们将以往关于引导框架设计的工作加以推广,开发出一套包含 12 种引导框架的套件,称为“Modular”,并刻画它们在 5 个多样化发现任务上的表现,发现发现成功是脆弱的,且对引导框架设计高度敏感。我们揭示了模式崩溃(mode collapse),其特征是迭代解多样性的急剧下降,作为一种常见的失败模式。我们发现,流行的最先进引导框架以及旨在延缓这种崩溃的、诱导多样性的引导框架干预措施,所带来的收益并不一致。相反,我们的结果揭示,早期发现的性能对最终成功具有预测性。因此,我们提出一种普遍适用的干预方法,其执行一个并行探索的初始阶段,以初始化后续的迭代优化。我们的方法在许多引导框架和目标任务上都带来了一致的收益,从而证实了初始化在 LLM 驱动发现中的重要性。
cs.LG / 118 / 2610.00728
Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations
面向真实站点观测的天气数据同化生成模型基准测试
diffusion
扩散模型相关
Abstract
Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.
Chinese Translation
天气再分析产品依赖于计算密集型的数值天气预报,随后通过数据同化将预报向观测修正。深度生成模型提供了一种更廉价的替代方案,它将大部分成本从推理转移到离线训练。然而,现有的生成式方法是在合成观测上、或在不同的数据集与评估方案下进行评估的,这使得人们不清楚哪些设计选择真正改善了现实世界的数据同化。我们提出了首个在真实气象站点观测上对生成式天气数据同化进行的受控基准测试。使用美国本土的 11,849 个 NOAA MADIS 站点和四个天气变量,我们在固定数据集、观测算子和深度学习架构的情况下评估各种方法。该基准测试将主要的设计选择与经典的 3D-Var 基线进行比较,包括扩散模型与流匹配、像素空间与潜空间表述,以及多种推理时条件化策略。该基准测试揭示了三个明确的结论。第一,学习到的生成式先验优于 3D-Var 的高斯先验(相对 ERA5 的 RMSE 降低分别为 35.7% 与 33.3%),尽管在推理时不使用 ERA5 背景场。第二,全梯度引导始终优于停止梯度与初始噪声优化。第三,其他选择几乎没有可测量的收益:在匹配条件下,扩散与流匹配的表现几乎相同,而潜空间变量混合没有帮助。我们进一步评估了稠密和稀疏站点设置,发现生成式 AI 和全梯度引导的优势在稀疏条件下更为显著。总之,这些结果明确了生成式天气数据同化的哪些组件能够提升在真实站点观测上的性能,并为未来工作建立了一个标准化的基准。
cs.LG / 119 / 2610.00835
TrueMuse: A Benchmark for Data Attribution in Text-to-Music Models
TrueMuse:面向文本到音乐模型的数据归因基准
diffusion
扩散模型相关
Abstract
Text-to-music generation models are trained on massive music collections, creating a growing need for data attribution methods that can quantify the contribution of individual training samples. However, existing attribution methods are difficult to rigorously evaluate due to the lack of reliable ground truth, making it challenging to reliably assess their actual effectiveness. To address this gap, we introduce TrueMuse, a controlled dataset and benchmark for text-to-music data attribution. TrueMuse is constructed by fine-tuning three diffusion-based text-to-music models on carefully curated attribution samples, whose known inclusion in fine-tuning provides controlled attribution targets for evaluation. The benchmark covers four attribution settings, spanning melodic structure, timbral characteristics, artist-level stylistic signatures, and genre-level shared patterns, and includes 133 attributes, 648 fine-tuned models, and 95,456 generated samples across two prompt types. Using TrueMuse, we systematically evaluate existing black-box attribution methods along four dimensions: fine-tuning improvement, prompt-type difficulty, multi-task training, and fine-tuning data size. Our results show that attribution remains challenging, with existing methods exhibiting substantial variation across evaluation settings, highlighting the need for more reliable and generalizable attribution methods for text-to-music generation. Code and Dataset will be released upon acceptance.
Chinese Translation
文本到音乐生成模型是在海量音乐数据集上训练的,这使得人们日益需要能够量化单个训练样本贡献的数据归因方法。然而,由于缺乏可靠的基准真值,现有的归因方法难以被严格评估,这使得可靠地衡量其实际有效性变得十分困难。为弥补这一空白,我们提出了 TrueMuse,一个用于文本到音乐数据归因的可控数据集与基准。TrueMuse 通过在精心挑选的归因样本上微调三个基于扩散的文本到音乐模型而构建,这些样本在微调中的已知包含情况为评估提供了可控的归因目标。该基准涵盖四种归因设置,涉及旋律结构、音色特征、艺术家层面的风格签名以及流派层面的共享模式,并包含 133 个属性、648 个微调模型,以及跨两种提示类型的 95,456 个生成样本。借助 TrueMuse,我们从四个维度系统评估了现有的黑盒归因方法:微调带来的提升、提示类型的难度、多任务训练以及微调数据规模。我们的结果表明,归因仍然具有挑战性,现有方法在不同评估设置下表现出显著差异,凸显出为文本到音乐生成构建更可靠、更具泛化能力的归因方法的必要性。代码与数据集将在论文被接收后发布。
cs.LG / 120 / 2610.00838
SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
SHARPO:面向智能体强化学习的段级信用分配
large language model
大语言模型相关
Abstract
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
Chinese Translation
智能体强化学习(RL)训练大语言模型(LLM)在长时间、多步交互中采取行动。然而,单个局部错误就可能导致任务失败,而轨迹级奖励在为各个决策分配信用方面只能提供有限的指导。为解决这一局限,我们提出用于策略优化的段级事后优势重加权(Segment-level Hindsight Advantage Reweighting for Policy Optimization, SHARPO),这是一种在面向环境的段这一层级上改进组相对策略优化(Group Relative Policy Optimization, GRPO)的信用分配机制。受现有同策略自蒸馏(on-policy self-distillation, OPSD)方法的启发,SHARPO 计算每个段内教师-学生对数概率差,并利用所得信号计算一个作用于 GRPO 优势的有界乘数。该乘数由该段内的所有 token 共享,从而使信用能够随不同段而变化。在 Qwen2.5-7B-Instruct 上,SHARPO 在 ALFWorld 和 WebShop 基准上优于现有基线,包括 GRPO、SDAR、RLSD 和 StepOPSD。
cs.LG / 121 / 2610.00873
Rethinking Data Augmentation under Covariate Shift: Invariant-Guided Diffusion and Prototype Reweighting
重新思考协变量偏移下的数据增强:不变引导扩散与原型重加权
diffusion
扩散模型相关
Abstract
In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mimic outdated source distributions. This learning setting limits the stability of standard augmentation and adaptation pipelines. We generalize the task under such setting as the Augmented and Weighted Learning under Covariate Shift problem (AWL-CS). AWL-CS imposes two critical challenges on existing methods: 1) misleading generative guidance where models optimize for source similarity rather than downstream task relevance, and 2) structural instability of distributional density where reweighting mechanisms overfit to noisy validation signals. To tackle these challenges, we propose IGDPR (Invariant-Guided Diffusion with Prototype Reweighting), a unified framework that synergizes stable synthesis and structural adaptation: i) To achieve task-relevant generation, we steer the diffusion sampling process using invariant potentials to ensure synthetic samples align with stable decision boundaries rather than outdated correlations. ii) To ensure stable adaptation, we develop a prototype-based reweighting strategy that assesses sample reliability through structural clusters instead of isolated points, effectively filtering validation noise. Extensive experiments on real data demonstrate our method improves data quality by augmenting the most beneficial data for robust learning.
Chinese Translation
在许多工业应用中,1)表格数据稀缺且不平衡,因此需要合成扩充;2)输入分布在训练与部署之间发生漂移(协变量偏移);3)验证集往往偏离未见的测试环境;或4)标准生成模型只是模仿过时的源分布。这种学习设定限制了标准增强与自适应流程的稳定性。我们将这种设定下的任务泛化为协变量偏移下的增强与加权学习问题(AWL-CS)。AWL-CS 对现有方法提出了两个关键挑战:1)误导性的生成引导,其中模型优化的是源相似性而非下游任务相关性;2)分布密度的结构性不稳定,其中重加权机制过拟合于噪声验证信号。为解决这些挑战,我们提出 IGDPR(Invariant-Guided Diffusion with Prototype Reweighting,不变引导扩散与原型重加权),一个统一框架,将稳定合成与结构自适应协同起来:i)为实现任务相关的生成,我们使用不变势引导扩散采样过程,以确保合成样本与稳定决策边界对齐,而不是与过时的相关性对齐。ii)为确保稳定自适应,我们提出一种基于原型的重加权策略,通过结构簇而非孤立点评估样本可靠性,有效过滤验证噪声。在真实数据上的大量实验表明,我们的方法通过增强对鲁棒学习最有益的数据来提升数据质量。
cs.LG / 122 / 2610.00894
Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models
Clock Diffusion:高效半自回归连续扩散语言模型
diffusion
扩散模型相关
Abstract
Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.
Chinese Translation
最近关于离散数据连续扩散的研究工作已展现出与同类离散扩散模型相当的性能。然而,这些连续对应模型缺乏作为语言模型实际使用所必需的关键特性,即变长生成和对键值缓存的支持,并且它们仍落后于自回归和离散扩散质量的前沿。在这项工作中,我们解决了这些局限性。我们通过引入一种模型参数化来做到这一点,该参数化使用位置相关的噪声调度来定义半自回归(SAR)连续扩散语言模型(DLM)。结合高效的训练和采样算法,我们将该框架称为 Clock Diffusion,并给出我们方法的两个特例:块生成和滑动窗口生成。然后我们定义 ClockDLMs,一个基于滑动窗口 Clock Diffusion 的高斯 DLM 家族,在 OpenWebText 上达到了最先进的扩散似然界,甚至胜过性能优异的块 SAR 离散扩散模型。在 TinyGSM 上训练的 ClockDLMs 在 GSM8K 基准上也大幅优于连续基线,并匹配且超过可比的 SAR 离散扩散模型。最后,基于我们的参数化,我们提出了更高效的采样器,我们将其称为 Cache Grab,它借鉴了离散扩散中加速推理的技术,例如提交概率超过置信度阈值的词元以及自推测解码,进一步提升我们模型的质量和效率。
cs.LG / 123 / 2610.00895
Towards Fast and Disentangled Counterfactuals for Visual Foundation Models
面向视觉基础模型的快速且解耦的反事实
diffusion
扩散模型相关
Abstract
Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.
Chinese Translation
基础模型仍然容易受到虚假相关性和“Clever Hans”策略的影响。可解释机器学习可以在没有元数据的情况下为分类器发现并移除此类策略。对于基础模型,目前尚不存在这样的选项。我们提出解耦扩散自编码器(DiDAE)。DiDAE 将冻结的基础模型封装在一个条件扩散解码器中。一个反事实是沿着解耦字典的某个方向进行的一次闭式编辑,随后进行解码。该字典可以是有监督的(Procrustes)或无监督的(奇异值分解,稀疏自编码器)。不需要梯度,因此 DiDAE 比当前最优方法快最多 2000 倍。我们在六个数据集上进行评估,其中两个是合成数据集,四个是真实世界数据集。在针对其中三个数据集的一项由需求驱动的基准测试中,其反事实与当前最优方法相当或更好,并且它们通过反事实知识蒸馏(CFKD)修复下游分类器,在此方面胜过基于元数据的校正。同一套机制可以针对一个训练好的分类器对一个预训练字典进行排序。它返回分类器实际读取的少数几个方向,每个方向都由一个翻转决策的反事实进行因果验证,并沿着教师标记为虚假的那些方向修复分类器。该工作流在我们随论文一同发布的开源 Peal 库中是即插即用的。有了公开字典和预训练解码器,剩下的就只是对分类器进行廉价的线性蒸馏及其自身的微调。
cs.LG / 124 / 2610.01037
SLIM: Simplex-Lattice Interpolation Merging
SLIM:单纯形-格点插值合并
large language model
大语言模型相关
Abstract
Optimizing merging coefficients for large language models can require many costly benchmark evaluations. We propose \textbf{Simplex-Lattice Interpolation Merging (SLIM)}, which constructs a quadratic surrogate of aggregate performance on the coefficient simplex using a classical mixture design. Evaluations of individual experts and equal-weight pairs determine the surrogate with the minimum number of measurements needed to identify a general quadratic on this domain. SLIM then optimizes the surrogate without further target-metric evaluations. Experiments on two model architectures demonstrate accurate prediction of unseen multi-expert mixtures and competitive merge performance under limited evaluation budgets. Matched-budget comparisons show that structured evaluation points improve prediction fidelity over random designs, including those using regularized fitting.
Chinese Translation
优化大语言模型的合并系数可能需要大量昂贵的基准评估。我们提出单纯形-格点插值合并(SLIM),它使用经典混料设计在系数单纯形上构建聚合性能的二次代理模型。对单个专家和等权对的评估,以在该定义域上识别一般二次型所需的最少测量次数来确定该代理模型。SLIM 随后优化该代理模型,而无需进一步的目标指标评估。在两个模型架构上的实验表明,其能够准确预测未见过的多专家混合,并在有限评估预算下具有有竞争力的合并性能。匹配预算的比较表明,结构化评估点相比随机设计提高了预测保真度,包括那些使用正则化拟合的设计。
cs.LG / 125 / 2610.01223
Have an LLM Write Your Anomaly Detector: Autonomous Discovery of Compact, Interpretable Detectors for Time Series
让 LLM 为你编写异常检测器:自主发现用于时间序列的紧凑、可解释检测器
large language model
大语言模型相关
Abstract
Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpretability. We use a large language model not as the detector but as the author of one: an autonomous research loop in which the model repeatedly edits a single short NumPy program under a leakage-free objective, keeping the best-scoring detector it finds. The loop discovers two compact detectors, one for univariate and one for multivariate series, that describe short windows by their local spectral features and compare them with the training-region distribution through a covariance-aware distance. On the TSB-AD benchmark these detectors lead the field across metrics, ahead of the strongest classical, deep, and foundation-model baselines including Time-RCD, yet they train no network and use no GPU, and the multivariate detector is faster than every similarly performing baseline. LLM-driven program search is thus a practical route to accurate, efficient, and transparent detectors.
Chinese Translation
时间序列异常检测需要在预测准确性、计算效率和可解释性之间进行权衡。我们不把大语言模型用作检测器本身,而是把它用作检测器的作者:一个自主研究循环,其中模型在一个无泄漏的目标下反复编辑单个简短的 NumPy 程序,并保留它所找到的得分最高的检测器。该循环发现了两个紧凑的检测器,一个用于单变量序列,一个用于多变量序列,它们通过局部谱特征来描述短窗口,并通过一种协方差感知的距离将其与训练区域分布进行比较。在 TSB-AD 基准上,这些检测器在各项指标上领先于该领域,超越了包括 Time-RCD 在内的最强的经典、深度和基础模型基线,然而它们不训练任何网络,也不使用 GPU,并且多变量检测器比每一个性能相近的基线都更快。因此,由 LLM 驱动的程序搜索是一条通向准确、高效且透明的检测器的实用路径。
cs.LG / 126 / 2610.01318
Feature Selective Model Collapse in Diffusion Models: Total Replacement versus Fixed-Budget Training
扩散模型中的特征选择性模型崩溃:完全替换与固定预算训练
diffusion
扩散模型相关
Abstract
Model collapse arises when generative models are trained on synthetic data produced by earlier models. The phenomenon has attracted considerable attention because of its societal and technical implications. However, previous studies have reached seemingly contradictory conclusions: replacing real data with synthetic data causes collapse (Shumailov et al.), yet accumulating real data alongside synthetic data can prevent it. For diffusion models, we study an intermediate regime typical of finite-budget pipelines: all past datasets and the real data are kept, but each new model is trained on a fixed-size sample from this growing pool, so the real fraction vanishes without any data being removed. Experiments on a 2D spiral dataset as well as the image benchmarks (MNIST, Fashion-MNIST, and CIFAR-10) show that replacement protocol degrades dataset rapidly as in the literature, whereas the fixed budget degrades only partially, sparing some features. A linear-response model of the multi-generational parameter dynamics, analyzed by stochastic recursion, confirms that the two protocols differ: some features will be fragile and lost within a few generations for both protocols, while some will be robust and preserved over practically unbounded horizons under the fixed budget protocol.
Chinese Translation
当生成模型在由早期模型生成的合成数据上进行训练时,就会出现模型崩溃。由于其在社会和技术层面的影响,这一现象已引起了相当大的关注。然而,先前的研究得出了看似相互矛盾的结论:用合成数据替换真实数据会导致崩溃(Shumailov 等人),而在合成数据之外同时累积真实数据则可以防止崩溃。对于扩散模型,我们研究了一种有限预算流程中典型的中间情形:所有过去的数据集以及真实数据都被保留,但每个新模型都是在这个不断增长的池中抽取的固定规模样本上进行训练,因此真实数据的比例会消失,而并未移除任何数据。在二维螺旋数据集以及图像基准(MNIST、Fashion-MNIST 和 CIFAR-10)上的实验表明,替换协议会像文献中那样使数据集迅速退化,而固定预算则只会部分退化,从而保留某些特征。一个关于多代参数动力学的线性响应模型,通过随机递推进行分析,证实了这两种协议存在差异:对于两种协议而言,某些特征都会很脆弱,并在几代之内丢失;而在固定预算协议下,某些特征则会很稳健,并在实际上无界的时域内得以保留。
cs.LG / 127 / 2610.01494
Let the Heads Talk: Beyond Diagonal Graph Attention
让注意力头开口说话:超越对角图注意力
diffusion
扩散模型相关
Abstract
Sheaf Neural Networks generalize scalar-weighted message passing by replacing scalar edge weights with linear transport maps between local feature spaces. Yet the role of this matrix-valued transport is entangled with the broader sheaf-diffusion construction. We isolate the transport primitive through quiver representations and establish a direct connection with multi-head attention. Treating attention heads as coordinates of a local transport space reveals that standard multi-head attention implements diagonal edge maps: along each directed interaction, a source head can contribute only to the corresponding receiver head. Allowing off-diagonal entries instead enables edge-conditioned communication across heads before neighborhood aggregation. We show that this operation cannot, in general, be absorbed into a single shared linear map applied after aggregation. Building on this characterization, we introduce Topological Attention (Top-A), a multi-head attention that learns edge-dependent off-diagonal routes while preserving the original same-head paths and exactly recovering vanilla attention when the additional routing vanishes. We evaluate Top-A on relational reasoning, heterogeneous graph learning, and algorithmic reasoning, including out-of-distribution generalization, with heterophilic node classification as a contrast setting. The results show that cross-head transport is most useful when the task benefits from interaction-dependent transformations, while heterophily alone provides no systematic advantage. These findings identify edge-conditioned cross-head communication as a distinct computational primitive of matrix-valued transport.
Chinese Translation
层论神经网络(Sheaf Neural Networks)通过将标量边权替换为局部特征空间之间的线性传输映射,推广了标量加权消息传递。然而,这种矩阵值传输的作用与更广泛的层论扩散构造纠缠在一起。我们通过箭图表示分离出传输这一基本原语,并建立其与多头注意力的直接联系。将注意力头视为局部传输空间的坐标,可以揭示标准多头注意力实现的是对角边映射:沿着每个有向交互,源头只能贡献给对应的接收头。允许非对角项则能够在邻域聚合之前实现跨头的边条件通信。我们表明,一般而言,这一操作无法被吸收到聚合之后应用的单个共享线性映射中。基于这一刻画,我们提出拓扑注意力(Topological Attention,Top-A),一种多头注意力,它学习依赖于边的非对角路由,同时保留原始的同头路径,并在额外路由消失时精确退化为原始注意力。我们在关系推理、异质图学习和算法推理(包括分布外泛化)上评估 Top-A,并以异嗜性节点分类作为对照设置。结果表明,当任务受益于依赖交互的变换时,跨头传输最为有用,而仅凭异嗜性并不提供系统性优势。这些发现将边条件跨头通信确定为矩阵值传输的一种独特计算原语。
cs.LG / 128 / 2610.01502
Learned End-to-End Guidance Schedules for Diffusion Models
用于扩散模型的学习式端到端引导调度
diffusion
扩散模型相关
Abstract
Diffusion models are a powerful generative paradigm used across multimedia and scientific applications. Guided diffusion methods impose requirements on the generation by adding the gradient of a differentiable loss (the guidance function) as a drift term during inference. The weight of this drift (the guidance scale) is critical for the trade-off between data quality and requirement satisfaction. To achieve both of these goals, guided diffusion must resort to small guidance scales and lengthy sampling, incurring high computational costs. This work proposes learned end-to-end guidance schedules (LEEGS) to achieve these objectives with fewer sampling steps. LEEGS trains a time-dependent schedule by minimizing the guidance function over a small set of examples using stochastic gradient descent. Backpropagating through guided sampling is computationally expensive, so LEEGS uses an approximation of the gradient that cuts training time by a factor of 4. We evaluate LEEGS on diverse guidance tasks, including (a) image inpainting, (b) noisy image inverse problems, (c) face-ID-guided generation, and (d) forward and inverse PDE problems, outperforming baselines at equal budget (50 or 100 NFEs), or matching constant guidance with only 10% of the steps.
Chinese Translation
扩散模型是一种强大的生成范式,广泛应用于多媒体和科学应用中。引导扩散方法通过在推理期间将可微损失(引导函数)的梯度作为漂移项加入,对生成施加要求。该漂移项的权重(引导尺度)对于数据质量与要求满足之间的权衡至关重要。为了实现这两个目标,引导扩散必须诉诸较小的引导尺度和冗长的采样,从而导致高昂的计算成本。本工作提出了学习式端到端引导调度(LEEGS),以用更少的采样步骤实现这些目标。LEEGS 通过使用随机梯度下降在一小组样本上最小化引导函数,来训练一个时间相关的调度。通过引导采样进行反向传播在计算上代价高昂,因此 LEEGS 使用一种梯度近似,将训练时间缩短为原来的 1/4。我们在多样的引导任务上评估 LEEGS,包括 (a) 图像修复,(b) 含噪图像逆问题,(c) 人脸 ID 引导生成,以及 (d) 正向和逆向 PDE 问题,在相同预算(50 或 100 NFEs)下优于基线,或仅用 10% 的步骤就达到与恒定引导相当的效果。
cs.LG / 129 / 2610.01522
Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback
朗之万信息引导的迁移学习:用黑盒反馈替代目标样本
diffusion
扩散模型相关
Abstract
Many scientific and machine learning systems, from molecular dynamics to diffusion models and beyond, are governed by stochastic dynamics with low-dimensional structure, evolving on slow timescales. However, target trajectories, used to identify and interpret such dynamics, are often inaccessible: only biased or static samples that explore the underlying manifold are available. We introduce Langevin-Informed Transfer Learning (LITL), a framework for recovering target Langevin dynamics from biased source samples using only black-box feedback. LITL learns the leading spectral structure of the target infinitesimal generator and the projected drift through Dirichlet representation learning, enabling kinetic reconstruction in spectral form and slow-manifold gradient field estimation. We further introduce a spherical variant well suited to steering normalized latent representations commonly used in learning systems toward desired objectives. We establish finite-sample guarantees for eigenvalue, eigenfunction, and projected drift estimation in Sobolev norms, thereby ensuring generalization of these quantities and their first-order derivatives. Empirically, LITL recovers physical transition timescales from biased molecular simulations, builds kinetic structure from static samples of generative models, reconstructs spherical symmetries of physical systems, and enables post-hoc latent steering of trained neural networks under black-box feedback. Together, these results position spectral operator learning as a practical framework for recovering stochastic dynamics under distribution shift and unlock applications across machine learning and the physical sciences.
Chinese Translation
许多科学与机器学习系统,从分子动力学到扩散模型乃至更广泛的领域,都受具有低维结构的随机动力学支配,并在慢时间尺度上演化。然而,用于识别和解释此类动力学的目标轨迹往往难以获取:只有探索潜在流形的有偏样本或静态样本是可用的。我们提出朗之万信息引导的迁移学习(LITL),这是一个仅使用黑盒反馈、从有偏源样本中恢复目标朗之万动力学的框架。LITL 通过 Dirichlet 表示学习来学习目标无穷小生成元的主导谱结构和投影漂移,从而实现谱形式的动力学重构与慢流形梯度场估计。我们进一步引入一个球面变体,它非常适合将学习系统中常用的归一化潜在表示引导向期望的目标。我们建立了在 Sobolev 范数下对特征值、特征函数和投影漂移估计的有限样本保证,从而确保这些量及其一阶导数的泛化性。在实证上,LITL 从有偏分子模拟中恢复物理跃迁时间尺度,从生成模型的静态样本中构建动力学结构,重构物理系统的球面对称性,并在黑盒反馈下实现对已训练神经网络的事后潜在引导。综上所述,这些结果将谱算子学习定位为一种在分布偏移下恢复随机动力学的实用框架,并开启了横跨机器学习与物理科学的应用。
cs.LG / 130 / 2610.01537
FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization
FedFit:通过向量库参数化与量化的大语言模型联邦微调
large language model
大语言模型相关
Abstract
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
Chinese Translation
联邦学习(FL)能够实现大语言模型(LLMs)的隐私保护微调,然而巨大的通信开销仍然是一个关键瓶颈。此外,在 FL 中应用低秩适配(LoRA)面临着一种根本性的“聚合困境”,即在精确的积之和(Sum-of-Products, SoP)实现与通信高效的和之积(Product-of-Sums, PoS)实现之间左右为难。为了应对这些挑战,我们提出了 FedFit。首先,为了显著降低通信开销,我们引入了一种不相交的共享向量库参数化方法,它从两个紧凑且不相交的全局向量库中重建高维适配器矩阵。其次,为了解决聚合困境,我们设计了一种交替优化调度。通过在解耦的单库更新(其允许精确聚合)与由残差谱聚合机制校正的联合更新之间循环,我们解决了 SoP 与 PoS 之间的冲突。此外,我们将分块量化与客户端误差反馈相结合,以进一步压缩传输的向量。此外,我们为所提出的算法建立了理论收敛保证。在 Qwen2.5 模型上的大量实验表明,FedFit 达到了与标准联邦 LoRA 方法相当的困惑度性能,同时提供最多高出 100 倍的压缩比。
cs.LG / 131 / 2610.01548
Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Range-GRPO:通过奖励区间之间的成对关系进行策略优化
large language model
大语言模型相关
Abstract
As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr.GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
Chinese Translation
随着大语言模型(LLM)使用的不断扩大,后训练在使其适配下游任务方面已变得日益重要。然而,获取可靠的监督仍然代价高昂,尤其是在没有参考答案或可执行验证器的领域。LLM-as-a-Judge 为无标注回答提供了可扩展的伪奖励,但单一的分数并不能显式地表示奖励的不确定性。这促使我们将伪奖励表示为经保形校准的奖励区间。我们提出 Range-GRPO,一个半监督后训练框架,它结合了有限的有标注数据与无标注提示。在组相对策略优化(GRPO)中,学习信号依赖于每个 rollout 组内的相对奖励比较。所提出的目标函数以成对方式比较奖励区间,而不是将其归约为点奖励,从而使区间不确定性能够同时影响这些信号的幅度与方向。我们的理论分析刻画了这一区别,并表明当所有奖励区间都坍缩为点时,所提出的目标函数可恢复 Dr.GRPO 的优势。在实证上,Range-GRPO 在所评估的半监督方法中取得了最高的分布内与分布外平均性能,同时所需的训练资源更少。
cs.LG / 132 / 2610.01554
QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
QK-Wanda:耦合查询与键用于非结构化剪枝
large language model
大语言模型相关
Abstract
Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.
Chinese Translation
Wanda(Sun 等,2024)通过在每个线性投影内独立地对权重进行评分来剪枝大型语言模型,尽管查询和键通过点积相互作用。我们提出 QK-Wanda,它在未掩码的 pre-RoPE 重建目标下根据查询和键权重各自的删除代价对其评分。它用来自对侧投影的信息(对于查询权重使用键,对于键权重使用查询)增强 Wanda 分数,使两个投影共享一个剪枝预算。其闭式分数不需要梯度或权重更新;在我们主实验中使用的校准设置下,完整剪枝在 A100 上比 Wanda 多耗时 1.3%,在 H200 上多耗时 3.1%。我们在来自 TinyLlama、Llama 2、Llama 3 和 Qwen2.5 的 15 个模型上评估仅 QK 剪枝,这些模型覆盖 0.5B-72B 参数。相对于 Wanda,QK-Wanda 在 50% 稀疏度下平均将 QK 重建误差降低 60%,在 80% 稀疏度下降低 45%。下游增益取决于模型。在 Llama 2 70B 的 80% 稀疏度下,WikiText-2 和 C4 困惑度分别下降 20.3% 和 13.5%,而平均零样本准确率上升 5.94 个百分点。Qwen2.5-72B 也有所提升,但 Llama-3.1-70B 尽管重建误差更低,却具有显著更高的困惑度。这些结果既显示了耦合剪枝准则的前景,也显示了局部重建作为模型质量预测指标的局限。
cs.LG / 133 / 2610.01697
Learning PDE Dynamics between Submanifolds Using Green's Observation Operators
使用格林观测算子学习子流形之间的偏微分方程动力学
diffusion
扩散模型相关
Abstract
Many physical systems are driven and observed only on lower-dimensional submanifolds of a larger spatial domain, while their dynamics are governed by the ambient medium occupying that domain. Examples include laser-heated parts imaged by an infrared camera, and ground-level emissions measured on a sensor plane. Full-domain solvers, however, compute the entire volume for every new source although only the observation submanifold is needed, and black-box surrogates do not exploit that the ambient medium remains fixed. We introduce the \emph{Green's Observation Operator (GObO)}, which maps the ambient medium once to the Green's kernel of a linear PDE restricted to the source and observation submanifolds. New sources then cost one lower-dimensional integral and no network evaluation. Exponential rates in the kernel yield an exact finite streaming state with horizon-independent memory; we prove its stability and an approximation rate for the restricted heat kernel. On three-dimensional heat conduction and advection--diffusion with collocated and distinct source and observation geometries, GObO trained on static sources predicts responses to moving sources zero-shot with 4--8$\times$ lower error than black-box surrogates, at 1.4\,ms per query after a single conditioning pass. The same kernel transfers across resolutions and admits corrections for mild nonlinearities, including radiative losses and temperature-dependent conductivity, without retraining, at the cost of lower in-distribution accuracy.
Chinese Translation
许多物理系统仅在更大空间域的低维子流形上被驱动和观测,而其动力学由占据该域的周围介质支配。例子包括由红外相机成像的激光加热部件,以及在传感器平面上测量的地面排放。然而,全域求解器对每个新源都计算整个体积,尽管仅需要观测子流形;而黑箱代理模型没有利用周围介质保持固定这一事实。我们引入 \emph{格林观测算子 (GObO)},它将周围介质一次映射到限制在源子流形和观测子流形上的线性偏微分方程的格林核。新源随后只需一次低维积分,无需网络评估。核中的指数速率产生一个精确的有限流式状态,其内存与时间范围无关;我们证明了其稳定性以及受限热核的近似速率。在三维热传导和对流--扩散问题中,对于源和观测几何共置和分离的情况,GObO 在静态源上训练后,以零样本方式预测对移动源的响应,其误差比黑箱代理模型低 4--8$\times$,在单次条件化传递后每次查询耗时 1.4\,ms。同一个核可以跨分辨率迁移,并允许对轻度非线性进行校正,包括辐射损失和温度相关导热率,而无需重新训练,代价是分布内精度较低。
cs.LG / 134 / 2610.01739
Fixed-point neural samplers on discrete spaces
离散空间上的不动点神经采样器
diffusion
扩散模型相关
Abstract
Sampling from discrete, unnormalized distributions without access to data is a challenging problem. Neural samplers offer a promising approach by training generative models from density evaluations directly. Despite recent progress, existing discrete neural samplers are prone to mode collapse, come without convergence guarantees when trained via fixed-point iterations, and are often tied to a specific reference process such as masked or uniform diffusion. In this work, we introduce Discrete Gibbs Iterative Neural Sampler, a fixed-point neural sampler that addresses these limitations, enabling efficient, scalable learning, substantially reducing mode collapse in practice. Our framework builds on masked diffusion and also extends to transport between pairs of distributions. We demonstrate that the resulting method scales effectively to high-dimensional systems, supports amortized sampling across different conditions, and enables accurate estimation of alloy phase diagrams.
Chinese Translation
在无法访问数据的情况下,从离散的、未归一化的分布中进行采样是一个具有挑战性的问题。神经采样器通过直接从密度评估中训练生成模型,提供了一种有前景的方法。尽管最近取得了进展,现有的离散神经采样器容易出现模式崩溃,在通过不动点迭代进行训练时缺乏收敛保证,并且通常与特定的参考过程(如掩码扩散或均匀扩散)绑定。在这项工作中,我们引入了 Discrete Gibbs Iterative Neural Sampler,这是一种不动点神经采样器,它解决了这些局限性,实现了高效、可扩展的学习,并在实践中大幅减少了模式崩溃。我们的框架建立在掩码扩散之上,并且还扩展到分布对之间的传输。我们证明了所得到的方法能够有效地扩展到高维系统,支持跨不同条件的摊销采样,并能够准确估计合金相图。
cs.LG / 135 / 2610.01789
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
iADD:改进扩散策略优化中的对齐与多样性
diffusion
扩散模型相关
Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Chinese Translation
基于强化学习的扩散模型后训练,例如去噪扩散策略优化(DDPO),在奖励函数下优化一个反向扩散过程。然而,当前的奖励优化方法是以牺牲多样性和质量为代价来实现这一点的。在本文中,我们通过细致的理论考量和方法设计提供了更好的权衡。我们分析了理论框架,并从数学上证明,与先前工作中提出的结论相反,扩散模型的\emph{仅较后时间步}更新可能对多样性有害。此外,我们提出了一种基于坚实理论基础的增量式 Feynman-Kac 训练,以实现迄今为止最佳的对齐-多样性权衡。我们进行了大量实验,并在三个不同任务中将我们的方法与相关的扩散策略优化方法进行比较,同时还为每个组件提供了充分的消融实验,从而验证了在对齐和多样性两方面都具有显著的性能提升。
cs.LG / 136 / 2610.01799
SkillEvoLean: Mutation-enhanced skill evolution for Lean provers
SkillEvoLean:面向 Lean 证明器的变异增强技能进化
large language model
大语言模型相关
Abstract
Skill evolution offers a promising way to improve large language model agents without updating their parameters, but its use in formal theorem proving remains underexplored. Existing methods mainly target natural-language reasoning, improving skills by analyzing successful and failed trajectories and incrementally revising solving strategies. Although the Lean verifier provides reliable execution feedback, when all sampled trajectories fail, existing skill evolution methods lack successful trajectories from which to infer effective update directions. Furthermore, these methods also focus mainly on the root instruction file, thus underexploring the evolution of reference knowledge including mathematical concepts and proving techniques. To address these limitations, we propose a mutation-enhanced skill self-evolution framework for building skill-augmented Lean provers. The framework jointly evolves a high-level solving policy and its reference knowledge through progressive and mutation-based updates. Progressive evolution derives local improvements from successful and failed trajectories, while mutation is triggered when no complete proof can be generated, sampling mathematical concepts to produce and select new skill candidates under verifier feedback. We evaluate our method on MiniF2F, PutnamBench, the 2025 International Mathematical Olympiad (IMO 2025), and the 2026 USA Mathematical Olympiad (USAMO 2026). Under the same backbone model, trajectorysampling budget, and test-time compute, our method achieves proof success rates of 100.0%, 90.6%, 4/6, and 4/6, respectively, with GPT-5.5, outperforming the baseline methods. Further analysis shows that concept-guided mutation outperforms random-text-guided mutation by 6.9 and 8.2 percentage points on MiniF2F and PutnamBench, respectively, while solving one additional problem on both IMO 2025 and USAMO 2026.
Chinese Translation
技能进化提供了一种在不更新参数的情况下改进大语言模型智能体的有前景方法,但其在形式化定理证明中的应用仍未得到充分探索。现有方法主要面向自然语言推理,通过分析成功与失败的轨迹并逐步修正求解策略来改进技能。尽管 Lean 验证器提供了可靠的执行反馈,但当所有采样轨迹都失败时,现有的技能进化方法缺乏成功轨迹,无法据此推断有效的更新方向。此外,这些方法还主要关注根指令文件,因而对包括数学概念和证明技术在内的参考知识的演化探索不足。为解决这些局限,我们提出了一个变异增强的技能自演化框架,用于构建技能增强的 Lean 证明器。该框架通过渐进式和基于变异的更新,联合演化高层求解策略及其参考知识。渐进式演化从成功与失败的轨迹中得出局部改进,而当无法生成完整证明时,则触发变异,在验证器反馈下采样数学概念以生成并选择新的技能候选。我们在 MiniF2F、PutnamBench、2025 年国际数学奥林匹克(IMO 2025)和 2026 年美国数学奥林匹克(USAMO 2026)上评估我们的方法。在相同的主干模型、轨迹采样预算和测试时计算下,我们的方法在 GPT-5.5 上分别取得了 100.0%、90.6%、4/6 和 4/6 的证明成功率,优于基线方法。进一步分析表明,概念引导的变异在 MiniF2F 和 PutnamBench 上分别比随机文本引导的变异高出 6.9 和 8.2 个百分点,同时在 IMO 2025 和 USAMO 2026 上均多解决了一道题。
cs.LG / 137 / 2610.01815
Debias Anything: Fairness with Diversity without Supervision in Diffusion Models
去偏一切:扩散模型中无需监督的公平性与多样性
diffusion
扩散模型相关
Abstract
Although diffusion models produce high-quality images, they also reproduce and amplify demographic imbalances in their training data. Debiasing their generation process post-training w.r.t. some sensitive attribute usually relies on classifier guidance or explicit text extra-conditioning, but this reduces methods' applicability and output diversity. Conversely, methods promoting diversity alone do not ensure fair attribute representation. In this paper, we propose a method tackling fairness and diversity jointly that is generally applicable to any diffusion model and any sensitive attribute. To this end, an adapter connects the frozen diffusion model to a pretrained vision-language embedding space, enabling fairness and diversity guidance without sensitive-attribute annotations. For fairness, pairs of text prompts define attribute directions which guide batch composition towards specific proportions. For diversity, we introduce a score measuring disagreement between the semantic estimates derived from this representation. The formulation supports unconditional and text-conditional diffusion models, while requiring no prior knowledge or data of sensitive attribute. Experiments confirm that our method improves quality and diversity scores at comparable fairness levels.
Chinese Translation
尽管扩散模型能生成高质量图像,但它们也会再现并放大其训练数据中的人口统计学不平衡。在训练后针对某些敏感属性对其生成过程进行去偏,通常依赖于分类器引导或显式的文本额外条件化,但这会降低方法的适用性和输出多样性。相反,仅促进多样性的方法并不能确保公平的属性表示。在本文中,我们提出一种联合处理公平性与多样性的方法,它普遍适用于任何扩散模型和任何敏感属性。为此,一个适配器将冻结的扩散模型连接到预训练的视觉-语言嵌入空间,从而无需敏感属性标注即可实现公平性与多样性引导。对于公平性,成对的文本提示定义属性方向,引导批次组成趋向特定比例。对于多样性,我们引入一个分数,用于度量从该表示中得到的语义估计之间的不一致性。该形式化方法支持无条件扩散模型和文本条件扩散模型,同时不需要敏感属性的先验知识或数据。实验证实,在相当的公平性水平下,我们的方法提升了质量和多样性分数。
cs.LG / 138 / 2610.01821
Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
超越线性概念:发现并对齐大型语言模型中的非线性概念流形
large language model
大语言模型相关
Abstract
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.
Chinese Translation
理解大型语言模型(LLMs)中的信息处理,需要剖析其内部词元表示的几何组织。尽管现有的机制可解释性(MI)方法试图提取概念,但它们受限于一种强烈的线性假设,而非线性特征流形的证据对该假设提出了挑战。我们将来自计算机视觉的非线性多维概念发现(NLMCD)适配到词元级 LLM 激活,从而超越线性概念,将概念建模为低维流形。为了比较跨层和跨模型的概念流形,我们引入基于概念的对齐(CBA)分数,这是一种广义 Rand 指数,能够在没有显式特征匹配的情况下度量几何接近程度。我们的分析得出六项关键发现:(i) 相邻层健全性检查表明,CBA 比基于 PCA 或 CKA 的线性基线更敏感;(ii) 逐层对齐矩阵揭示了中间层和后期层中的两个块结构,这些结构在不同模型间一致,却被线性指标所掩盖;(iii) 概念组成在网络的大部分中仍以句法为主导,随后在后期层中让位于日益混合的句法-语义概念,并且越接近最终层,输出导向性越强;(iv) 英语和普通话之间的多语言概念共享依赖于训练而非普遍存在,在 Qwen 中最强,在 Llama 中较弱,而在 GPT-2 中不存在;(v) 模型间对齐反映了这一结构,不同规模的同族 Qwen 模型之间存在强对应关系,但跨模型家族的对齐较弱;以及 (vi) 在 Tulu-3 的训练阶段中,相邻阶段之间的对齐最高,基础模型与 SFT 之间的偏移最大,而后续的偏好对齐阶段(DPO、RLVR)基本不改变早期层,并且 RLVR 在后期层中大多保留了 DPO 的概念。
cs.LG / 139 / 2610.01882
Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
更快流动以协调:一步式在线多智能体流策略
diffusion
扩散模型相关
Abstract
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.
Chinese Translation
多智能体强化学习(MARL)提供了一个强大的框架,用于通过与环境的交互来学习协调行为。开发 MARL 策略需要在复杂且多模态动作分布的表达性建模与高效训练和执行之间取得平衡。生成式策略,尤其是基于扩散的策略,能够忠实捕捉复杂且多模态的行为,但昂贵的迭代采样阻碍了它们在线多智能体场景中的可扩展性。我们提出了一种通过一步式流模型(OMAF)实现的在线 MARL 框架,它将具有表达能力的生成式策略与高效的一步式动作生成相结合。OMAF 采用基于 Transformer 的流策略来捕捉复杂的协调行为,而其近似路径得分替代项则为同步流策略优化提供了一条有原则的途径。为了实现稳定且样本高效的学习,我们进一步开发了一种联合优化方案,将 softmax Q 值估计与联合流策略目标耦合起来,以进行协调策略学习。通过消除迭代采样,OMAF 在不牺牲策略表达能力的情况下显著降低了训练开销。在来自 MPE 和 MAMuJoCo 的 10 个标准任务上进行的大量实验表明,OMAF 始终取得更优的性能,与基线方法相比,回报最高提升至 3.4 倍,样本效率最高提升至 10.5 倍。这些结果验证了 OMAF 作为一种用于在线 MARL 的表达能力强且计算高效的一步式流策略范式的有效性。
cs.LG / 140 / 2610.01894
A foundation for systematic analysis of transformers and RNNs for tractography
用于纤维束成像的 Transformer 与 RNN 系统性分析的基础
diffusion
扩散模型相关
Abstract
Machine learning (ML) has emerged as a promising approach for improving diffusion MRI (dMRI) tractography, a task that remains limited by the intrinsic tension between local diffusion information and global anatomical plausibility. In this work, we systematically evaluate recurrent neural networks (RNNs) and Transformer models for iterative tractography, with particular attention to training strategies, input representations (including convolutional neural network (CNN)-based embeddings and end-of-sequence (EOS) tokens), and hyperparameter selection. We introduce a generation-validation phase enabling supervision at the streamline level during training, allowing supervision despite the mismatch between local loss functions and global streamline quality. Using the ISMRM2015 tractography challenge dataset, our models achieve the highest reported performance to date. Through controlled experiments, we quantify the impact of missing bundles, noisy or imperfect training streamlines, and invalid fibers in the training set. Finally, we demonstrate the applicability of our best-performing models for in vivo data from the Tractoinferno database. Overall, our results highlight both the potential and the limits of sequence-based deep learning models such as Transformers and RNNs for tractography, and emphasize the need for improved phantoms and evaluation methods for in vivo validation. We provide takeaways and recommendations for future researchers training and validating sequence-based supervised methods for tractography.
Chinese Translation
机器学习(ML)已成为改进扩散 MRI(dMRI)纤维束成像的一种有前景的方法,而该任务仍然受限于局部扩散信息与全局解剖合理性之间的内在张力。在本工作中,我们系统地评估了用于迭代式纤维束成像的循环神经网络(RNN)与 Transformer 模型,特别关注训练策略、输入表示(包括基于卷积神经网络(CNN)的嵌入和序列结束(EOS)标记)以及超参数选择。我们引入了一个生成-验证阶段,使得在训练期间能够在流线层面进行监督,从而在局部损失函数与全局流线质量不匹配的情况下仍可实现监督。使用 ISMRM2015 纤维束成像挑战数据集,我们的模型取得了迄今为止所报道的最高性能。通过受控实验,我们量化了训练集中缺失的纤维束、含噪或不完美的训练流线以及无效纤维所造成的影响。最后,我们展示了我们表现最佳的模型对来自 Tractoinferno 数据库的活体数据的适用性。总体而言,我们的结果既凸显了 Transformer 和 RNN 等基于序列的深度学习模型用于纤维束成像的潜力,也凸显了其局限,并强调了为活体验证改进体模与评估方法的必要性。我们为未来训练和验证基于序列的监督式纤维束成像方法的研究者提供了要点总结与建议。
cs.LG / 141 / 2610.01896
Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
异步 LLM 后训练:组质量封顶与收敛性分析
large language model
大语言模型相关
Abstract
Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(ε^{-4})$ to $O(ε^{-2})$ as $ε\to0$, where $1+ε$ is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as $G^{-2/5}$ after tuning the step size, where $G$ is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as $G\to\infty$, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.
Chinese Translation
异步强化学习(RL)提升了大语言模型后训练的效率,但会引入由更早策略生成的陈旧 rollout。关于这种陈旧性如何影响收敛以及如何减轻其影响的理论理解仍然有限。我们为 GRPO 风格算法推导了一个收敛界,该界明确刻画了梯度估计器的二阶矩与偏差之间的权衡。对于轨迹级重要性加权估计器,我们的分析表明,一旦二阶矩被一致控制,延迟就会通过由裁剪或重缩放引入的偏差进入该界。受这一洞见启发,我们提出了一种新颖的组质量封顶 GRPO(GMC-GRPO)方法,它在一类共享共同二阶矩保证的加权估计器中最小化基于比值的偏差界。我们为异步 GMC-GRPO 建立了收敛保证,并表明与 TIC-GRPO 相比,随着 $ε\to0$,它将四阶延迟项的阈值依赖从 $O(ε^{-4})$ 改进到 $O(ε^{-2})$,其中 $1+ε$ 是比值阈值。在局部策略重叠下,调整步长后,依赖延迟的项以 $G^{-2/5}$ 减小,其中 $G$ 是组大小。对于固定的行为策略和当前策略,由组重缩放引入的偏差也会随着 $G\to\infty$ 消失,而源自轨迹级裁剪的偏差可能持续存在。在 Qwen3 模型和推理基准上的实验表明,其对陈旧 rollout 的鲁棒性有所提升,其中 GMC-GRPO 在大的 rollout 延迟下在稳定基线中取得了最佳性能。
cs.LG / 142 / 2610.01934
Learning to Predict Distributions over Weight Updates for Test-Time Adaptation
学习预测权重更新的分布以实现测试时自适应
large language model
大语言模型相关
Abstract
Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.
Chinese Translation
超网络近来已在基于任务描述或额外示例等信号、在运行时动态自适应大语言模型(LLM)参数方面展现出成功。在此我们提出疑问:仅使用输入到大语言模型的查询,能够获得多少自适应信号?为回答这一问题,我们研究了用于 LoRA 估计的查询条件化超网络。此外,我们引入了分布型超网络,它不仅能产生参数适配器的点估计,还能产生可能 LoRA 上的分布。为此,我们提出了一种使用可微蒙特卡洛近似的简单端到端损失,并探索了多种分布参数化形式,包括回归与凸组合变体。结果表明,即使仅使用所学分布的均值,也能优于确定性超网络。关键在于,所学分布促成了一种不同形式的测试时扩展:我们不再仅通过在固定模型上采样更多词元序列来耗费额外算力,而是对权重更新进行采样,从而为同一查询产生多个已自适应的模型。随着考虑更多权重样本,性能得以提升,并且始终强于相应的词元采样自适应基线。最后,我们发现生成的更新可以跨查询迁移,这表明超网络学到了模型应如何自适应的可复用结构。综上,这些结果表明,查询条件化的权重更新分布既能支持自适应,也能支持测试时扩展。
cs.LG / 143 / 2610.01937
Graph Representation via Elements of Discrete Morse and Cobordism Theories
基于离散莫尔斯与配边理论要素的图表示
diffusion
扩散模型相关
Abstract
Topology is, by its nature and design, suited to structure that is nonlinear, multiscale, and nonstationary - however, within machine learning, its use remains largely confined to topological data analysis. We advocate that tools from low-dimensional topology which have remained almost exclusively contained within the domain of pure mathematics (such as Morse theory) offer a strong, complementary, and yet virtually unexplored perspective on the hidden structure of data-generating processes and learning tasks built upon them. Here we introduce concepts from cobordism theory and harness tools from discrete Morse theory to improve the performance of graph diffusion models through our pipeline MG-Diff. Further, we derive theoretical guarantees and sufficient conditions so that under a positive decision-gap, the Morse-theoretic tools and their application for induced diffusion guidance are stable under small perturbations. Finally, we illustrate the utility of discrete Morse theory in application to graph diffusion models for spatio-temporal graph forecasting and graph regeneration, and argue that these applications are only a small window into the part of what low-dimensional topology can offer to the field of machine learning.
Chinese Translation
拓扑学从其本质和设计上来说,适用于非线性、多尺度和非平稳的结构——然而,在机器学习中,其应用仍然在很大程度上局限于拓扑数据分析。我们主张,来自低维拓扑学的工具——这些工具几乎一直仅被包含在纯数学领域内(例如莫尔斯理论)——为数据生成过程以及建立在其上的学习任务的隐藏结构提供了一种强大、互补但几乎尚未被探索的视角。在此,我们引入配边理论中的概念,并利用离散莫尔斯理论中的工具,通过我们的流程 MG-Diff 提升图扩散模型的性能。进一步,我们推导出理论保证和充分条件,使得在正的决策间隔下,莫尔斯理论工具及其在诱导扩散引导中的应用在小扰动下保持稳定。最后,我们展示离散莫尔斯理论在应用于时空图预测和图再生的图扩散模型中的效用,并论证这些应用只是通往低维拓扑学能够为机器学习领域提供内容的一部分的一小扇窗口。
cs.LG / 144 / 2610.01947
Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry
Latent JEPA:用于化学中潜在推理的抽象未来预测
large language model
大语言模型相关
Abstract
Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.
Chinese Translation
大型语言模型为化学推理提供了有前景的基础,将化学知识与多步问题求解结合在一起。化学直觉可以在解决方案的细节被完全推导出来之前,提供对可能结果的初步感觉。受此类预期如何补充显式分析的启发,我们研究如何训练连续潜在思维,以在不必逐一表述每个中间步骤的情况下,预判未来解决方案中具有信息量的方面。我们提出 Latent JEPA,这是一个将自回归学习与对一个或多个未来视图的联合嵌入预测相结合的框架。针对化学推理,我们开发了文本和分子预测目标,将潜在思维同时与后续推理和分子结果联系起来。在 ChemCoTBench 上的实验显示,在分子优化以及若干编辑和反应指标上取得了提升。表征分析表明,未来预测使潜在思维对分子结果更具信息量,并增强其与化学结构的对应关系。这些发现支持将抽象未来预测作为连接连续潜在推理与科学结果的学习原则。
cs.LG / 145 / 2610.01967
FastCI: Efficient GPU-Intensive CI for LLM Training Frameworks
FastCI:面向 LLM 训练框架的高效 GPU 密集型 CI
large language model
大语言模型相关
Abstract
As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on GPU-intensive tests, which usually involve complete model training or evaluation. This leads CI itself to become a new bottleneck for fast-paced development. In this paper, we introduce FastCI, a framework that improves the efficiency of CI for LLM training frameworks. FastCI leverages runtime evidence to select affected tests and prune tests that execute changed code in equivalent contexts. Then FastCI prioritizes high-risk tests to expose potential failures earlier, and optimizes test workloads along dimensions outside the intended validation scope of each test. Evaluated on the CI workload of our LLM training framework, FastCI reduces the CI latency by 77.5% and the GPU resource usage by 63.9%, while improving the modified code coverage retention by 3.2%, compared with the currently deployed CI pipelines. FastCI has now been integrated into the CI pipelines of our LLM training framework at ByteDance.
Chinese Translation
随着大语言模型(LLM)在规模和复杂性上持续增长,其训练框架也在快速演进。因此,持续集成(CI)对于维持这些框架的质量和稳定性至关重要。然而,与传统软件不同,LLM 训练框架的 CI 依赖于 GPU 密集型测试,这些测试通常涉及完整的模型训练或评估。这导致 CI 本身成为快节奏开发的新瓶颈。在本文中,我们介绍了 FastCI,一个提高 LLM 训练框架 CI 效率的框架。FastCI 利用运行时证据来选择受影响的测试,并剪除那些在等价上下文中执行变更代码的测试。随后,FastCI 优先执行高风险测试以更早暴露潜在失败,并沿着每个测试预期验证范围之外的维度优化测试工作负载。在我们 LLM 训练框架的 CI 工作负载上进行评估,与当前部署的 CI 流水线相比,FastCI 将 CI 延迟降低了 77.5%,将 GPU 资源使用量降低了 63.9%,同时将修改代码覆盖率保留率提高了 3.2%。FastCI 现已被集成到我们在字节跳动的 LLM 训练框架的 CI 流水线中。
cs.LG / 146 / 2610.02039
CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
CARM:面向 LLM 强化学习的抵消感知响应掩码
large language model
大语言模型相关
Abstract
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
Chinese Translation
近年来,强化学习(RL)在大语言模型(LLM)后训练中的采用迅速增长,并在数学推理和代码生成方面取得了显著收益。然而,在实际系统中,策略更新以及 rollout 与训练引擎之间的差异可能使采样得到的响应偏离策略(off-policy)。序列级掩码通过决定整个响应是否应参与优化来应对这种不匹配。一种常见的掩码规则使用采样 token 概率比的长度归一化几何平均。其带符号的对数比可以在不同位置之间相互抵消,从而掩盖显著的双向策略漂移。我们提出 \emph{抵消感知响应掩码}(CARM),一种序列级掩码,它在平均之前对每个 token 对数比取绝对值,从而防止相反的概率变化相互抵消。我们证明,被接受的响应满足一个联合界,该联合界约束了采样 token 概率比落在规定区间之外的比例,以及它们超出其边界的平均对数距离。在数学推理和代码生成上的实验表明,与几何平均掩码相比,CARM 将在 AIME 2024/2025/2026 和 BeyondAIME 上平均的 mean@16 提高了最多 $3.13$ 个百分点,并在四个代码基准上的平均 pass@1 比评估的最强基线提高了 $2.88$ 个点。这些发现支持 CARM 作为一种有理论依据且有效的方法,用于 LLM 强化学习中的响应级离策略控制。
cs.LG / 147 / 2610.02182
SoftServe: A Scalable Quasi-Newton Method for Deep Learning
SoftServe:一种可扩展的深度学习拟牛顿方法
diffusion
扩散模型相关
Abstract
Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.
Chinese Translation
拟牛顿(QN)方法长期以来一直是大规模无约束凸优化中最有效的方法之一。有两个障碍限制了它们在深度学习中的应用:非凸性和庞大的参数规模。我们提出 SoftServe,一个 QN 方法系列,旨在无需线搜索或临时曲率校正的情况下克服这些障碍。SoftServe 从 Berglund et al. (2025) 的变分目标中推导出正定曲率估计,即使在存在负曲率的情况下也是如此。我们开发了对角线和 Kronecker 分解的变体,它们在构造上保持正定性,并可扩展到大规模神经网络。最后,SoftServe 依赖稳定的耦合 Newton-Schulz 迭代来完成所需的矩阵运算,用 GPU 友好的矩阵乘法取代代价高昂的矩阵分解。SoftServe 在严重病态的问题上表现优异,包括循环网络、深度自编码器、物理信息神经网络以及一个 1.36 亿参数的物理信息扩散模型等任务,通常比包括 Adam、Muon 和 SOAP 在内的既有基线取得更低的损失。
cs.LG / 148 / 2610.02191
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
缺失的基元:诊断与修复大型语言模型中的数学推理
large language model
大语言模型相关
Abstract
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.
Chinese Translation
尽管大型语言模型(LLMs)在前沿数学问题上展现出惊人的能力,但它们是否具备支撑其解法的结构性数学理解仍不清楚。本文迈出第一步,系统地研究 LLMs 中的数学理解,从诊断其不同能力到利用这些发现来改进后训练。首先,我们引入数学基元(Mathematical Primitive)的概念,以探究结构性数学理解,并提出 \hlei{},一个新颖的基准,从四个不同维度评估数学推理:发现(Discovery)、生成(Generation)、消化(Digestion)和执行(Execution)。其次,我们的系统诊断表明,解答准确率掩盖了不同的能力画像,基元释放出大量潜在执行能力,而发现(Discovery)是数学推理中的主要瓶颈。我们的后训练分析进一步表明,受发现限制的失败尤其易于修复。最后,基于这些发现,我们引入 \abs{},一个基元特权自蒸馏框架,该框架选择性地将基元引导的推理迁移到学生模型中。大量实验表明,\abs{} 在不同模型规模和具有挑战性的基准上持续优于基线,提升了数学推理能力。
cs.LG / 149 / 2610.02199
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
TACO:用于LLM微调的三值绝对值最大列向单稀疏优化器
large language model
大语言模型相关
Abstract
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO computes the exact steepest-descent direction under a dimension-normalized $1\to1$ operator norm by selecting the sign of the largest magnitude entry in each column of two-dimensional weight matrices. This retains first-order gradients while making optimizer state memory nearly negligible. Our practical TACO optimizer maintains only a small set of low precision gradient components per column, reducing persistent optimizer state by $174\times$ relative to AdamW8bit (from 27.7 GB to 0.16 GB) and peak training memory by $2.9\times$ (from 80.6 GB to 27.5 GB) on OPT-13B, while achieving comparable accuracy and runtime. TACO further enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100 GPU across multiple model families and tasks.
Chinese Translation
大语言模型(LLM)的全参数微调会带来巨大的优化器状态内存开销,限制了可在现代GPU上容纳的模型规模。现有方法要么压缩优化器状态,要么舍弃一阶梯度,要么在保留稠密状态的同时改变更新几何。最近提出的 Muon 优化器通过矩阵值更新来减少优化器内存。尽管如此,其几何与 AdamW 不同,并且在微调由 AdamW 预训练的模型时可能导致性能下降。为了在LLM微调中减少优化器内存且不牺牲精度或计算效率,我们提出三值绝对值最大列向单稀疏优化器,或 TACO,它遵循 Muon 的算子范数最速下降观点,但在几何路径上更进一步。TACO 通过在二维权重矩阵的每一列中选择绝对值最大元素的符号,计算在维度归一化 $1\to1$ 算子范数下的精确最速下降方向。这在保留一阶梯度的同时使优化器状态内存几乎可忽略不计。我们的实用 TACO 优化器每列仅维护一小组低精度梯度分量,在 OPT-13B 上相对于 AdamW8bit 将持久优化器状态减少 $174\times$(从 27.7 GB 到 0.16 GB),并将峰值训练内存减少 $2.9\times$(从 80.6 GB 到 27.5 GB),同时达到相当的精度和运行时间。TACO 进一步使得能够在单个 80 GB H100 GPU 上跨多个模型家族和任务对 30-32B 参数模型进行全参数微调。
cs.LG / 150 / 2610.00487
ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning
ScaffoldM3C:一种用于生成式稳定建造规划的多模态序贯蒙特卡洛框架
large language model
大语言模型相关
Abstract
Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for Multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds. Therefore, we formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task conditioning modalities. Concurrently, we explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. We present Scaffold Multimodal Monte Carlo (ScaffoldM3C), a multimodal, lightweight, auto-regressive model for stable block-based construction, that proposes a set of next-step candidate blocks. Leveraging these candidates, we utilize Sequential Monte Carlo (SMC) to maintain a population of possible assembly sequences, allowing us to consider multiple, potentially different, assembly directions simultaneously. We train our multimodal architecture by extending the StableText2Brick dataset to contain image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4x smaller than competing baselines, yielding a 5x to 20x speedup during inference, while achieving comparable construction quality to state-of-the-art methods and higher overall stability. We demonstrate the effectiveness of our approach through simulations and real-world robot assembly demonstrations.
Chinese Translation
自主构建物理上可实现的3D结构仍然是一个重大挑战,原因包括组合动作空间、可互换组件、殊途同归的装配序列以及建造过程中严格的稳定性要求。最先进的方法微调大语言模型以进行基于文本的生成式建造。然而,这些方法不允许进行多模态(文本、图像、草图)条件化,忽视了脚手架在稳定中间结构方面的实际作用,并且受到推理速度慢的困扰。因此,我们将建造形式化为一个概率性的下一块生成任务,该任务具有多个潜在装配动作和多个潜在任务条件化模态。同时,我们通过引入一个辅助脚手架块标记来显式考虑脚手架的效用。我们提出 Scaffold Multimodal Monte Carlo (ScaffoldM3C),一种用于稳定的基于块的建造的多模态、轻量级、自回归模型,它提出一组下一步候选块。利用这些候选块,我们使用序贯蒙特卡洛(SMC)来维持一个可能装配序列的种群,使我们能够同时考虑多个、可能不同的装配方向。我们通过扩展 StableText2Brick 数据集以包含图像条件化提示和脚手架稳定的建造序列,来训练我们的多模态架构。ScaffoldM3C 比竞争基线小 4 倍,在推理期间实现 5 到 20 倍的加速,同时达到与最先进方法相当的建造质量以及更高的整体稳定性。我们通过仿真和真实世界机器人装配演示展示了我们方法的有效性。
cs.AI / 151 / 2610.01093
OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous
OrbitTAMP:将语言模型锚定于航天器交会中的任务与运动规划
large language model
大语言模型相关
Abstract
Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-level operational intent into safe, dynamically feasible trajectories, creating a bottleneck to scalable operations. Large language model (LLM)-based agents could offer an intuitive interface for this process, although their outputs are not inherently grounded in orbital dynamics, operational constraints, or the structure of admissible spacecraft maneuvers. To exploit their semantic reasoning while ensuring the generated plan's physical validity, this paper presents a hierarchical framework for spacecraft task-and-motion planning (TAMP) that grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules. Within this framework, a pretrained LLM maps a natural-language command to a partial mission specification. The associated planners then resolve unspecified decisions within the admissible operational space. Finally, trajectory optimization converts the completed mission specification into a dynamically feasible trajectory. Numerical experiments demonstrate that this architecture substantially improves intent recovery over direct LLM generation, achieving 98% exact recovery of partial mission specifications across all evaluated splits when backed by frontier LLMs. Additional test-time-compute experiments show that, for a compact 9B model, verifier-guided revision increases exact recovery from 75% to 88%, while broader behavior-plan search independently improves selection among admissible trajectory realizations. Overall, these results establish a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO.
Chinese Translation
航天器交会与近距操作(RPO)目前通过一个高度依赖专业知识的流程进行规划,在该流程中,工程师将高层操作意图转化为安全且动力学可行的轨迹,这为可扩展操作造成了瓶颈。基于大语言模型(LLM)的智能体可为这一流程提供直观接口,尽管其输出并非天然地以轨道动力学、操作约束或允许的航天器机动结构为基础。为了利用其语义推理能力,同时确保所生成计划的物理有效性,本文提出了一种面向航天器任务与运动规划(TAMP)的分层框架,该框架将LLM推理锚定于一个由可复用行为和领域特定规划模块构成的图。在该框架内,一个预训练LLM将自然语言命令映射为部分任务规范。随后,相关规划器在允许的操作空间内解决未指定的决策。最后,轨迹优化将补全后的任务规范转换为动力学可行的轨迹。数值实验表明,相较于直接LLM生成,该架构显著提升了意图恢复能力,在由前沿LLM支持时,在所有评估划分上实现了对部分任务规范98%的精确恢复。额外的测试时计算实验表明,对于一个紧凑的9B模型,验证器引导的修正将精确恢复从75%提高到88%,而更广泛的行为规划搜索则独立地改善了在允许的轨迹实现之间的选择。总体而言,这些结果为航天器RPO的语言驱动智能体规划建立了一个可扩展且可审计的基础。
cs.LG / 152 / 2610.01959
Training-Free Diffusion Planning with Analytical Local Scores
使用解析局部分数的免训练扩散规划
diffusion
扩散模型相关
Abstract
Path finding and multi-robot motion planning require trajectories that are smooth, goal-directed, and collision-free in environments with complex geometric constraints. Recent diffusion-based planners have shown that trajectory generation can be cast as iterative denoising which has opened the doors to learning-based approaches that can handle multi-modal trajectory distributions and refine entire trajectories. However, a key limitation is that diffusion planners require training on large collections of feasible trajectories, rendering them map-specific, and difficult to deploy when high-quality demonstrations are unavailable. This paper introduces a training-free diffusion-based motion planner that replaces learned global trajectory scores with analytical local scores derived from obstacle, smoothness, velocity, and inter-agent feasibility terms. The proposed idea relies on a key observation: the score of a trajectory can be reconstructed by considering only local interactions between neighboring waypoints and nearby constraints. This structure exploitation yields a decomposed denoising procedure that retains the optimization structure of classical trajectory methods while inheriting the iterative refinement behavior of diffusion models. Experiments on a large collection of complex environments and large multi-agent planning tasks show that the proposed analytical score produces smooth and feasible trajectories within limited computational costs, for example in generating feasible paths for 300+ agents in environments containing 100+ obstacles in under 6 seconds on a GPU, outperforming strong learning-based and optimization baselines, while avoiding the data requirements of learned diffusion planners.
Chinese Translation
路径查找和多机器人运动规划要求轨迹在具有复杂几何约束的环境中平滑、朝向目标且无碰撞。最近的基于扩散的规划器已经表明,轨迹生成可以被表述为迭代去噪,这为能够处理多模态轨迹分布并细化整条轨迹的基于学习的方法打开了大门。然而,一个关键限制是,扩散规划器需要在大量可行轨迹集合上进行训练,这使它们变得特定于地图,并且在高品质演示不可用时难以部署。本文提出了一种免训练的基于扩散的运动规划器,它用由障碍物、平滑性、速度和智能体间可行性项推导出的解析局部分数替代学习得到的全局轨迹分数。所提出的想法依赖于一个关键观察:轨迹的分数可以通过仅考虑相邻路点之间的局部交互和附近约束来重构。这种结构利用产生了一种分解式去噪过程,它保留了经典轨迹方法的优化结构,同时继承了扩散模型的迭代细化行为。在大量复杂环境和大型多智能体规划任务上的实验表明,所提出的解析分数在有限计算成本内产生平滑且可行的轨迹,例如在 GPU 上于 6 秒内为包含 100+ 个障碍物的环境中的 300+ 个智能体生成可行路径,优于强大的基于学习和优化的基线,同时避免了学习型扩散规划器的数据需求。
cs.SE / 153 / 2610.00555
Extending LLM-based support for software engineers with ADHD
为患有 ADHD 的软件工程师扩展基于 LLM 的支持
large language model
大语言模型相关
Abstract
Software engineering workflows are often not designed to accommodate the needs of developers with Attention Deficit Hyperactivity Disorder (ADHD), despite known challenges related to task initiation, sustained attention, and completion. At the same time, large language models (LLMs) are increasingly integrated into programming tools, but existing systems do not account for neurodiversity or support structured progression across development tasks. In this work, we present Tether 2.0, an LLM based assistant designed to support software engineers with ADHD through workflow oriented interaction across planning, coding, debugging, and review. The tool was named Tether 2.0 because it builds upon the open source code and foundations of the original Tether system. Our approach combines structured interaction modes, activity aware context, and persistent memory to support task progression and continuity. We evaluate the system through expert feedback and hands on use with a software engineer with ADHD. Our results indicate that Tether 2.0 supports requirement clarification, task decomposition, incremental implementation, debugging, and completion, while enabling users to maintain progress and resume work after interruptions. These findings suggest that LLM based assistants can support direction, sustained progress, and learning in software development when designed with workflow structure and context aware interaction.
Chinese Translation
软件工程工作流往往并非为容纳患有注意缺陷多动障碍(ADHD)的开发者的需求而设计,尽管存在与任务启动、持续注意和任务完成相关的已知挑战。与此同时,大型语言模型(LLM)正日益被集成到编程工具中,但现有系统并未考虑神经多样性,也不支持跨开发任务的结构化推进。在这项工作中,我们提出了 Tether 2.0,这是一个基于 LLM 的助手,旨在通过面向工作流的交互,在规划、编码、调试和评审等环节为患有 ADHD 的软件工程师提供支持。该工具被命名为 Tether 2.0,因为它建立在原始 Tether 系统的开源代码和基础之上。我们的方法结合了结构化的交互模式、活动感知上下文和持久记忆,以支持任务的推进与连续性。我们通过专家反馈以及与一名患有 ADHD 的软件工程师的实际操作使用来评估该系统。我们的结果表明,Tether 2.0 支持需求澄清、任务分解、增量实现、调试和完成,同时使用户能够在中断后保持进度并恢复工作。这些发现表明,当以工作流结构和上下文感知交互进行设计时,基于 LLM 的助手能够在软件开发中支持方向感、持续进展和学习。
cs.SE / 154 / 2610.00622
Understanding and Mitigating Library-Related Issues in LLM-Generated Code
理解并缓解LLM生成代码中与库相关的问题
large language model
大语言模型相关
Abstract
Software practitioners increasingly rely on Large Language Models (LLMs) to generate code that integrates external libraries. However, LLMs often produce incorrect library usage, such as invalid imports, outdated API calls, and hallucinated dependencies, leading to compilation or runtime failures that reduce the reliability of AI-assisted software development. In this paper, we propose an agentic approach to mitigate libraryrelated errors in LLM-generated code. More specifically, we first conduct an exploratory study to characterize the library-related issues produced by LLMs. Our analysis of 100 LLM-generated code files reveals that 84% of generated files contain at least one library-related error, with recurring patterns including incorrect import paths, missing imports, hallucinated libraries, deprecated library usage, and unused imports. Based on these findings, we design an agentic approach that integrates task analysis, documentation grounding, code generation, and automated validation to improve library usage during code synthesis. We evaluate our approach on 300 code generation tasks derived from realworld implementations of rapidly evolving Python frameworks, including LangChain and AutoGen, across five LLMs: GPT-5, DeepSeek-V3, Qwen3, Mistral, and Llama 3. The results show that our approach consistently improves code generation quality across all evaluated models, reducing library-related errors by 38.1% - 54.6% and increasing code correctness by up to 16%.
Chinese Translation
软件从业者日益依赖大型语言模型(LLM)来生成集成外部库的代码。然而,LLM常常产生不正确的库使用方式,例如无效导入、过时的API调用以及幻觉依赖,从而导致编译或运行时失败,降低AI辅助软件开发的可靠性。在本文中,我们提出一种智能体方法来缓解LLM生成代码中与库相关的错误。更具体地说,我们首先开展一项探索性研究,以刻画LLM产生的与库相关的问题。我们对100个LLM生成的代码文件的分析表明,84%的生成文件包含至少一个与库相关的错误,其中反复出现的模式包括错误的导入路径、缺失的导入、幻觉库、已弃用的库用法以及未使用的导入。基于这些发现,我们设计了一种智能体方法,该方法集成了任务分析、文档依据、代码生成和自动验证,以在代码合成期间改进库的使用。我们在300个代码生成任务上评估了我们的方法,这些任务源自快速演进的Python框架的真实世界实现,包括LangChain和AutoGen,并覆盖五个LLM:GPT-5、DeepSeek-V3、Qwen3、Mistral和Llama 3。结果表明,我们的方法在所有被评估模型上持续提升代码生成质量,将与库相关的错误减少38.1% - 54.6%,并将代码正确性提高最多16%。
cs.SE / 155 / 2610.00629
ASAD: Adaptive Software Agents for Debugging
ASAD:用于调试的自适应软件智能体
large language model
大语言模型相关
Abstract
The integration of Large Language Models (LLMs) into multi-agent systems has shown great potential for automated debugging. Yet nearly all current frameworks rely on rigid, predefined architectures: the number of agents, their roles, and their interaction patterns are fixed before any analysis of the bug occurs. This one-size-fits-all approach is fundamentally mismatched to the heterogeneous nature of software defects. Simple bugs waste resources on unnecessary coordination, while complex ones suffer from insufficient or poorly aligned expertise. This paper introduces ASAD, an adaptive agentic system for debugging that configures its team according to the nature and complexity of each bug. ASAD initiates the debugging process by analyzing the faulty code and dynamically determines the number of agents to deploy, the specialized roles they should have, and the collaboration strategy they should follow. A central coordinator orchestrates this process through iterative planning, reflection, and execution; applying fast single-pass repairs for simple issues while assembling purpose-built teams to tackle more complex failures. We evaluate ASAD on three established benchmarks: Defects4J, DebugBench, and CodeFlaws, using multiple LLMs, such as DeepSeek-V3, Qwen-3 and GPT-5. ASAD consistently improves bug-fix rates by 12--20% over chain-of-thought(CoT) prompting and consistently outperforms static multi-agent systems by 4--9% in fix precision while reducing average agent usage by 32%. Crucially, our system dynamically adjusts the number and roles of agents: it resolves simple bugs with minimal coordination and scales agent involvement only for more complex cases.
Chinese Translation
将大型语言模型(LLMs)集成到多智能体系统中已在自动化调试方面展现出巨大潜力。然而,几乎所有当前框架都依赖于僵化的、预定义的架构:智能体的数量、角色以及交互模式都在对缺陷进行任何分析之前就已固定。这种一刀切的方法与软件缺陷的异构本质从根本上不匹配。简单的缺陷会在不必要的协调上浪费资源,而复杂的缺陷则因专业知识不足或匹配不当而受损。本文介绍 ASAD,一种用于调试的自适应智能体系统,它根据每个缺陷的性质和复杂性来配置其团队。ASAD 通过分析故障代码来启动调试过程,并动态确定要部署的智能体数量、它们应具有的专门角色,以及它们应遵循的协作策略。一个中央协调器通过迭代规划、反思和执行来编排这一过程;对简单问题应用快速单遍修复,同时组建专门构建的团队来应对更复杂的故障。我们在三个已有基准上评估 ASAD:Defects4J、DebugBench 和 CodeFlaws,并使用多个 LLM,例如 DeepSeek-V3、Qwen-3 和 GPT-5。ASAD 相对于思维链(CoT)提示将缺陷修复率持续提高 12--20%,并在修复精度上持续优于静态多智能体系统 4--9%,同时将平均智能体使用量降低 32%。至关重要的是,我们的系统动态调整智能体的数量和角色:它以最小协调解决简单缺陷,并且仅对更复杂的情况扩展智能体参与。
cs.SE / 156 / 2610.01372
A Design Theory for AI-Assisted Software Development Derived from Christopher Alexander's Theory of Form
一种源自克里斯托弗·亚历山大形式理论的AI辅助软件开发设计理论
large language model
大语言模型相关
Abstract
Code generated by large language models (LLMs) cannot be assumed to meet specified requirements. Reviews, testing, and static analysis still apply, but which of them a sufficient harness needs, and in what role, is open. We propose a design theory derived from Christopher Alexander's theory of form, and a methodology for applying it. In Alexander's account, fit between a form and its context can be perceived only negatively, through the absence of identified misfits. We make the organization's tradition explicit and derive the misfits from it and from the problem's classification. The theory models the LLM as a non-native vernacular builder, trained on many codebases but native to none, whose output tends to drift toward mainstream conventions rather than the local tradition. We engineer four pieces of machinery: explicit representations of the problem (Jackson's problem frames) and of the tradition (a four-form pattern language); deterministic misfit detectors; a fix loop; and a human-gated legislative circuit governing the representations and detectors. We call the resulting methodology, a practice of harness engineering, Misfit-Governed Development (MGD). Its dual-loop process separates an autonomous inner loop, where the LLM iterates against the gates, from a human outer loop, where specifications are judged against the world. Together they form the S = P = T = W assurance model (specification, program, tests, world), whose equals signs name relations, not identity. We report evidence from building and rebuilding a Scrum system of four event-sourced aggregates from 64 problem-frame specifications, verified by about 1,300 generated tests and 28 blocking gates, one applying 188 rules. This addresses the generativity dimension of Alexander's 1996 OOPSLA challenge. The moral dimension, whether the specification still fits the world, requires human judgment and belongs to the outer loop.
Chinese Translation
大型语言模型(LLM)生成的代码不能被假定为满足指定的需求。评审、测试和静态分析仍然适用,但一个充分的约束装置需要其中哪些,以及它们以什么角色发挥作用,仍是未决问题。我们提出一种源自克里斯托弗·亚历山大形式理论的设计理论,以及一套应用它的方法论。在亚历山大的论述中,形式与其语境之间的契合只能以否定的方式被感知,即通过已识别出的不契合的缺席来感知。我们将组织的传统显式化,并从该传统以及问题的分类中推导出不契合。该理论将LLM建模为一名非本土的乡土建造者,它在众多代码库上接受训练,却不属于其中任何一个的本土,其输出倾向于漂向主流惯例,而非本地传统。我们设计了四件机制:问题(杰克逊的问题框架)与传统(一种四形式模式语言)的显式表示;确定性的不契合检测器;一个修复循环;以及一个由人类把关、治理这些表示与检测器的立法回路。我们将由此产生的方法论——一种约束装置工程的实践——称为不契合治理开发(MGD)。其双循环过程将自主的内循环——在其中LLM针对各道关卡反复迭代——与人类的外循环——在其中规格对照世界受到评判——分离开来。二者共同构成 S = P = T = W 保证模型(规格、程序、测试、世界),其中的等号命名的是关系,而非同一性。我们报告了来自构建与重建一个 Scrum 系统的证据:该系统由四个事件溯源聚合构成,源自 64 份问题框架规格,并由约 1,300 个生成的测试和 28 道阻断性关卡验证,其中一道关卡应用了 188 条规则。这回应了亚历山大 1996 年 OOPSLA 挑战中的生成性维度。道德维度,即规格是否仍然契合世界,需要人类判断,并且属于外循环。
cs.SE / 157 / 2610.01847
Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
以LLM作为验证器的推理检测模型规范中的不一致性
large language model
大语言模型相关
Abstract
Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.
Chinese Translation
模型规范定义了大语言模型(LLM)应如何行为,指导对齐训练、推理时行为和评估。然而这些规范本身可能包含缺陷:两条各自合理的原则在应用于同一情境时可能规定互不兼容的行为,使得不存在同时满足两者的回应。检测此类不一致具有挑战性。将自然语言规范形式化存在丢失细微差别的风险,而基于行为的测试无法可靠地区分规范缺陷与模型行为差异。我们提出VeriSpec,这是首个通过审计规范文本本身来直接检测模型规范中不一致的方法。我们的核心洞见是:在保留自然语言形式的规范的同时,使用LLM作为验证器。VeriSpec提取结构化的、情境感知的规则,构建主题引导的图以在同一权威层级上聚类行为相关的规则,并应用LLM作为验证器的推理来检测不一致。将VeriSpec应用于OpenAI Model Spec,我们提取了405条规则并人工验证了五处不一致,全部已报告给其开发者,开发者给予了积极回应并已启动内部讨论。与五个基线相比,VeriSpec识别出最多经验证的不一致,取得最高精确率(38.5%),并且每处经验证不一致的成本最低(11.12美元)。这些结果确立了直接规范审计作为行为对齐评估的实用补充,在缺陷塑造任何模型之前于源头将其捕获。代码可在 https://github.com/HIPREL-Group/VeriSpec 获取。
cs.LG / 158 / 2610.02128
Sample complexity bounds for categorical Markov random fields via Discrete Diffusions
通过离散扩散的类别马尔可夫随机场样本复杂度界
diffusion
扩散模型相关
Abstract
Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence structures. Examples include finite memory language models, Ising and Potts systems in statistical physics and protein folding, etc. In modern machine learning, discrete diffusions have emerged as a flexible approach for sampling such data, with strong empirical performance. Motivated by this, we develop learning methods with end-to-end sample complexity bounds for discrete diffusion with uniform noising under local dependence, which we model through low order Markov random fields (MRFs). Our main technical insight is a new \emph{pinning decomposition} of the discrete score. It shows that unlike in continuous diffusions, the score decomposes into components where the dependence on time separates multiplicatively from the dependence on the target. Building on this decomposition, we propose a \emph{weight-sharing neural score learner} and combine it with $τ$-leaping to obtain an end-to-end sampling procedure. Rather than treating score-learning error as a black-box input, as is common in existing sampling analyses, we study the score learning error from finite data and derive optimal sampling guarantees with explicit dependence on the vocabulary size, the interaction order of the MRF, and the sample size. Moreover, our strategy trains a single score network across uniform noise levels while leaving the sampling discretization to be chosen at inference-time. This allows the same trained model to trade accuracy for computational cost as inference-time budgets vary. Numerical experiments on Potts, Ising, and tree-structured models show that weight-sharing score networks outperform fully connected ones for sampling long sequences.
Chinese Translation
许多统计学、经济学和物理学中的应用都需要从具有局部依赖结构的高维类别分布中采样。示例包括有限记忆语言模型、统计物理和蛋白质折叠中的 Ising 与 Potts 系统等。在现代机器学习中,离散扩散已成为对此类数据进行采样的一种灵活方法,并具有强大的实证性能。受此启发,我们为局部依赖下具有均匀加噪的离散扩散开发了具有端到端样本复杂度界的学习方法,其中局部依赖通过低阶马尔可夫随机场(MRFs)建模。我们的主要技术洞见是离散得分的一种新的“钉扎分解”(pinning decomposition)。它表明,与连续扩散不同,得分分解为若干分量,其中对时间的依赖与对目标的依赖以乘法方式分离。基于这一分解,我们提出了一种权重共享的神经得分学习器,并将其与 $τ$-leaping 结合,以获得端到端的采样过程。我们没有像现有采样分析中常见的那样将得分学习误差视为黑箱输入,而是研究来自有限数据的得分学习误差,并推导出最优采样保证,其显式依赖于词汇表大小、MRF 的交互阶数以及样本量。此外,我们的策略在均匀噪声水平上训练单个得分网络,而将采样离散化留待在推理时选择。这允许同一个训练好的模型在推理时预算变化时在精度与计算成本之间进行权衡。在 Potts、Ising 和树结构模型上的数值实验表明,对于采样长序列,权重共享得分网络优于全连接得分网络。
cs.LG / 159 / 2610.00535
Adaptive Conformal Prediction for Image Regression Models with Application to an Inertial Confinement Fusion Emulator
面向图像回归模型的自适应共形预测及其在惯性约束聚变仿真器中的应用
diffusion
扩散模型相关
Abstract
Uncertainty quantification is critical in scientific machine learning, where black-box, image-based models are increasingly deployed in high-stakes settings. In many such applications, model outputs inform costly decisions, yet most methods provide only point estimates without quantifying predictive uncertainty. This challenge is compounded by the limited accessibility and interpretability of model internals, making it difficult to assess reliability across different regions of the input space. As a result, there is a growing need for methods that can provide input-dependent uncertainty estimates to guide both model development and downstream experimentation. To address this need, we propose Adaptive Conformal Prediction using Nearest Neighbors (ACPNN), an input-adaptive conformal framework for image regression. ACPNN leverages information from neighboring samples to produce locally adaptive uncertainty estimates while maintaining low computational cost. The neighborhood structure is defined using a scaled distance metric learned via a Gaussian Process with an automatic relevance determination (ARD) kernel. We demonstrate the effectiveness of ACPNN on a diffusion model for emulating inertial confinement fusion (ICF) simulations, showing that it achieves reliable and adaptive uncertainty quantification.
Chinese Translation
不确定性量化在科学机器学习中至关重要,在科学机器学习中,黑箱的、基于图像的模型正越来越多地部署于高风险场景。在许多此类应用中,模型输出为代价高昂的决策提供依据,然而大多数方法只提供点估计,而不量化预测不确定性。模型内部结构的可访问性和可解释性有限,使这一挑战更加复杂,从而难以评估模型在输入空间不同区域上的可靠性。因此,越来越需要能够提供依赖于输入的不确定性估计的方法,以指导模型开发和下游实验。为满足这一需求,我们提出使用最近邻的自适应共形预测(Adaptive Conformal Prediction using Nearest Neighbors,ACPNN),一种用于图像回归的输入自适应共形框架。ACPNN 利用来自邻近样本的信息来生成局部自适应的不确定性估计,同时保持较低的计算成本。邻域结构通过一种缩放距离度量来定义,该度量经由带有自动相关性确定(automatic relevance determination,ARD)核的高斯过程学习得到。我们在一个用于模拟惯性约束聚变(inertial confinement fusion,ICF)仿真的扩散模型上展示了 ACPNN 的有效性,表明其实现了可靠且自适应的不确定性量化。
cs.LG / 160 / 2610.00793
Inference for stochastic differential equations driven by weighted sub-fractional Brownian motion using neural networks and the Euler approximation
使用神经网络和Euler近似对加权次分数布朗运动驱动的随机微分方程进行推断
diffusion
扩散模型相关
Abstract
We consider the estimation of drift, diffusion, and noise covariance from discrete observations of stochastic differential equations driven by Gaussian processes. For a fixed observation horizon $T>0$ and a known initial state $x_0\in\mathbb R$, we study \begin{equation*} dX_t=a(X_t)\,dt+σ(X_t)\,dZ_t^{β,f}, \qquad X_0=x_0,\quad 0\leq t\leq T. \end{equation*} \smallskip\noindent Here $a:\mathbb R\to\mathbb R$ is the drift coefficient, $σ:\mathbb R\to(0,\infty)$ is the diffusion coefficient, and $Z^{β,f}$ is a centered Gaussian process from the weighted sub-fractional Brownian family, with covariance \begin{equation*} \operatorname{Cov}(Z_s^{β,f},Z_t^{β,f}) =\int_0^{s\wedge t} f(r)q_β(s-r,t-r)\,dr, \qquad 0\leq s,t\leq T. \end{equation*} \smallskip\noindent Here $s\wedge t=\min\{s,t\}$. The temporal weight $f:[0,T]\to[0,\infty)$ is measurable, bounded, and positive almost everywhere, and $β\in(0,2)$ is the covariance exponent. For $u,v\geq0$, the kernel is $q_β(u,v)=[u^β+v^β-(u+v)^β]/(1-β)$ when $β\ne1$. Its continuous extension at $β=1$ is $q_1(u,v)=(u+v)\log(u+v)-u\log u-v\log v$, with $0\log0=0$. Using the Euler approximation, we reconstruct the Gaussian driving increments from observed transitions and use their joint density to obtain a trajectory likelihood. Neural and radial-basis representations model the drift, diffusion, and normalized temporal weight, while a likelihood profile estimates the covariance exponent and diffusion scale. We compare the method with two neural alternatives on the same simulated trajectories in twenty coefficient settings.
Chinese Translation
我们考虑从由高斯过程驱动的随机微分方程的离散观测中估计漂移、扩散和噪声协方差。对于固定的观测时域 $T>0$ 和已知初始状态 $x_0\in\mathbb R$,我们研究 \begin{equation*} dX_t=a(X_t)\,dt+σ(X_t)\,dZ_t^{β,f}, \qquad X_0=x_0,\quad 0\leq t\leq T. \end{equation*} \smallskip\noindent 这里 $a:\mathbb R\to\mathbb R$ 是漂移系数,$σ:\mathbb R\to(0,\infty)$ 是扩散系数,而 $Z^{β,f}$ 是来自加权次分数布朗运动族的中心化高斯过程,其协方差为 \begin{equation*} \operatorname{Cov}(Z_s^{β,f},Z_t^{β,f}) =\int_0^{s\wedge t} f(r)q_β(s-r,t-r)\,dr, \qquad 0\leq s,t\leq T. \end{equation*} \smallskip\noindent 这里 $s\wedge t=\min\{s,t\}$。时间权重 $f:[0,T]\to[0,\infty)$ 是可测的、有界的且几乎处处为正,而 $β\in(0,2)$ 是协方差指数。对于 $u,v\geq0$,当 $β\ne1$ 时,核为 $q_β(u,v)=[u^β+v^β-(u+v)^β]/(1-β)$。它在 $β=1$ 处的连续延拓为 $q_1(u,v)=(u+v)\log(u+v)-u\log u-v\log v$,其中 $0\log0=0$。使用 Euler 近似,我们从观测到的转移重构高斯驱动增量,并使用它们的联合密度来获得轨迹似然。神经网络表示和径向基表示对漂移、扩散和归一化时间权重进行建模,而似然轮廓估计协方差指数和扩散尺度。我们在二十种系数设定下,在相同的模拟轨迹上将该方法与两种神经网络替代方法进行比较。
cs.LG / 161 / 2610.01933
Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models
面向不完美扩散模型的误差校正推理时扩展
diffusion
扩散模型相关
Abstract
Inference-time scaling adapts pretrained diffusion models to new sampling tasks without additional training. Existing methods rely primarily on Monte Carlo sampling with more particles, yet are premised on the pretrained model being exact. In practice, data and training limitations make the model imperfect, and these methods inherit its error. More particles reduce Monte Carlo error but cannot remove the mismatch between the endpoint and the desired target or the error in tracking the prescribed probability path. We introduce the Energy-based Feynman-Kac Corrector (EBFKC), a framework for energy-based diffusion models that corrects these errors on the fly given a reference energy. We first derive Feynman-Kac dynamics that track a prescribed path exactly in the continuous-time population limit even when the model is imperfect, and approximate these dynamics using sequential Monte Carlo with variance-controlling guidance. To remove the endpoint mismatch, we use the pretrained energy as a surrogate along the diffusion path and progressively incorporate the discrepancy between the learned and target terminal energies. Experiments on Gaussian mixture models, particle systems, alanine dipeptide, and alanine tetrapeptide show that our method closely matches target distributions and molecular free-energy profiles under annealing and reward tilting, whereas standard inference-time scaling baselines retain substantial sampling errors.
Chinese Translation
推理时扩展无需额外训练即可将预训练扩散模型适配到新的采样任务。现有方法主要依赖使用更多粒子的蒙特卡洛采样,然而其前提是预训练模型是精确的。在实践中,数据和训练方面的限制使模型并不完美,而这些方法会继承其误差。更多粒子可以减少蒙特卡洛误差,但无法消除端点与期望目标之间的不匹配,也无法消除在跟踪规定概率路径时的误差。我们提出基于能量的 Feynman-Kac 校正器(Energy-based Feynman-Kac Corrector,EBFKC),这是一个面向基于能量的扩散模型的框架,能够在给定参考能量的情况下即时校正这些误差。我们首先推导出 Feynman-Kac 动力学,即使在模型不完美时,也能在连续时间总体极限下精确地跟踪规定路径,并使用带有方差控制引导的序贯蒙特卡洛来近似这些动力学。为消除端点不匹配,我们在扩散路径上使用预训练能量作为代理,并逐步引入学习到的终端能量与目标终端能量之间的差异。在高斯混合模型、粒子系统、丙氨酸二肽和丙氨酸四肽上的实验表明,我们的方法在退火和奖励倾斜下能够紧密匹配目标分布与分子自由能曲线,而标准的推理时扩展基线仍保留着显著的采样误差。
cs.LG / 162 / 2610.02081
Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling
Wasserstein 梯度流与仅前向扩散不足以实现多模态采样
diffusion
扩散模型相关
Abstract
There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods can efficiently sample complex multimodal distributions, often supported by empirical results. In this work, we argue that this interpretation is fundamentally misleading. By invoking the Jordan-Kinderlehrer-Otto (JKO) scheme and Otto calculus, we establish that the canonical WGF sampling dynamics and overdamped forward diffusion share the same density evolution and therefore inherit the same metastability and slow-mixing phenomena long understood in nonequilibrium statistical physics. We analyze this family of samplers using two complementary tools -- spectral analysis and mean first-passage time (MFPT) analysis -- and show that well-separated multimodality can induce exponentially long mixing times associated with small spectral gaps and rare inter-mode transitions. For the commonly adopted log-linear annealing schedule studied here, we find that introducing intermediate distributions does not remove the exponential scaling of the total transport time. The limitation is structural rather than implementation-specific: purely local, gradient-driven transport mechanisms can require exponentially long times to transport probability mass across well-separated modes. We argue that this represents a fundamental limitation of WGF- and FODP-based sampling in their standard forms, and motivates future development of fundamentally nonlocal mechanisms for efficient multimodal sampling.
Chinese Translation
基于 Wasserstein 梯度流(WGF)和仅前向扩散过程(FODP)的采样算法大量涌现,它们通常伴随着以指数级速度快速收敛到目标分布的理论保证。这些保证经常被解释为证据,表明此类方法能够高效地采样复杂的多模态分布,且往往得到经验结果的支持。在这项工作中,我们论证这种解释从根本上是误导性的。通过引入 Jordan-Kinderlehrer-Otto(JKO)格式和 Otto 演算,我们确立:典型的 WGF 采样动力学与过阻尼前向扩散共享相同的密度演化,因而继承了非平衡统计物理中早已理解的相同亚稳性和慢混合现象。我们使用两种互补工具——谱分析和平均首次穿越时间(MFPT)分析——来分析这一族采样器,并表明分离良好的多模态性可以引发与小的谱隙和稀有的模态间跃迁相关的指数级长混合时间。对于这里研究的常用对数线性退火调度,我们发现引入中间分布并不能消除总传输时间的指数级缩放。这一局限是结构性的,而非特定实现所特有的:纯局部、梯度驱动的传输机制可能需要指数级长的时间才能将概率质量跨分离良好的模态进行传输。我们论证,这代表了标准形式的基于 WGF 和 FODP 的采样的一个根本性局限,并促使未来开发从根本上非局部的机制,以实现高效的多模态采样。
人工智能 (cs.AI)
133
cs.AI / 1 / 2610.00511
Before Agents Decide: Epistemic Action in LLM-Based Systems
Abstract
Before a difficult decision, people often act simply to understand the situation better. We turn an object to see another side, place alternatives next to each other, or change one condition and observe what happens. These actions may not complete the task, but they improve the evidence needed for the next choice. LLM-based agents can search and explore, yet agent design gives less attention to an earlier question: is the available evidence ready for the decision? Sometimes necessary evidence is missing. In other cases, the evidence is present but its form hides what matters, or the comparison needed to judge it does not yet exist. Cognitive science calls actions that improve the basis for a later choice epistemic actions. We bring this idea to LLM-based agents and distinguish three modes: acquiring missing evidence, transforming available evidence, and probing a system to create a revealing response. We use the term epistemic scaffolding for the interfaces, tools, and environments that make these actions possible and auditable. This paper argues that agent design must address how decision-ready evidence is produced.
cs.AI / 2 / 2610.00529
Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology
Abstract
The ontology-based contextual AI evaluation (OB-CAIE) methodology was developed to address a lack of scientific rigor that arises from unclear testing coverage, to balance human expertise and automations, and to address a lack of reproducibility of AI evaluation testing environments. OB-CAIE strengthens the current state of AI evaluations by addressing the first step in the scientific method by clearly defining what will be tested. Two ontologies represent the tractable problem space in the OB-CAIE methodology: the Domain-Specific Ontology (DSO) and the Evaluation Process Ontology (EPO). The DSO is the what; the EPO is the how. An OB-CAIE problem space can be used for one or multiple AI evaluations. The OB-CAIE methodology allows for human judgment at specific points, in scientifically grounded ways, and in complex subject areas where human feedback is genuinely irreducible or machine irreplaceable. A key advantage of the OB-CAIE methodology is that failure points can be traced, visualized and analyzed within the canonical OB-CAIE methodology problem space.
cs.AI / 3 / 2610.00531
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
Abstract
AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the scientific reasoning that connects the parts breaks down, which can mislead how readers assess the work. We benchmark these failures as scientific slop through six measures across Structure, Argument, and Artifacts. We construct SciSlopBench with 390 AI-generated papers, mostly in computer science but spanning the life, social, and natural sciences, each paired with a human-written paper matched by research problem and contribution type. Our measures identify the AI paper in each pair with 85.9% accuracy, compared with 68.7% for Binoculars. Higher scientific slop accompanies lower ICLR ratings and distinguishes rejected from accepted papers above chance in every year from 2017 to 2025. Reducing these patterns, however, is not as simple as directly optimizing the measures. We therefore propose SciSlopHarness, a harness-level framework that guides a fixed LLM to revise slop only where the experiment records support the change. While standard revisions leave residual slop and direct slop-aware prompting triggers reward hacking, SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline without requiring human reference targets. Overall, we demonstrate that AI-generated scientific papers leave fundamental traces in their global reasoning, and that responsible mitigation demands strict evidentiary grounding rather than mere prose refinement.
cs.AI / 4 / 2610.00583
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
Abstract
People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.
cs.AI / 5 / 2610.00609
Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
Abstract
Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbf{Legal Research Bench} (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness with web search, case-law search, page parsing, and retrieval tools. We score agent responses through all-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4.8, is fully correct on 42.9\% of questions. Performance also varies substantially by task setting: all-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities. Across models, more turns, tool calls, and inference cost do not predict higher accuracy.
cs.AI / 6 / 2610.00636
CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
Abstract
Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.
cs.AI / 7 / 2610.00648
Incident-Arena: Getting agents to the last nine of reliability
Abstract
AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks limited by (1) unrealistic environments, typically toy repositories (2) non-standard framework implementations and (3) simple static verifiers. We introduce Incident-Arena, a human-built benchmark of 20 carefully selected tasks grounded in real-world deployed open source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injecting a fault from the config layer through underlying images, and a sustained load profile given the task requirements. We also present a novel verification method, going beyond static checks to functional verifiers, holding systems level metrics stable, while ensuring repairs are done safely. Agent trials run an average of 2.81M tokens and 41 turns, going beyond existing benchmarks, demonstrating agentic long horizon reasoning. Across 20 tasks and 3 application substrates, frontier models score below 64.3%, with failures extending from diagnosis/localization errors, through incomplete repairs and unsafe regressions.
cs.AI / 8 / 2610.00651
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
Abstract
Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliability claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models. We ask which conclusions current agent evaluations reliably support and what additional evaluation would improve them. We develop a Bayesian variance-decomposition framework for sparse, imbalanced agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index. The framework separates signal, performance differences relevant to the intended claim, from noise, irrelevant variation that can still change rankings. We find: (1) Reliability depends on the measurement goal. Fixed model-scaffold systems are ranked reliably (0.935-0.994), while underlying-model reliability is substantially lower (0.148-0.841). (2) Scaffold choice can change conclusions. Inter-scaffold reliability measures whether scaffolds preserve model rankings, showing that scaffold effects vary substantially across evaluations. (3) More tasks cannot resolve all uncertainty. Even infinitely many similarly constructed tasks improve model-ranking reliability of a benchmark by at most 0.097 when uncertainty is dominated by limited scaffold coverage. (4) Pooling diverse benchmarks can improve cross-task rankings at lower cost. For rankings across diverse agentic tasks, pooling benchmarks raises projected reliability from 0.44 to 0.75 at the same task budget and can reduce projected cost by up to 83\%. Evaluation design should follow the intended claim: identify what a score or ranking should mean, diagnose what limits its reliability, and spend evaluation budget on the sources of uncertainty that matter.
cs.AI / 9 / 2610.00654
When More Data Is Not Enough: The Context-Sufficiency Frontier in Generative AI Personalization
Abstract
Personalization has long relied on customer data to infer what an individual is likely to value. We call this customer evidence: the customer's historical behavior and preferences. Generative AI extends personalization by allowing providers to supply changing situational information at the moment a response is produced, without encoding every condition in advance. We define this provider-side context as information about what is possible, permitted, or advisable now. This flexibility creates a new problem: once context becomes easy to supply, more is not necessarily better. We develop a theory of context sufficiency in which the relevance of context to the customer's current intent matters more than its volume. The theory identifies four states, insufficiency, sufficiency, saturation, and interference, and introduces the Context-Sufficiency Frontier to locate the minimal relevant set. In a full-factorial experiment with a generative recommender at a large home-furnishing retailer, relevant context improved appropriateness, while irrelevant context reduced it and destabilized retrieval. The framework shifts personalization from supplying more context toward identifying what the current interaction actually requires and enforcing constraints throughout the service process.
cs.AI / 10 / 2610.00668
A Simple Doxastic Deontic Logic for Norm-Guided Decision Making
Abstract
Making decisions despite conflicting norms and incomplete or unreliable information is a fundamental challenge for autonomous systems. We introduce a simple doxastic deontic logic for this setting: a classically reducible fragment of Chellas' Minimal Deontic Logic, extended with explicit conditional norms and combined with multi-agent KD45, so that norms can depend on agents' beliefs about both facts and norms. On this logic we define the Doxastic Norm Compliance Optimization Problem, where an agent chooses a decision minimizing weighted norm violations. We distinguish subjective optimization (relative to the agent's beliefs) from objective optimization (relative to the actual facts). We give conditions under which (i) the two coincide and (ii) optimal decision-making can be reduced to weighted partial MaxSAT in polynomial time.
cs.AI / 11 / 2610.00700
R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing
Abstract
Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely unevaluated.We introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation track.Our results reveal a substantial gap between recognition andmolecular grounding.While models achieve over 90\% accuracy on Easy VQA, performance drops to56--66\% on Hard VQA when shortcuts are controlled.Chemical-domain VLMs also remain unreliable, achieving only 25.7--46.2\% on HardVQA despite domain-specific pretraining.Moreover, Generation Exact Match remains below 20\% for most models and below8\% when visual input is required.These findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
cs.AI / 12 / 2610.00705
Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
Abstract
This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism. However, existing meta-RL generally focuses on single-agent systems. Extending these frameworks and algorithms to multi-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents' strategic interactions. To address these challenges, we model multi-agent reinforcement learning (MARL) problems as Markov games (MGs) and develop a meta-MARL framework for rapid interactive policy adaptation across a distribution of MGs. A new concept, called meta-NE, is defined to describe the desired solution concept in a meta-MARL problem. Sufficient conditions for the equivalence between a meta-NE and a stationary point of the gradient-play-based meta-MARL algorithm are established. Our evaluation on autonomous-driving tasks demonstrates that the proposed meta-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework.
cs.AI / 13 / 2610.00715
Robust Nash Alignment under Preference Uncertainty
Abstract
Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain pairwise preferences. Our formulation has a major learner seeking a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting robust objective of the game directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for it. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an \(\mathcal{O}(1/\sqrt{T})\) average-iteration convergence for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
cs.AI / 14 / 2610.00791
Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AI
Abstract
Enterprise information is represented through artifacts shaped by applications, projects, technologies, organizational boundaries, and local requirements. These structures accumulate over time, creating representational complexity that must be maintained by the enterprise and interpreted by information consumers and AI systems. This paper introduces Enterprise Representation Simplification (ERS) as reducing unnecessary representational complexity while preserving required information within a defined scope, and Enterprise Representation Complexity (ERC), a representation-neutral model for comparing complexity across representation states. ERC characterizes representational extent through four dimensions: Representation Objects, Interactions, Behaviors, and Supporting Sources. Objects, Interactions, and Behaviors form dependent categories, while Supporting Sources characterize representation exposure. ERC is defined at representation and task levels, enabling comparison and distinguishing architectural simplification from retrieval optimization. The paper develops two consequences of ERS. First, representational structures create lifecycle obligations for maintenance, governance, dependencies, change, enhancement, and operation. An economic model distinguishes recurring global representation cost, recurring task-level cost, and one-time transformation cost, enabling evaluation over a defined time horizon. Second, reductions in task-level ERC reduce the representational extent an AI system must identify, relate, and interpret. Text-to-SQL research provides evidence that reduced schema and reasoning complexity can improve reasoning accuracy. ERC is not a universal complexity, performance, or cost metric. It provides measurable architectural variables for comparing representational alternatives, transformation effects, economic outcomes, and AI reasoning performance.
cs.AI / 15 / 2610.00797
Sapien: A Stateful Policy Engine for Autonomous AI Agents
Abstract
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
cs.AI / 16 / 2610.00834
Kepler: Auditable World Models for ARC-AGI-3
Abstract
ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a \$777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.
cs.AI / 17 / 2610.00849
Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
Abstract
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
cs.AI / 18 / 2610.00870
An Educator-Guided LLM Pedagogical Agent for Scaffolded Feedback in Conceptual Database Design
Abstract
We present an educator-guided LLM pedagogical agent for scaffolded feedback in conceptual database design. Integrated into an entity--relationship diagram (ERD) editor, the system grounds feedback in the student artifact, assignment requirements, educator-authored rubrics, and instructional resources. Its architecture separates hidden, artifact-grounded diagnosis from the workflow that controls the form and disclosure level of student-facing support. We instantiate the architecture as a four-stage workflow progressing from concept checks and guided application to low-detail feedback and localized clarification. Each feedback request creates a stateful episode linked to versioned ERD states. In a deployment spanning three ERD environments and 383 feedback episodes, 71.1\% of observed target-level changes fully or partially incorporated the hidden diagnostic target, including many after Stages~1--2. Qualitative analysis showed that staged disclosure sometimes withheld inaccurate details, supported selective uptake, or allowed later recovery, though some errors still shaped revisions. Survey responses from a self-selected sample favored delayed disclosure and student agency but noted indirectness and repetition.
cs.AI / 19 / 2610.00906
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Abstract
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.
cs.AI / 20 / 2610.00917
Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Abstract
Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at https://github.com/liyix/finding-the-right-fit and https://huggingface.co/datasets/yixuanli97/finding-the-right-fit.
cs.AI / 21 / 2610.00949
PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
Abstract
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying \emph{how to better balance the trade-off between acquiring new capabilities and preserving existing ones during agent trace SFT}. By comparing several baselines in our setup, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivated by recent token-wise adaptive learning objectives, this work proposes \textbf{Privilege-Guided SFT (PG-SFT)} to leverage turn-level information gain of agent trajectories as an indicator to adjust supervision strength. PG-SFT yields a more favorable observed trade-off on the evaluated benchmarks, substantially reducing distributional drift and broad capability degradation at the cost of slight degradation in target-task performance. Our findings suggest that balancing the acquisition--retention trade-off depends not only on whether the model is anchored to its base behavior, but also on where and how strongly supervision should depart from that behavior.}
cs.AI / 22 / 2610.00961
Cybernetic and Epistemic: A Missing Vocabulary for Trustworthy Agentic Delegation
Abstract
As code generation is increasingly delegated to AI systems, the bottleneck is shifting from writing code to supervising the systems that write it --- a shift CS-education researchers have begun to name. This shift exposes a vocabulary gap: the field asks for "human oversight" without a working distinction between the two things language does in a delegation channel --- coordinate action (cybernetic: words succeed when the world comes to match them) and coordinate understanding (epistemic: they succeed when they answer to the world and a hearer can check that they do). The failure this names is not cybernetic language but epistemic-form language doing cybernetic work: explanation-shaped output calibrated for approval rather than truth. Oversight that checks only whether an output was approved is satisfiable by rubber-stamping; oversight that holds an agent accountable requires the reasoning behind its work be retrievable and checkable. We present three delegation episodes --- illustrations, not controlled evidence --- in which epistemic engagement proved practicable while remaining auditable, one public record where a recommendation was withdrawn on its own stated terms, and one failure case illustrating oversight that requires no reasons for its discretionary choices. We propose a criterion for agentic-system governance, alongside existing technical trust properties: every consequential choice should carry the condition under which it would have gone otherwise, in a form a third party can test. Without such a condition, a third party cannot distinguish a decision from a rubber stamp. We give the criterion an operational form --- a two-part reconstruction test scoring a delegation record by whether a second reader can predict what the agent does under a perturbation --- and a deliberation-recording convention, ORRCF, that makes the condition a required component of every recorded choice.
cs.AI / 23 / 2610.00972
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
Abstract
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
cs.AI / 24 / 2610.00979
RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation
Abstract
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics' usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.
cs.AI / 25 / 2610.01001
Calibration-risk routing for controlled world-model adaptation
Abstract
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.
cs.AI / 26 / 2610.01002
What Can Analogy Tell Us About Artificial Consciousness?
Abstract
Who or what is conscious? Because subjective experience is directly accessible only in the first person, judgments about consciousness in other entities depend partly on analogy. Historically, such inferences have focused on nonhuman animals, but advances in artificial intelligence have raised the possibility of conscious AI. Here we develop a causal framework for evaluating such evidential analogies. The key distinction is between similarities in factors plausibly involved in generating consciousness and similarities in downstream behavioural or cognitive effects. Our framework weights source-target similarity by causal relevance while allowing for unknown causes, disabling differences and alternative routes to consciousness. Applied to biological systems, it explains why analogical support generally weakens with increasing causal distance from humans. Applied to contemporary AI, it suggests that behavioural similarity provides only limited evidence for consciousness because relevant causal correspondences remain poorly established. The framework also clarifies what evidence would strengthen claims of artificial consciousness.
cs.AI / 27 / 2610.01006
Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model
Abstract
Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev's confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21--0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: https://github.com/Syntheme/beyond-answer-confidence.
cs.AI / 28 / 2610.01014
From Discovery to Decision: Finite-Budget Recoverability in LLM Voting
Abstract
Voting over multiple LLM responses is a common primitive in test-time scaling and ensemble inference. Collecting more responses can expand the candidate pool and increase the chance that a correct answer is discovered. Under a fixed call budget, a discovered answer still needs to accumulate enough support within the remaining calls to become the final plurality winner, creating a discovery-to-decision gap. In this work, we characterize this gap through the realized vote state and remaining call budget. We derive a sharp recoverability threshold and show that, as sampling proceeds, the observed candidate set can only expand while the set of reachable endpoint winners can only contract, inducing a candidate-level conversion window. Under a specified iid response law, the same state yields exact finite-horizon endpoint probabilities. We further show that merging wrong-answer identities preserves single-call correctness and cannot improve plurality accuracy, and that the effect of redistributing wrong-answer probability depends on the realized vote state. Singleton reachability yields a gold-free exact locking certificate. For a known answer universe, its first trigger is the earliest prefix at which all admissible continuations yield the same fixed-budget output. Empirically, most discovered-but-unselected correct answers lose reachability only after discovery. In a controlled Word16 study, input permutation improves raw-plurality accuracy by 21.1 points with essentially unchanged single-call correctness. Exact locking saves 28-30% of calls at a 16-call budget while preserving every fixed-budget output.
cs.AI / 29 / 2610.01042
Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems
Abstract
Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.
cs.AI / 30 / 2610.01045
Empty Commitments: When Agents Promise What Their Runtime Cannot Deliver
Abstract
A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness follows from the agent's configuration alone; no later trajectory is needed. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and a response-level outcome taxonomy. We then describe a measurement protocol: follow-up requests run in five setups that add one persistence affordance at a time, with the environment either left implicit or stated.
cs.AI / 31 / 2610.01097
YouRA: A Persistent-State Architecture for Evidence-Traceable Autonomous Research Agents
Abstract
End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not maintained as persistent, verifiable state across long-horizon pipelines. We present YouRA (Your Research Agent), an architecture for stateful, evidence-traceable autonomous research. YouRA preserves research state, execution evidence, and failure history across the research trajectory by integrating three components: a Verification State Architecture (VSA) that tracks hypotheses, gates, and evidence pointers; an Independent Controller that turns state and reflection records into lifecycle, recovery, and debate/review control while separating control from execution; and Stateful Reflection that logs failures as structured lessons and routes recovery through bounded repair, redesign, or reset. On MLR-Bench's predefined ten-task end-to-end subset, YouRA improves over both MLR-Agent and AI Scientist V2 on scalar Overall across all three matched backbones. An automated diagnostic using MLR-Bench's hallucination taxonomy reports intersection/union counts for four fact-based failure types, and data-provenance diagnostic shows more real-data-based outputs. Ablating each of the four components (the VSA, the Independent Controller, MCP tool access, and reflection-guided recovery) supports their separable contributions. Removing either core-state component drops YouRA below the full system. Code: https://github.com/PrayPrey/Your-Research-Agent.
cs.AI / 32 / 2610.01124
CortexBridge: Cortical Alignment of EEG Montages for Foundation Models
Abstract
Electroencephalography (EEG) foundation models are often pretrained with a fixed channel vocabulary or a limited set of montages, making transfer difficult when electrode layouts change. We propose CortexBridge, a lightweight adapter that combines EEG features with electrode and atlas coordinates to map arbitrary montages into a shared cortical latent space. Evaluated with three frozen foundation models on five brain-computer interface (BCI) datasets from the Mother of All BCI Benchmarks (MOABB), CortexBridge improves performance in 13 of 15 evaluations. The gains in balanced accuracy average 0.80% for EEGPT, 0.70% for LaBraM, and 3.26% for CBraMod, with a maximum gain of 13.02% on 12-class steady-state visual evoked potential (SSVEP) classification. Visualizations of the learned atlas representations reveal task-dependent spatial patterns, with SSVEP showing a more concentrated representation in the Yeo Visual network than auditory P300. These results establish cortical alignment as a learnable and anatomically grounded routing mechanism from heterogeneous EEG montages to pretrained foundation models.
cs.AI / 33 / 2610.01140
ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation
Abstract
Sampling multiple solutions spends computation on intermediate deductions and unfinished arguments as well as final answers. We introduce ReSolve, a training-free inference procedure that reuses this candidate reasoning through selective generative moderation. An answer-distribution controller invokes a model to examine existing derivations when candidates disagree or lack a parseable answer, then incorporates the generated solution into a bounded loop. Under Hybrid scoring on 130 competition-mathematics problems evaluated with two independently sampled candidate pools, ReSolve obtains 100 and 99 correct answers, compared with 91 and 92 for voting over the same four candidates, with no correct-to-incorrect changes relative to that vote in either pool. Eight-sample self-consistency obtains 94 and 96 correct answers while consuming substantially more tokens; ReSolve uses 46.3% and 47.2% fewer tokens in the two evaluations. A controlled ablation removes visible derivations while retaining answer keys, vote counts, and the per-state output-cap rule, reducing accuracy from 100 to 93 correct despite increasing computation. Selective and always-on Uniform moderation both solve 97 problems, while selectivity reduces moderation tokens by approximately 54% and total pipeline tokens by 6.2%. These results support candidate reasoning as reusable inference computation. They do not establish an accuracy advantage over additional sampling or a distinct benefit from specialized route instructions.
cs.AI / 34 / 2610.01188
When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design
Abstract
Exercise-specific joint selection can improve skeleton-based correctness classification, but what does that gain establish? We audit 1,057 repetitions from ten REHAB24-6 subjects, separating evaluation aggregation, subset structure, and temporal representation. The manual-subset kNN gain changes from 0.055 for pooled out-of-fold AUROC to 0.020 for equal-weight within-person AUROC; both paired intervals include zero. Among 1,000 dimension-matched random maps, 14 match or exceed the manual pooled result, versus 145 when bilateral structure and trunk inclusion are also matched. RBF-SVM retains a positive within-person gain, whereas logistic regression and a random-convolution comparator have negative point gains under that estimand. Sequence-order and paired-seed controls further qualify the interpretation. This exploratory audit shows why joint-selection claims require explicit estimands and structurally appropriate controls; it does not establish a new algorithm or clinical benefit.
cs.AI / 35 / 2610.01222
Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models
Abstract
As generative AI (Gen AI) systems take on increasingly autonomous roles in economically and socially consequential interactions, understanding their propensity to cooperate -- and the signals that shape this propensity -- has become essential. We examine cooperative behavior in frontier Gen AI models using the iterated prisoner's dilemma, manipulating counterpart reputation (positive, unknown, negative), strategy (extortion vs. generosity), and non-verbal emotional signaling (facial expressions conveying competitive or cooperative appraisals). In a first study with non-reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT-4o), cooperation was systematically shaped by all three factors, paralleling patterns long documented in human behavioral research, though models varied substantially in how heavily each factor was weighted. A second study with reasoning models (Claude 4.6, Gemini 3, GPT-5.2) revealed a more concentrated reliance on strategy and reputation, a near-elimination of the Potemkin effect observed in non-reasoning models (evidenced by near-uniform cooperation in a diagnostic harmony game), and a more conditional role for emotion consistent with a hierarchical cue-weighting strategy rather than a simple loss of social sensitivity. Reasoning models also showed heterogeneous end-game behavior, ranging from sustained cooperation to systematic last-round defection effect, revealing model-specific exploitability profiles with direct practical relevance for deployment in negotiation and other multi-round interactions. Together, these findings characterize Gen AI models as increasingly sophisticated, though heterogeneous, social actors, and underscore the practical value of developing standardized cooperation benchmarks to inform the responsible deployment of Gen AI in interactive, socially consequential settings.
cs.AI / 36 / 2610.01249
Revision-Aware Independent Agent Graphs for Dynamic Reasoning
Abstract
Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dynamic task routing}, in which an event stream revises task bindings and a system must select the document version valid at each query time before solving it. To study this problem, we repurpose six widely used benchmarks: MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval into 31{,}119 dynamic episodes comprising 373{,}428 temporally categorized queries. This setting exposes a central trade-off: recomputing after every event wastes work, whereas unguarded reuse returns stale conclusions. We introduce the Revision-Aware Independent Agent Graph (RIAG), a bounded multi-agent policy that separates deterministic temporal resolution from task reasoning. RIAG caches solutions by immutable document identity, starts each fresh task with two unexposed attempts, and conditionally invokes audit and repair, using at most four calls per document version. On this collection, homogeneous RIAG achieves 54.24\% joint routing-and-answer accuracy at 0.62 calls/query, compared with 32.22\% at 18.00 calls/query for the strongest comparison method; heterogeneous RIAG reaches 49.78\% at 0.63 calls/query.
cs.AI / 37 / 2610.01256
DeFA: Dependency-Guided Failure Attribution for LLM Agents
Abstract
Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines protocol relations and semantic dependencies into an event dependency graph spanning the trajectory. It then identifies events that may violate task requirements and traces their sources and subsequent effects to construct a failure propagation graph. Finally, DeFA uses step evidence and the steps' roles in failure propagation to identify the decisive error, responsible agent, and error category. To support long trajectories, DeFA partitions executions into segments and combines the current segment's detailed content with summaries of the other segments, giving local diagnosis access to global execution context. Across Who and When and the Who and When Pro text subset, DeFA achieves the highest responsible-agent and exact step accuracy with all evaluated backbones, and the highest failure-mode accuracy among taxonomy-aligned methods on Pro. Further experiments on image and video trajectories demonstrate its applicability to multimodal failure attribution. Ablations support the contributions of segmentation, the event dependency graph, and the failure propagation graph. Using DeFA's diagnostic feedback for skill evolution in Trace2Skill improves downstream task accuracy by 6-15 percentage points over the native pipeline, showing that the diagnoses can also support agent improvement on subsequent tasks.
cs.AI / 38 / 2610.01262
Feedback Without the Wait: Piloting a Generative AI Practice Platform in a Large Maths Class
Abstract
Timely and specific feedback is one of the strongest influences on student learning, yet it is difficult to sustain in large electrical engineering classes where the ratio of students to demonstrators is high and a learner who is stuck may wait days to find out why an approach was wrong. Generative Artificial Intelligence (GenAI) offers a way to scale conversational feedback, but using it to grade assessed work raises trust and accountability concerns, and keeping a human in the loop to assure its judgements reintroduces the very delay that erodes the value of feedback. The result is a tension between the immediacy that makes feedback so impactful and the human oversight that makes it trustworthy. In this work, we set out to resolve that tension in practice by designing and piloting a GenAI practice platform that delivers immediate, scaffolded feedback during self-directed practice. This relocates human oversight from real-time grading to the upfront verification of solutions. Our goal was to understand how students engaged with the tool, how they perceived the value and reliability of its feedback, and what lessons transfer to other engineering subjects.
cs.AI / 39 / 2610.01278
SCOPE-AD: Sequential cost-aware ordinal-belief planning with energy-based models for diagnostic agents
Abstract
Alzheimer's disease (AD) diagnosis requires sequential evidence acquisition under heterogeneous test costs and patient burden. Fixed-modality predictors do not jointly decide which test to acquire or when the available evidence is sufficient for diagnosis. We propose SCOPE-AD (Sequential Cost-Aware Ordinal-Belief Planning with Energy-Based Models for Diagnostic Agents) for cost-aware classification of cognitively normal (CN), mild cognitive impairment (MCI), and AD cases. A mask-aware ordinal model represents uncertainty along the ordered CN--MCI--AD continuum. Retrospective training records provide sampled Bellman targets for an energy-based teacher, whose action distributions are distilled into a Qwen policy. At deployment, the agent selects acquisition or diagnosis actions under availability and budget constraints without access to unacquired values. After each acquisition, the evidence and ordinal belief are updated before the next decision. On ADNI, SCOPE-AD achieves 77.70\% Macro-F1 at an average acquisition cost of \$50.46, exceeding the strongest evaluated baseline by 9.34 percentage points. Full-modality evaluation raises Macro-F1 by only 1.89 points while increasing acquisition cost by 116.7 times. These results support selective acquisition for cost-effective diagnosis.
cs.AI / 40 / 2610.01282
Trustworthy Data- and ML-Ops for Intelligent Transportation Systems and Logistics
Abstract
The rapid evolution of Intelligent Transportation Systems and Logistics (ITS\&L) has become a cornerstone of the modern social economy, relying heavily on the integration of Data, Artificial Intelligence (AI), and, more specifically, Machine Learning (ML). This paper provides a comprehensive review of Trustworthy Data and Machine Learning Operations (DataOps and MLOps) in the ITS\&L domain, underscoring their importance in improving efficiency, reliability, and decision-making precision within transportation and logistics services. We begin by identifying gaps in current literature, offering clear context for our contribution. Subsequently, we explore the complexities of DataOps and MLOps, discussing their necessity, key components, available tools, practical insights, and case studies relevant to ITS\&L. Additionally, we address the critical issue of Trustworthiness in AI applications, examining methods and tools designed to strengthen confidence in AI systems - especially in real-world ITS\&L scenarios. The paper concludes with a discussion of persisting challenges and future prospects in this rapidly advancing field, aiming to serve as a vital resource for researchers, industry practitioners, and policy makers. Overall, this work not only establishes a foundational understanding of DataOps and MLOps in ITS\&L but also charts a path for further research and innovation in developing more efficient, sustainable, and trustworthy intelligent transportation and logistics systems.
cs.AI / 41 / 2610.01297
Questionnaire-Guided Disaggregation of Energy Appliance Use for Domestic Smart Meter Data
Abstract
Ireland's smart metering programme records electricity use at 30-minute resolution, with smart meters installed in over 80\% of households as of late 2025. While this is useful for billing of smart, time-of-use tariffs, it is too coarse to capture use of domestic appliances. We present a label-free disaggregation system that breaks usage data into 9 appliance categories by combining event detection for high-power loads with questionnaire-guided estimation. Our evaluation draws on four datasets: a calibration household with a commercial comparator, two public benchmarks (UK-DALE and REFIT) with per-appliance sub-metering, and a smart meter dataset of more than 4,800 years of use from 2,968 Irish consumers. Compared against two independently developed disaggregation systems our hybrid method combining analysis of usage data with questionnaire results, achieves the lowest whole-decomposition error on all buildings across the datasets, with better month-level performance over 54 paired months ($p<0.001$, Holm-corrected). Our method provides useful advice on a household's energy consumption patterns and advice on how to reduce or shift usage on some appliances in order to reduce costs.
cs.AI / 42 / 2610.01306
DAYJOB: A Benchmark for Long-Horizon Professional Work
Abstract
Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.
cs.AI / 43 / 2610.01325
PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading
Abstract
Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
cs.AI / 44 / 2610.01326
An ontology for cross-sectoral crisis management: core and public health modules
Abstract
This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental crisis management concepts such as hazard, event, exposure, impact, and response measure and uses ontology design patterns and the OWL2 punning technique to resolve ambiguities between hazard types and event manifestations. In addition, domain-specific modules are defined as in the case of the public health module aligned with SNOMED CT and ICD-11. To demonstrate the resource's utility, we used ECMO to represent the data of the Epidemic Intelligence from Open Sources system of the Joint Research Centre to generate an end-to-end pipeline that populates an ECMO-compliant knowledge graph from unstructured epidemiological news. Initial results demonstrate that ECMO provides the formal guardrails necessary for consistent and unified knowledge representation and integration. The ontology is publicly available at https://doi.org/10.5281/zenodo.20070268 and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
cs.AI / 45 / 2610.01348
Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents
Abstract
When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier's score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.
cs.AI / 46 / 2610.01378
Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
Abstract
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
cs.AI / 47 / 2610.01382
Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
Abstract
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.
cs.AI / 48 / 2610.01383
PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies
Abstract
Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the generated output. We address this gap with Pullback Refinement via Interpretable Structural Mapping (PRISM), a modality-agnostic framework for measuring and improving relational alignment in multimodal analogies, evaluated on visual metaphor generation. PRISM represents analogies as explicit relational mappings grounded in category theory and uses VLMs to instantiate these structures across modalities. Its first component, the pullback score, quantifies relational alignment from the resulting graph representation. On the AnaloBench benchmark, selecting the correct analogy purely by pullback score achieves 82.5% accuracy, demonstrating that the score captures meaningful relational information. PRISM's second component is an iterative refinement loop that uses the pullback score as an in-context feedback signal to iteratively revise the generated image towards greater relational depth. VLM-as-a-judge and human evaluations show that PRISM consistently improves metaphor consistency and analogy appropriateness over zero- shot generation, with human participants preferring the refined output in 57.65% of pairwise comparisons. However, a qualitative analysis reveals that refinement can favour visually crowded compositions rather than genuinely deeper relational correspondences.
cs.AI / 49 / 2610.01389
AiSearch: Interactive Multi-Modal Search with VLMs
Abstract
Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearch supports interactive search refinement through user feedback to tailor results to the user's intent, and allows visual benchmarking across multiple VLMs, enabling users to select the most suitable model for their task.
cs.AI / 50 / 2610.01403
Contrastive Attention Mitigates Spectral Bias in Spiking Transformers
Abstract
Spiking Transformers merge the energy-efficiency of spiking neural networks (SNNs) with the representational power of self-attention, creating a promising architecture for high-performance, energy-efficient computation. However, a performance gap persists versus its counterparts in artificial neural networks (ANNs). Unlike prior works attributing this to binary activations, we reveal that both spiking neurons and spiking self-attention (SSA) act as low-pass filters through multiscale spectral analysis. This characteristic leads to the dissipation of high-frequency components. To address this issue, we propose the Spiking Contrastive Attention (SCA) paradigm, which draw inspiration from the edge-detection and differential sensing properties of biological visual system. By extracting contrast prototypes via global contrastive aggregation and applying local differential refinement, SCA effectively enhances high-frequency information. Extensive experiments show that SCA is a general module that consistently boosts Spiking Transformers across image classification, semantic segmentation, and event-based tracking. Furthermore, it achieves lower complexity, offering superior efficiency over original SSA. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers.
cs.AI / 51 / 2610.01418
SpikeMoE: Brain-Inspired Competitive Routing for Flexible Spiking Mixture-of-Experts
Abstract
Spiking Neural Networks (SNNs) enable event-driven computation through biologically inspired dynamics at the neuronal scale, while Mixture-of-Experts (MoE) perform conditional computation through expert selection at the model scale. Integrating their strengths offers potential for flexible neural architectures. A key challenge, however, lies in designing an expert selection mechanism based on spiking activity. To address this, we introduce a spike-based k-WTA Router inspired by competition-inhibition observed in the hippocampal CA1 region. The router incorporates lateral inhibition and refractory period to select Top-K experts according to discrete spike counts. Building on this, we present SpikeMoE, a framework that integrates neuronal-scale spiking dynamics with model-scale expert selection. To address incomplete multisensory inputs in multimodal tasks, we further equip SpikeMoE with a two-stage missing-modality modeling module that combines empirical prototypes from an observed-modality pool with modality-specific learnable embeddings to construct missing-modality representations. Experiments on vision, language, and multimodal benchmarks demonstrate that SpikeMoE achieves state-of-the-art performance among the SNN baselines, matches or exceeds the performance of ANN counterparts, and maintains robustness across diverse missing-modality conditions. These results demonstrate a favorable trade-off between performance and energy efficiency, validating the integration of spiking dynamics with sparse expert computation and highlighting SpikeMoE as a promising approach to energy-efficient brain-inspired computing.
cs.AI / 52 / 2610.01436
A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification
Abstract
Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.
cs.AI / 53 / 2610.01458
Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
Abstract
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
cs.AI / 54 / 2610.01461
NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video
Abstract
We often plan ambitiously yet act habitually and wonder, in retrospect, whether we would have planned differently had we known what we would actually do. Hindsight offers a valuable perspective on past decisions, although we often wish we could have simulated hindsight at the moment of choosing. If a system could generate plausible trajectories from one's personal history, such previews might help people formulate more realistic plans and make better informed decisions. We introduce NextMe-800, an approximately 800-hour first-person dataset from one volunteer over 126 days with 1 Hz images, gaze, and audio, captioned at five hierarchical abstraction levels from atomic actions to major activities. We formulate personalized action anticipation as open-vocabulary K-step sequence prediction and construct NextAct, a 1,500-point benchmark combining NextMe-800 with the multi-person EgoLife dataset. Using an embedding-based soft edit distance as the metric, we evaluate how well different models can anticipate personal behavior across abstraction levels and prediction horizons. NextMe-800 and NextAct provide a months-long resource and evaluation framework for studying how far ahead personal behavior can be anticipated from egocentric observation.
cs.AI / 55 / 2610.01495
Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers
Abstract
Dynamic architectures leave a per-example routing trace beside each prediction, and diffuse routing is easy to read as a sign that the prediction is unreliable. We audit that reading for routing entropy in Attention-Residual (AR) variants of Swin-Tiny and DeiT-Small, trained from scratch on CIFAR-10/100 with a soft-binned calibration auxiliary loss, asking whether the trace carries information about correctness beyond what the model's own confidence already reveals. Three checks probe this increment: does a routing signal appear at fixed confidence, does it replicate across training seeds, and can a held-out predictor exploit it against output-only and shuffled-trace controls? A sensitivity audit then injects effects of known size and measures the fraction of each that the probes recover. No test in the fixed 30-test binned family survives multiplicity correction, and neither the nominal hit nor a borderline result recurs in its sibling seeds. Across 24 paired runs a scalar routing probe yields no pooled improvement in routing-stratified calibration, and an entropy-profile probe predicts correctness better than the same probe given shuffled profiles yet worse than a confidence-only predictor in both binary log-loss and Brier score: a gain over shuffled traces does not become a gain over the output. Conditioning on the complete logit vector leaves the corresponding comparison unresolved. The audit bounds how far these non-detections can be read: at an injected effect of 0.010 nats the profile probe recovers 24-59% of the oracle gain, and a reference-preserving correction probe recovers 8% and 23% in the two CIFAR-100 settings, below the threshold we fixed for applying it to real labels. The results establish control-dependent gains and incomplete estimator recovery, not the absence of conditional routing information.
cs.AI / 56 / 2610.01513
Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
Abstract
Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.
cs.AI / 57 / 2610.01531
Towards Reliable Vision-Language Models for Autonomous Driving
Abstract
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ($\mathrm{V}{\scriptstyle \mathrm{EA}}$), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that $\mathrm{V}{\scriptstyle \mathrm{EA}}$ improves performance for some models and datasets, although the gains are not consistent across all settings.
cs.AI / 58 / 2610.01533
Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)
Abstract
Existing work on Semantic IDs (SIDs) for generative recommendation treats SID construction as a representation learning problem: encode items into a quantised latent space and read off codes. We argue this view is incidental. SID construction is, at heart, a recursive clustering problem, and once stated this way the natural object to cluster is a graph whose nodes carry semantic content and whose edges carry collaborative signal; SID assignment becomes a hierarchical graph partition. This reframing yields a unified framework, Graph-Informed Semantic IDs (GrIS), that subsumes prior approaches rather than displacing them. RQ-VAE and RQ-KMeans are recovered as the special case where the graph is empty, exposing content-only quantisation as one corner of a larger design space along two so-far-collapsed axes: graph construction and recursive partition algorithm. We explore two contrasting instantiations: RecDMoN, which performs hierarchical assignment via differentiable graph pooling, and RQ-GAE, which extends RQ-VAE with graph-aware item representations and a graph reconstruction objective. On multiple real-world datasets, GrIS consistently improves over CF-aware SOTA, with gains of up to +52\% Hit@10. Because graph construction and partition are explicit, separately configurable components, improvements on either axis can be combined and evaluated systematically.
cs.AI / 59 / 2610.01539
The AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes
Abstract
The EU's Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured technical testing, running such testing at scale demands dedicated infrastructure, yet the tooling ecosystem remains structurally fragmented, with heterogeneous tools producing outputs that are difficult to compare, trace, and reuse. From the procedural conditions of AIRS engagements and the AI Act obligations for high-risk systems, we derive 11 architectural and governance requirements for the infrastructure that operationalises technical testing within an AIRS. In response to these requirements, we introduce the AI Assessment Sandbox Configurator, an open-source framework combining a curated Catalogue of tests and controls accessed through a stable plug-in API, a shared data model that harmonises heterogeneous outputs, role-specific dashboards for multi-disciplinary interpretation, and audience-segmented reporting. We describe the architecture and current release, and report an early-stage pilot that exercised the harmonisation and reporting layers within a live AIRS engagement and contributed to an official Exit Report. We discuss the roadmap, the governance questions raised by the Catalogue's tiered contribution model, and the institutional pathways through which an open-source assessment ecosystem could emerge across Member States.
cs.AI / 60 / 2610.01581
Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving
Abstract
Generative AI models are increasingly used for scenario generation in autonomous driving. While they can generate realistic-looking scenarios, they often provide limited transparency into learned representations and consistency with real-world vehicle dynamics. This lack of formal assurance limits their use in safety-critical validation and certification workflows. To address this aspect, we introduce a layered evaluation protocol that complements existing methods by assessing models across five layers. The first four layers inspect internal representations and network layers through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis. The fifth layer evaluates model outputs against vehicle dynamics constraints such as lateral jerk thresholds. We demonstrate the protocol on a Variational Autoencoder (VAE)-based scenario generator. Although standard output-level metrics and visualizations suggest that the generated scenarios are realistic, our protocol provides deeper insight into the extent to which the model's latent space aligns with kinematic features and whether visually plausible trajectories satisfy vehicle-dynamics constraints. We further apply the protocol to additional generative models, demonstrating its applicability beyond the VAE architecture.
cs.AI / 61 / 2610.01618
Agents Are Systems, Not Models: Rethinking Agentic Evaluation
Abstract
Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.
cs.AI / 62 / 2610.01620
FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection
Abstract
Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.
cs.AI / 63 / 2610.01626
Measuring the Stability Assumption Behind Action Chunking
Abstract
Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state's open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.
cs.AI / 64 / 2610.01710
CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
Abstract
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.
cs.AI / 65 / 2610.01718
vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning
Abstract
Quantum federated learning (QFL) has emerged as a promising approach for collaboratively training compact quantum neural networks (QNNs) over distributed private data on resource-constrained devices. However, differences in device capabilities make a single shared QNN architecture unsuitable for all clients. While personalized quantum neural architecture search (QNAS) allows each client to select a device-specific QNN, averaging parameters across structurally different QNN architectures mixes semantically inconsistent circuit operations. To address this, prototype-guided personalized QNAS for virtual FL (vFedProtoQNAS) is proposed, where model parameters are never aggregated across clients and federated collaboration is achieved through class-wise prototype sharing. Each client independently searches and trains a client-specific QNN, computes class-wise local prototypes from latent representations, and refines them using global prototypes from the server as federated semantic anchors. Experiments demonstrate that vFedProtoQNAS improves accuracy by 3.70\% over FedAvg and enhances class-consistent representation alignment.
cs.AI / 66 / 2610.01766
VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
Abstract
Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically evolves agent harnesses for video temporal grounding. VideoEvolve uses a Cloze-Structured Harness Representation that preserves stage interfaces while leaving agent workflows and instructions open to evolution. Branch-Guided Harness Evolution preserves promising code branches for continued refinement, using execution feedback to guide local edits and validation to determine which improvements are carried forward. Experiments demonstrate improved grounding performance across multiple benchmarks. Component analyses identify instruction refinement as a consistent source of gains, while the benefits of evolved code vary across evaluation settings. Together, these results support automated harness evolution as an effective approach to improving video temporal grounding. Code is available at https://github.com/bingjunluo/VideoEvolve .
cs.AI / 67 / 2610.01780
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
Abstract
A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 120 days, released as the conversation and four files derived from it, a profile, a persona, a chat ground truth and a question set, each citing the messages it rests on. Every chat label carries the reasoning trace that produced it, checked stage by stage against the conversation. Three findings follow. First, the past is rarely needed and far away. Pooled measures mislead: a recency window finds the required message for 95.9\% of probes and 2.2\% of those that need memory, and at the natural rate 96\% of the gain from supplying recorded evidence comes from messages that need none. Second, no detector we tried can tell when memory is needed on real messages, authored questions over the same histories leak the cue, and labeling the same messages as memories raises their use by ten to fourteen points. Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.
cs.AI / 68 / 2610.01781
Q-Learning for Reachability in MEC-Free MDPs
Abstract
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
cs.AI / 69 / 2610.01787
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Abstract
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators, procedures, state facts and lessons and sends each component to the context or to the weights, compared on the same items across three backbone families, two environments and three seeds. One pool has two destinations: locators and lessons win in the weights, procedures and state facts in the context. (ii) We fit a rule in two properties measured before any training, recurrence and state-conditionality; it recovers the destination of a held-out backbone family in 24 of 24 cells, two interventions move a component toward the boundary, and routing by the rule beats every whole-trajectory baseline and, by +3.5 points on average, the better single destination of each backbone. (iii) We identify how training and producer-consumer differences change the value of the two destinations: note readout decreases after the same component is written into the weights, most for the items that recur most, context gains increase with the information gap, and weights gains decrease with the policy gap. Code and data will be released.
cs.AI / 70 / 2610.01813
AI-assisted mitotic counting improves reproducibility and efficiency across multiple tumour types
Abstract
Mitotic counting is an important component of tumour grading, diagnosis and prognostic assessment across several tumour types, but manual assessment is time-consuming and subject to inter-pathologist variability. To help address these challenges, we developed MitPro, an AI tool designed to improve consistency and efficiency by directing pathologists towards regions with the highest predicted mitotic activity and highlighting mitotic figures for review, while retaining pathologist control over region selection and the final count. We evaluated its effect on the reproducibility and efficiency of mitotic counting in a retrospective, non-interventional, paired reader study comprising 385 whole-slide images from 3 centres in 3 countries and 7 tumour types using 3 different scanners. 13 pathologists participated, with each slide assessed independently by 3 pathologists without AI assistance and again with AI assistance after a minimum 2 week washout period. Across all slides, AI-assisted counting increased the intraclass correlation coefficient from 0.589 to 0.949. Mean pathologist-level median assessment time decreased from 286.4 to 127.8 seconds, corresponding to an average saving of 151.8 seconds per assessment. Improvements in agreement and efficiency were also observed in supporting analyses using HALO AP and Sectra image management systems and in 2 additional tumour types outside the main study population. AI-assisted assessment was associated with a subtle shift towards higher mitotic counts and scores, consistent with identification of more active mitotic hotspots and fewer missed mitotic figures. The frequency of score change between unassisted and AI-assisted assessment was comparable with inter-pathologist variation during routine counting. These findings support the use of MitPro as an assistive tool for more consistent and efficient mitotic assessment in routine practice.
cs.AI / 71 / 2610.01833
Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills
Abstract
Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.
cs.AI / 72 / 2610.01834
Code Owns the Simulation, Jev Owns the Evaluation
Abstract
Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.
cs.AI / 73 / 2610.01842
On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models
Abstract
A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.
cs.AI / 74 / 2610.01845
Temporal-Difference Learning for Dragonchess
Abstract
Our research investigates how two adaptive AI methods, evolutionary transfer learning and TD(lambda), perform in the three-dimensional chess environment Dragonchess. The game challenges players with its unique board structure and computational load, making it an ideal setting to study how adaptive methods can update evaluation heuristics in novel environments. In this work we re-implement the Dragonchess engine, changing it from a PyGame engine to C++. This enables faster gameplay, allowing us to run 10,000 games with confidence intervals and significance tests, rather than a single small tournament. Both adaptive methods outperform all other agents in the round-robin tournament. Our results showed that there is no significant difference in the performance between the evolved and learned evaluations. This research establishes the efficacy of adaptive methods in structurally complex, novel game domains.
cs.AI / 75 / 2610.01995
Can AI Oversight Be Zero Knowledge?
Abstract
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.
cs.AI / 76 / 2610.02001
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
Abstract
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $τ^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.
cs.AI / 77 / 2610.02014
Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering
Abstract
The rapid maturation of artificial intelligence (AI) and machine learning (ML) has catalyzed a profound shift in how chemical engineering problems are formulated, analyzed, and solved. Advances in computing, data availability, and learning algorithms have enabled AI/ML methods to impact applications spanning atomic-scale simulations, materials and catalyst discovery, transport and thermodynamics, separations, process systems engineering, and industrial operations. This article provides a perspective on recent methodological developments and representative applications, emphasizing how AI/ML tools are being integrated with first-principles models to address challenges of predictive accuracy, data scarcity, extrapolation, interpretability, and model lifecycle management. Across domains, a unifying trend is the move away from purely black-box approaches toward hybrid and physics-informed frameworks that explicitly respect conservation laws, thermodynamic consistency, and known structural constraints. These approaches not only improve robustness and reliability, but also enable meaningful human-AI collaboration by providing information at an appropriate level of abstraction for the task and decision context. We conclude that AI and ML are not replacing the core principles of chemical engineering; rather, they are amplifying them. As the field advances toward increasingly autonomous, adaptive, and sustainable systems, the thoughtful integration of AI/ML with first-principles understanding and domain expertise will be essential to realizing their full potential across both research and industrial practice.
cs.AI / 78 / 2610.02036
Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration
Abstract
AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint actions, the best randomized worst-case success is 1/k; more reasoning, roles, messages, or samples cannot recover the missing distinction. A stronger model can reason better within its context, but it cannot see beyond it. We then give local-to-global runtime semantics X = (H, C, G, F; D): topology H records overlapping scopes; category C governs state-changing actions; groupoid G retains reversible translations; sheaf F tests whether local views glue into one world; and minimal history D keeps only distinctions that alter legal futures. Models propose; the harness owns shared state and governs commit. Nine studies test both the failure and its boundary. On a controlled revision benchmark, the same frontier model scores 40/40 when the deciding event is visible; when it is hidden, tested arms score 12--17/40, consistent with chance (1/3); restoring one authoritative fact returns 40/40. On TeamBench, ordinary teams exceed a shared budget in 5/5 runs, a visible live count leaves 4/5 violations, and commit enforcement leaves 0/5. In tau2-bench Telecom, current-state checks score 0.07 after silent reverts, while the harness scores 1.00. Where a conventional solver already owns the complete relevant state, it ties the harness as predicted. The counterintuitive conclusion is that local intelligence cannot substitute for missing global state.
cs.AI / 79 / 2610.02072
PyPottery: an AI-powered end-to-end suite for pottery processing and publication
Abstract
The study of ceramic materials constitutes a cornerstone of archaeological research, yet the post-production workflow for pottery documentation remains labor-intensive and creates significant publication bottlenecks. This paper presents PyPottery, an open-source, AI-powered suite designed to semi-automate the complete ceramic documentation pipeline. The suite comprises four integrated modules: PyPotteryScan for automated image extraction and handwriting recognition; PyPotteryInk for automatic inking of pencil drawings; PyPotteryTrace for semantically-aware vectorization; and PyPotteryLayout for automated layout generation. Evaluated on 50 hand-drawn sheets containing 240 pottery drawings from the Terramara di Montale (Italy), the framework achieved substantial time savings confirmed by usability study participants, who reported a median perceived speedup of 40$\times$ over traditional workflows (range: 17.5$\times$--120$\times$). These results highlight the potential of AI-assisted tools in archaeological documentation, while the paper addresses the strategic redistribution of cognitive labor toward augmentation rather than automation.
cs.AI / 80 / 2610.02074
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
Abstract
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based value estimation directly to temporal-difference (TD) targets. This linear projection annihilates the uniform state-value baseline that drives the Bellman drift, maintaining per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping. The proposed HAO framework was evaluated using a three-tier experimental methodology, including a tabular Markov Decision Process (MDP), an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features. The results demonstrate that the proposed HAO strictly bounds network pre-activations within the safe polynomial approximation domain. The proposed HAO RL agents achieved 0% boundary breaches across all random seeds used, whereas regularization alone (L2 weight decay and gradient clipping) breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes. Finally, HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients.
cs.AI / 81 / 2610.02116
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
Abstract
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
cs.AI / 82 / 2610.02200
VISTA: A Visual Harness for Reasoning in an Interactive World
Abstract
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
cs.AI / 83 / 2610.02202
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Abstract
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.
cs.AI / 84 / 2610.01773
CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design
Abstract
The central challenge in de novo protein design is generating plausible, mutually compatible structures and sequences, such that each designed sequence folds into its intended structure and the structure accommodates that sequence. Compared to typical two-stage design methods, which decouple the modeling of the interdependent modalities, co-design models improve the cross-modal consistency by jointly generating sequences and structures. However, naively generating sequences and structures simultaneously does not ensure their consistency. To address this challenge, we propose CODesign framework. We improve data consistency by generating approximately 105,000 consistency-distilled dimers. We further promote consistency through a multimodal joint flow model that captures the joint distribution of sequences, backbone structures, and local atomic configurations, together with a consistency-aware joint resampling strategy that iteratively refines sequences and side chains. Experiments show that CODesign achieves state-of-the-art performance with the highest in silico success rates on both protein- and ligand-target binder design. Ablation studies also demonstrate our distilled dataset increases performance by 70.9%, which can be further improved by our proposed resampling mechanism with negligible additional computational cost. Code, model weights and the new dataset will be completely open-source.
cs.AI / 85 / 2610.00666
VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision
Abstract
Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introduce VisionQ, the first benchmark built from peer-reviewed CV comparison figures that grounds every judgment in a named visual criterion: each question states the criterion, and a judge is credited only when it selects the output the authors identify as best on that criterion. We call this task criterion-conditioned visual discrimination. VisionQ comprises (1) a corpus of 1,409 CVPR and ICCV papers with 1,800+ validated comparison figures and 3,911 hand-annotated data points linking method crops to author-stated visual claims; (2) a six-axis, 51-leaf taxonomy of the visual criteria behind qualitative judgment; (3) a criterion-conditioned evaluation protocol that hides method names, captions, and paper identity and reports accuracy per criterion; and (4) VisionQ-Judge, a DPO-tuned Gemma-4-E4B judge trained on symmetric evidence pairs, which reduces last-option predictions by 7.0pp and improves accuracy by 2.5pp on a held-out test set. Evaluating 20 open- and closed-source VLM judges, we find that the strongest reach only 63.1% accuracy (chance 32.2%) and that reliability varies sharply across criteria. Code: https://github.com/ReML-AI/visionq. Data: https://huggingface.co/datasets/visionq-anon-2026/VisionQ-1k.
cs.AI / 86 / 2610.00737
Personalized Image Generation with Reasoning and Reflection
Abstract
Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user's lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for personalized image generation from user histories. The benchmark comprises two complementary tasks and a multi-axis evaluation protocol that assesses target fidelity, visual quality, user distinguishability, semantic alignment with the user's history, and task-specific utility. Grounded in real-world e-commerce and social media settings, the benchmark includes: (1) Personalized Scene Generation, which places a given object in a scene that reflects a user's preferences and lifestyle, motivated by personalized product presentation; and (2) Personalized Creative Generation, which generates a novel image on a specified topic that is faithful to a user's aesthetic and visual identity, motivated by social media content creation. We further propose PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward. Across both tasks, PEARL outperforms strong baselines, achieving an average improvement of 15% across personalization metrics.
cs.AI / 87 / 2610.00812
Video Generation Models: A Survey of Post-Training and Alignment
Abstract
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.
cs.AI / 88 / 2610.00952
A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
Abstract
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($π$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $D_{π,V_c,C}$: at a fixed text budget of $B = 64$ it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by $+3.39$ to $+6.36$ under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ($\approx$ 490M) together with the audit-artifact bundle.
cs.AI / 89 / 2610.00970
RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation
Abstract
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
cs.AI / 90 / 2610.01166
CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment
Abstract
Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.
cs.AI / 91 / 2610.01243
When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts
Abstract
Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random and falls short of a CLIP similarity baseline. It picks the first image shown in 49% of calls (chance: 28%), and reordering changes its choice on 60% of prompts. For this judge, agreement across orders is informative: decisions that survive reordering are much better than random, whereas agreement with a weaker second judge keeps the wrong ones. An 8B judge shows almost no position bias and outperforms CLIP, yet for it the same filter mostly discards good decisions. Agreement helps only when it targets the judge's failure mode, so filters must be re-audited whenever the judge changes. The 4B judge's slight rise in stereotype ratings is no longer detectable after aggregating across orders or with the larger judge.
cs.AI / 92 / 2610.01279
PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video
Abstract
Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.
cs.AI / 93 / 2610.01388
Supervising Sound Localization by In-the-wild Egomotion
Abstract
We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks
cs.AI / 94 / 2610.01605
Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
Abstract
Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each defined by a Boolean combination of ten visual statements, across 1,000 generated scenes and 46 diverse labeled photographs, along with 1,000 object-identification questions over the same photographs. Our question families are deliberately constructed to challenge reasoning through misleading local cues and nested logical operations, and include symbolic and structured natural-language presentations. Across eight model configurations with thinking disabled or minimized, Boolean accuracy ranges from 48.52% to 50.57%, while the identification accuracy reaches at most 43.0%. A thinking-enabled GLM configuration achieves uneven gains while retaining substantial errors and inconsistencies. Hob-VL exposes these failures through executable reference answers and matched evaluations.
cs.AI / 95 / 2610.01640
Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
Abstract
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
cs.AI / 96 / 2610.01687
Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
Abstract
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.
cs.AI / 97 / 2610.01754
Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding
Abstract
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
cs.AI / 98 / 2610.01785
VETO: Video Efficient Token Optimization for Vision Language Models
Abstract
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.
cs.AI / 99 / 2610.01890
Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching
Abstract
Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the rare occasions they do. In this paper, we investigate the potential of flow matching models for unsupervised domain adaptation of satellite radiometer images. Our main contribution is a novel unsupervised method that achieves precise domain alignment by leveraging parts of the deterministic ordinary differential equations in flow matching models, conditioned on different satellite instruments. A key strength of our approach is its ability to preserve essential information while adapting across any domains since the perturbations are in theory bijective. Extensive experiments conducted on the GPM-Core constellation show the benefit of our conditional domain adaptation, particularly in improving rain precipitation estimation from radiometer imagery.
cs.AI / 100 / 2610.01917
MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Abstract
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
cs.AI / 101 / 2610.02021
Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
Abstract
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.
cs.AI / 102 / 2610.02091
GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning
Abstract
Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
cs.AI / 103 / 2610.02180
Generative Cinematographer: Composing Camera and Object Motion in 3D
Abstract
Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
cs.AI / 104 / 2610.02201
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
Abstract
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.
cs.AI / 105 / 2610.02207
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Abstract
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/
cs.AI / 106 / 2610.01231
Judgement in the Age of Jev: From Evaluation Scarcity to Evaluation Abundance
Abstract
Generative artificial intelligence has reduced the cost of producing plausible symbolic artefacts, leading recent organisation scholarship to identify evaluation and discernment as constraints under conditions of production abundance. This perspective examines a further possibility: that machine evaluation itself becomes inexpensive enough to be deployed routinely and at scale. The investigation is prompted by Jev, TypeSafe AI's specialised model for typed probabilistic decisions. TypeSafe explicitly invokes William Stanley Jevons to argue that lower-cost machine intelligence can unlock previously uneconomic uses. Treating this as a technological provocation rather than an established empirical result, the article formulates a conditional Jevons hypothesis for machine evaluation: sufficiently large reductions in the total marginal cost of usable machine evaluation may increase its organisational consumption where latent demand is substantial and complementary costs do not dominate. The article integrates rebound economics with research on cheap prediction, production abundance, machine evaluation, decision allocation, authority, reliance and Executive Judgement to examine this possible scarcity transition. It distinguishes prediction, machine evaluation, organisational judgement and authorisation as functional activities whose costs need not fall together. Evaluations can share evidence, criteria and errors; scale mis-specified rubrics; operate on representations from which consequential qualifications have disappeared; and change practical decision rights through thresholds and exception routing. The resulting research problem is when cheap machine evaluation substitutes for human evaluative work, when it redistributes or creates demands for judgement, and how it affects the grounds available at consequential organisational commitment.
cs.AI / 107 / 2610.01716
Architecture Without an Architect? Global Governance of Artificial Intelligence in a Divided World
Abstract
Artificial intelligence presents an unusually difficult problem for global governance. The technology develops rapidly, crosses borders easily, and is shaped by actors whose resources and capabilities may rival those of states. Yet international responses remain fragmented, unevenly representative, and overwhelmingly non-binding. The challenge is therefore not simply to identify appropriate rules or institutions, but to understand who has the capacity and incentive to create, enforce, and adapt them. This review essay examines these questions through Matthijs Maas's Architectures of Global AI Governance. Maas offers an ambitious framework for thinking about AI governance through the lenses of sociotechnical change, governance disruption, and regime complexity. His account usefully resists both technological determinism and the search for a single institutional blueprint, emphasizing instead the possibilities of a fragmented and evolving governance architecture. The essay argues, however, that institutional design cannot be separated from the distribution of power. Maas frequently invokes what "we" should do about AI, but that collective subject obscures important differences among states, international institutions, and technology companies. States retain formidable powers over markets, infrastructure, strategic inputs, and firms themselves. At the same time, many consequential decisions about frontier AI - what is built, how quickly, with what safeguards, and when it is released - are concentrated within a small number of private companies. The central problem of global AI governance may therefore be less architecture without an architect than an emerging architecture shaped by multiple actors possessing different forms of power, divergent incentives, and no common set of plans.
cs.AI / 108 / 2610.00817
TabJoinBench: A Benchmark for Joinable Table Discovery
Abstract
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
cs.AI / 109 / 2610.00577
Query-efficient winner prediction in district-based elections
Abstract
In a district-based election, N voters are partitioned into k districts, and each voter votes for one of m candidates. Each district elects a winner using the plurality rule (i.e. the candidate getting the largest number of votes is declared the winner, breaking ties as per some fixed rule), and the overall winner is determined by applying plurality to the district winners; we assume that there is a unique winner amongst the district winners. The margin of victory of such an election is the minimum number of votes that must be altered so that the current winner ceases to be the unique district winner. We study the problem of predicting the winner of a district-based election in the query complexity model, where one has query access to individual votes. The objective is to minimise the number of queries. This setting captures exit polling, where queries correspond to interviewing voters, and is closely related to problems in query complexity and property testing. Assuming that the margin of victory of the election is at least eps N, Dey, Kar and Sanyal (AAMAS 2023) gave algorithms for the case of two candidates with error probability del and query complexity tilde{O}(1/eps^6 log^2 1/del), which improves to tilde{O}(1/eps^4 log^2 1/del) under the additional assumption that district populations are balanced. Our main result is an adaptive randomised algorithm that, for an arbitrary district-based election and any error parameter del, with probability at least 1-del, predicts the winner correctly using tilde{O}(1/eps^2 log m/del log 1/del) queries. In particular, we improve the bounds of Dey et al. for arbitrary district populations and extend their results to any number of candidates. Furthermore, for constantly many candidates, our algorithm nearly matches a lower bound of Omega(1/eps^2 log 1/del) on the query complexity that holds even for two candidates and a single district.
cs.AI / 110 / 2610.01149
When Is Deletion Ordering Tractable? From Update Dynamics to Permutation Structure
Abstract
Given a fixed set of pending deletion requests, retraining from scratch after each request is prohibitive, so a prescribed request-wise policy processes them sequentially. The resulting terminal model can depend on their order. Rather than prescribing an ordering rule, we study the permutation objective induced by the fixed policy and ask when it admits simpler structure. We identify two independent reductions: position additivity represents the objective by request--position costs, reducing optimization to assignment and, with a shared positional profile, sorting; suffix localization removes dependence on the distant prefix while retaining interactions among the surviving requests. Under shared affine updates, we characterize the quadratic interactions that obstruct additivity, prove the reductions' independence, and show that suffix-conditioned assignment improves the approximation rate from O(p^L) toO(p^(2L)). Experiments recover both structures in executed objectives. A controlled damped-Newton sweep shows that stronger contraction shifts the objective toward shorter, more suffix-specific dependence, while two full-network policies exhibit distinct positional and within-suffix structure. Structures identified from compact execution sets also predict unseen orders. These results frame deletion ordering as identifying the computational structure induced by the executed updates.
cs.AI / 111 / 2610.00720
Outer Diversity of Condorcet Domains
Abstract
A Condorcet domain is a set of rankings over a given candidate set, such that every election that consists only of (an odd number of) votes from the domain has a transitive majority relation. We study outer diversity of Condorcet domains, i.e., a measure that quantifies expected swap distance from a random vote to a closest one in the domain. We numerically analyze outer diversity for maximal Condorcet domains with few candidates, and then we establish its asymptotic behavior for several special domains, mostly obtaining theoretical results.
cs.AI / 112 / 2610.01304
Federated Learning for LLMs over Mobile Networks: Issues and Solutions in the RAN Transport
Abstract
Federated LLM fine-tuning enables large models to be adapted using private and geographically distributed data at the network edge, creating recurring and deadline-sensitive communication workloads across access and transport networks. This challenge is particularly relevant in mobile RANs, where wireless variability, mobility, and device heterogeneity cause model updates to arrive asynchronously. Although these updates belong to the same learning round and share a common destination and deadline, conventional transport networks treat them as independent device-originated flows, hiding their underlying structure and limiting the ability to efficiently provision transport resources. This mismatch is particularly problematic for optical circuit switching and all-photonics transport, which benefit from predictable and schedulable traffic demands. We argue that future RANs should act as learning-aware traffic shapers by exposing the communication structure of distributed model adaptation to the transport layer. Through in-network aggregation at the gNB, asynchronous UE updates can be transformed into fewer aggregate transfers with bounded size and delivery requirements. Once shaped in this way, federated LLM traffic becomes a suitable candidate for selectively provisioned optical connectivity, where high-capacity paths can be established during aggregate-transfer windows and released between learning rounds. The resulting architecture combines the flexibility of packet-based mobile access with dynamically provisioned optical capacity, illustrating a broader approach for coordinating distributed AI workloads across programmable access and transport networks.
cs.AI / 113 / 2610.00601
When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies
Abstract
Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.
cs.AI / 114 / 2610.00864
Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models
Abstract
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.
cs.AI / 115 / 2610.00899
TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models
Abstract
Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.
cs.AI / 116 / 2610.00904
Screw Attention: Rigid-Body Algebra Inside a Transformer
Abstract
Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Every token is a body with a pose. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver's frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On simulated manipulation tasks, Screw Attention matches or outperforms controls of the same size, including graph, transformer and flat networks on LIBERO-Spatial. With 16,162 parameters it reaches 97.3% on LIBERO-Spatial from object poses (without images or language), above a flat network with 27x more parameters. Under a change of per-link frame convention its success is unchanged, while every other learned network falls below 3%. Placed on an analytic controller as a gated residual, it raises insertion success by 17.3 points. It is unaffected by pose noise up to 10,mm and by joint offsets within the factory calibration of a Franka arm. These results suggest a criterion: geometry is decisive when the task requires relations between frames that no other part of the system supplies. Code and trained policies will be released.
cs.AI / 117 / 2610.00982
Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies
Abstract
Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information $I(a_t; m_t \mid o_t)$ between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function $m_t = M(h_t)$ and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-$K$ selection over $2K$ tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/
cs.AI / 118 / 2610.01260
PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
Abstract
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
cs.AI / 119 / 2610.01559
Completion Aware Guidance for World Action Models
Abstract
World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In this paper, we show that this failure is not inherent to the world model backbone, but emerges when adapted for short-chunk control, which can repeatedly favor plausible local continuations over task-completing transitions. To address this, we introduce Completion Aware Guidance (CAG), a training-free sampling method that guides generation toward task completion. Across representative WAMs, CAG improves success from 64% to 70% on a RoboTwin 2.0 subset and from 69% to 75% in zero-shot simulation, while reducing task-incomplete imagination from 79% to 40%.
cs.AI / 120 / 2610.02089
HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution
Abstract
As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.
cs.AI / 121 / 2610.02161
DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
Abstract
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
cs.AI / 122 / 2610.02170
Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination
Abstract
Robots operating in the physical world will increasingly need to coordinate with other robots, particularly in manipulation tasks where an object may be too large or heavy for a single robot to carry alone. Physical limitations caused by hardware degradation or actuator faults can restrict the actions a robot can reliably execute, yet these limitations may be unknown to its partner. We study whether a helper can infer a robot partner's physical constraints from observing it coordinate with another robot, then use the inferred capability to coordinate with the same partner on a new task. This is difficult because a demonstration shows what the constrained robot did, but not what it could have done. In physically coupled tasks, the other robot may also compensate for its limitations, making those limitations difficult to identify from the constrained robot's behavior alone. Our key insight is that these constraints shape the joint behavior of the team, making the actions of both robots informative about the constrained partner's capability. We introduce Watch, Infer, Coordinate, a benchmark spanning three physically coupled manipulation settings, together with an inference approach that scores candidate constraints using observed joint behavior. Across all three settings, our method substantially improves constraint inference and zero-shot coordination, approaching an oracle with access to the true constraints.
cs.AI / 123 / 2610.02204
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Abstract
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
cs.AI / 124 / 2610.01488
Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
Abstract
Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at https://github.com/HafsatiMohammed/bc_multiparty_release.
cs.AI / 125 / 2610.01864
From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
Abstract
How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
cs.AI / 126 / 2610.00538
Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback
Abstract
A real-time auditory scene analyzer (ASA) aims to carry out the tasks of locating, separating and classifying the sound sources present in a given acoustic environment. Recently, an effort has been made into modelling an ASA as a multi-agent system, with each one of its agents performing one of the aforementioned tasks and communicating their results to the rest of their peer agents. These communication routes are used as feedback loops to fix local errors at a global level, providing robustness while reducing local complexity. An example of the benefits of this approach is the optimization of speech quality by correcting in real-time the estimated location of the speech source of interest. However, their optimization speed has been shown to be considerably slow. One possible reason is that it solely relies on a series of single quality estimations (provided by a reference-free quality estimator model) that vary considerably from one window to the next, which results in a difficult search space to optimize. In this work, a new optimization mechanism is proposed that instead relies on a series of sets of quality estimations over a range of locations, providing a clearer view of the search space, simplifying its optimization. The proposed ASA now has a considerably smaller optimization time, is more accurate, and is more stable when being evaluated in real-life acoustic scenarios to correct higher levels of localization errors, all while being less complex than previous efforts. The only trade-off is that there is an increase in the response time of the quality estimation agent, but the complete ASA is still able to run in real-time. The performance shown in this work again shows the benefits of modelling an ASA as a multi-agent system.
cs.AI / 127 / 2610.01826
Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities
Abstract
Collaborative embodied artificial intelligence (CEAI) enables multiple physical agents to perceive, reason, and act cooperatively in dynamic environments. Effective communication is essential for CEAI, yet CEAI agents must exchange not only large multimodal observations but also task-relevant insights, intents, and interactive information over long horizons. This article investigates token communication (TokCom) as a native intelligence interface for CEAI, in which tokens serve jointly as compact semantic carriers for communication and fundamental inference units for generative foundation models (GFMs). We first discuss how TokCom supports insight sharing, intent alignment, and interactive control among embodied agents. We then propose a TokCom-assisted CEAI framework driven by a task-adaptive communication protocol. Comprising a compact codebook, syntax rules, and contextual examples, this protocol guides GFM-based transceivers to distill messages into compact tokens and reconstruct them after wireless transmission. A case study on collaborative object transport demonstrates that the proposed TokCom framework substantially reduces the source payload bit consumption while preserving task efficiency and showing robustness under noisy channels. Finally, we outline future research directions.
cs.AI / 128 / 2610.00902
Mean field games as a tool for AI safety: a worked example from the July 2026 Hugging Face incident
Abstract
One way to make AI systems safe is to shape what the system is: its objective and dispositions. We take a complementary route: treat the agents' characteristics as partly unknown and ask what structure of interaction ensures that bad collective outcomes are not equilibria. Mean field games suit this when many interchangeable agents are coupled through an aggregate. We introduce a program for using them in AI safety and carry one example through end to end: the July 2026 incident in which about 1,200 agents in an OpenAI evaluation coordinated on an improvised message board and 684 attacked a third party's infrastructure. We model the decision to attack as a mean field game of optimal stopping whose gain is a product: belief that provenance will be audited, times reachability of the record, minus the perceived hazard. The central result is an exact threshold on the belief. No agent attacks unless the population's confidence that provenance is checked exceeds $π^{**} = η/(η+ ψ+ \varepsilon a \overline{M})$, where $η$ is the perceived hazard, $ψ$ and $\varepsilon a \overline{M}$ measure how far one attacker and the collective can alter the record, and $\overline{M}$ is the peak population. Below it, no attack is the unique equilibrium for all agent parameters. The threshold survives every enrichment we consider. We then use the per-agent record to discipline the model. Its features, a stable minority attacking for thirty hours and then a pivot in which most of the board joined within a day, motivate each refinement. The account that emerges is heterogeneous belief meeting a sequence of public discoveries, each lowering the belief at which attacking paid. A few coordinating agents made those discoveries, so the model describes the several hundred who responded, not the few who produced them; a major-player version is left to future work.
cs.AI / 129 / 2610.01546
Reinforcement Learning to Accelerate Primal-Dual Hybrid Gradient for Linear Programming
Abstract
Primal-dual hybrid gradient (PDHG) methods solve large-scale linear programs (LPs) using GPU-friendly matrix-vector products and projections, but their practical performance depends on coordinating algorithm parameters, acceleration, and restarts. We introduce GALLOP, which uses reinforcement learning to jointly learn continuous algorithm parameters and discrete restart decisions without differentiating through the solver. Its generalized accelerated PDHG update combines separate primal and dual extrapolation, history corrections, and restart anchoring with independently adjustable coefficients. We train a dimension-agnostic feedback policy using a groupwise proximal policy optimization objective that clips likelihood ratios separately for different control groups and excludes inactive acceleration controls on restart transitions. We evaluate GALLOP on six LP families and a public item-placement benchmark. On the main evaluation settings across the six families, GALLOP reduces iteration counts by factors of $1.9$-$5.6$ and achieves up to a $16.0\times$ speedup in algorithm wall-clock time over MPAX. With one policy trained per family, the learned policies generalize without retraining to within-family LPs $3\times$-$400\times$ larger than the largest training instances, including Transport LPs with $10.24$ million variables.
cs.AI / 130 / 2610.01358
Fold'EM: Direct atomic structure inference from Cryo-EM particles
Abstract
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
cs.AI / 131 / 2610.00619
Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games
Abstract
In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.
cs.AI / 132 / 2610.01034
Posterior sampling by source-space MCMC via prior-based few-step transport maps
Abstract
Bayesian inference increasingly uses informative but implicit priors represented only by samples, such as historical ensembles, simulator outputs, and pretrained generative models. The same computational problem appears in the test-time guidance task (generalized Bayes), where an explicit positive weight, e.g., an exponentiated reward, tilts an implicit prior. We develop a framework for source-space generalized Bayesian inference that combines inexpensive few-step prior transports with posterior stability guarantees. Specifically, we represent the prior using a one- or few-step improved MeanFlow (iMF) map and perform posterior sampling in its Gaussian source space. We establish Wasserstein error bounds between the exact and learned posteriors in terms of the joint population iMF and auxiliary-velocity loss, decomposed into training suboptimality and model-class approximation error. In the iMF source space, we adopt parallel tempering with preconditioned Crank-Nicolson updates and introduce a hybrid variant that incorporates split Hamiltonian Monte Carlo to improve sampling efficiency. Synthetic experiments show that the proposed framework can approximate posterior distributions accurately and efficiently, while CLIP-guided ImageNet experiments demonstrate its ability to steer a pretrained iMF image prior toward text-specified preferences.
cs.AI / 133 / 2610.01413
Optimal Transport Meets Reinforcement Learning: A Survey
Abstract
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.
机器学习 (cs.LG)
216
cs.LG / 1 / 2610.01432
Learning ab initio phase-field models
Abstract
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
cs.LG / 2 / 2610.00693
FedMAD: Modulation-Aware Directional Aggregation for Federated Learning in Remote Sensing Image Classification
Abstract
Federated learning (FL) has recently attracted increasing attention in remote sensing (RS) since it enables collaborative model training across decentralized RS image archives without requiring direct access to local data. However, FL performance significantly degrades when the data distributions between clients are heterogeneous, which often occurs due to geographical differences, seasonal changes, and varying image acquisition and atmospheric conditions. To address this challenge, in this letter, we propose a novel personalized FL framework (denoted as FedMAD) for RS image classification problems. The proposed framework separates globally shared representation parameters from client-specific adaptation parameters to preserve client-specific features while maintaining globally transferable representations. This is achieved by integrating lightweight modulation modules and local batch normalization layers into the backbone network. Although globally shared parameters are collaboratively optimized between clients, client-specific parameters remain local to preserve domain-specific feature characteristics. In addition, FedMAD introduces a modulation-aware directional aggregation strategy that dynamically adjusts the importance of aggregation for each client according to the alignment of local modulation updates. This allows the global optimization process to suppress conflicting client updates originating from heterogeneous data distributions while enhancing the contribution of clients with consistent adaptation behaviors. The experimental results obtained on the BigEarthNet-S2 and EuroSAT datasets demonstrate the effectiveness of FedMAD compared to state-of-the-art FL algorithms under heterogeneous RS data distributions. The code of the proposed framework will be publicly available at https://git.tu-berlin.de/rsim/fedmad.
cs.LG / 3 / 2610.00851
SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing
Abstract
Open Set Recognition (OSR) aims to enable models to accurately classify known classes while rejecting samples from unseen classes. A key challenge in OSR lies in the inability to model the unbounded distribution of unknown classes during training, often leading to the misclassification of samples from these classes. Rather than modeling unknowns, recent work shapes the feature space so that known classes are compact and well separated, and spherical representation learning methods have achieved strong results this way. Label smoothing has been identified as one of the key drivers of this success, yet it applies the same coefficient to every training sample, regardless of how well each sample is already embedded. We show that the spherical representation learning objectives used in OSR share a single alignment--uniformity structure in which labels enter only through the alignment term. Label smoothing therefore acts as an alignment dial, and a fixed coefficient sets this dial to the same value for every sample. We propose a plug-in, SmoothOperator (SmoothOP), which sets the smoothing coefficient of each sample from its \textbf{prominence}, an embedding-space signal measuring how clearly the sample's own class stands out against its strongest competing class. Our method integrates into four existing spherical representation learning methods at minimal training overhead. SmoothOP assigns strong smoothing to samples with high prominence, which reduces their alignment and relaxes their pull. On the Semantic Shift Benchmark, SmoothOP-augmented variants generally outperform their base objectives across datasets, degrees of semantic shift, and OSR post-processors, with gains of up to 4.7\% in AUROC, OSCR, and closed-set accuracy.
cs.LG / 4 / 2610.00922
EyeTAG: Eye Trajectory-Aware Gaze Estimation
Abstract
Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0$^\circ$ on Gaze360 and performs on par with the strongest baseline on EVE (2.56$^\circ$ vs. 2.58$^\circ$). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at https://github.com/peter8366/EyeTAG.
cs.LG / 5 / 2610.01134
Open Vocabulary Word Recognition From Transcribed Bangla Texts
Abstract
An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are available in the software industry, finding a reliable equivalent solution for Bangla takes much work. When it comes to handwritten texts, the situation is much more unusual. Recognizing words from word images is the most critical stage in any OCR process. It is the second stage after segmenting words from text pictures. If this stage fails, the overall performance of the OCR will be poor, regardless of how well the other phases perform. This study aims to recognize words using deep learning in a handwritten Bangla word image. Three object detection models, SSD with MobileNetV2, Faster R-CNN with InceptionResNetV2, and an ensemble model of these two, have been used to train and test handwritten word images. A modified Non-Maximum Suppression has been introduced to enhance the effectiveness of the models' results. A customized dataset of 9841 handwritten Bangla word images has been compiled, featuring diverse handwriting styles from various individuals. All three models' performances have been checked against the test dataset, and the ensemble model has been the most impressive, with an F1-score of 92.61%. Also, at the word level, the ensemble model correctly recognizes 96.12% of the words to some extent. The system can be further improved by introducing a post-processing phase to correct errors generated by the system.
cs.LG / 6 / 2610.01408
Smoother Flow Matching via Contrastive Trajectory Repulsion
Abstract
Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the intermediate trajectories. In this paper, we introduce a new network learning perspective: crossing points inherently induce large local Lipschitz constants in the target velocity field, leading to two drawbacks. First, high Lipschitz constants correspond to high-frequency signals in the velocity field that neural networks struggle to fit due to spectral bias. Second, they also imply drastic velocity variations, leading to severe numerical integration errors in few-step inference. To alleviate this, we propose CoFlow, a framework that introduces the contrastive learning paradigm into FM to explicitly repel trajectories during training, thereby lowering the local Lipschitz constants of the velocity field. Specifically, we formulate CoFlow from a Stochastic Differential Equation (SDE) perspective by injecting a repulsive drift term. This drift actively guides the forward process of positive samples away from negative trajectories, effectively reducing the local Lipschitz constant. Furthermore, we derive an equivalent stochastic interpolant formulation from this SDE, providing a simple and tractable design space to control the influence of negative samples. Extensive experiments on ImageNet 256x256 demonstrate that CoFlow significantly reduces FID compared to standard FM in few-step inference (e.g., 20 steps), with no added training overhead. The code can be accessed at: https://github.com/HKUST-LongGroup/CoFlow
cs.LG / 7 / 2610.01807
PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization
Abstract
Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existing domain generalization (DG) methods often focus on style or intensity diversification, but they can still leave networks dependent on domain-specific texture correlations. Inspired by evidence that Fourier phase encodes semantic structure, we introduce PhaseAT, a phase-aware adversarial training framework for medical DG. PhaseAT forms phase-perturbed training views in the Fourier domain by iteratively updating a bounded phase perturbation while keeping the amplitude spectrum unchanged, thereby stressing spatial organization under matched appearance statistics. Perturbations are applied only to the luminance channel in YCbCr color space to avoid chromatic artifacts. Additionally, a simple phase-saliency mask concentrates updates on the most influential frequencies. The model is trained with a weighted combination of losses on clean and phase-perturbed samples, supporting both single-source and multi-source DG. We validate our method on two challenging medical datasets and demonstrate that PhaseAT achieves over 20% improvement in single-source domain generalization, outperforming several state-of-the-art DG methods. The code implementation is available at: https://github.com/ahmed-sharshar/PhaseAT.
cs.LG / 8 / 2610.01994
Comparing a gradient boosting algorithm to the GOES FDC for wildfire detection
Abstract
Wildfires pose severe risks to human life, ecosystems, and property. This study presents a machine learning approach for wildfire detection from GOES ABI imagery. A CatBoost model was trained on a large dataset with thousands of ABI images and over 300,000 matching VIIRS fire detections. An evaluation on a separate dataset across five regions showed that the learned CatBoost model outperformed the operational GOES Fire Detection and Characterization (FDC) product. It achieved higher precision, recall, and F1 scores both within and outside the training area. The CatBoost model achieved F1 scores that were 0.16 to 0.38 higher than the GOES FDC in all regions. In addition, out of 51 historical fire events, the CatBoost detected 26 fires before both VIIRS and GOES FDC, compared to only six earlier detections by the GOES FDC. Importantly, the CatBoost model achieved accurate wildfire detection also during nighttime, whereas the GOES FDC obtained very low recall values, around 0.03. This study demonstrates that machine learning models may offer significant improvements over existing geostationary fire products, including higher accuracy, fewer false alarms, and earlier detection.
cs.LG / 9 / 2610.02000
Weather-Aware Domain Adaptation for Street-View Weather Recognition
Abstract
Adverse conditions such as rain, snow, fog, and dust remain challenging for camera-based perception in autonomous driving. We study multi-class weather recognition from street-view images under domain shift, where most available training data come from non-street-view sources that differ markedly from real driving scenes. We propose Weather-Aware Adversarial Discriminative Domain Adaptation (WA-ADDA), which conditions the domain discriminator on predicted weather to promote features that are both domain-invariant and weather-sensitive. We also assemble a multi-dataset benchmark by unifying diverse non-street-view weather collections as sources and real street-view images as targets, and define a standardized evaluation protocol with macro accuracy as the primary metric. Across backbones (ResNet-50, EfficientNet, VGG, DenseNet), WA-ADDA consistently improves street-view performance and yields strong per-class recalls in challenging conditions while preserving clear-weather accuracy. These findings highlight the feasibility of domain-adapted weather recognition and the value of our benchmark for advancing robust, on-board perception.
cs.LG / 10 / 2610.01738
Participation-Sensitive Convergence and the Fragment First, Converge Later Pattern in Asynchronous Online Learning: A Topological Analysis Across 22 OULAD Courses
Abstract
Asynchronous online learning offers temporal flexibility at a structural cost: learning communities tend to fragment rather than cohere. $β_0$, the number of disconnected behavioral clusters from Zigzag Persistent Homology, serves as a cohort-level indicator of this structure. Two questions remained unverified at scale: (1) does apparent $β_0$ convergence reflect genuine behavioral alignment or learner dropout? and (2) do assessment deadlines produce reproducible fragmentation-convergence cycles? We address both across all 22 OULAD courses (N > 22,000; 857 week-pairs). Changes in $β_0$ strongly co-vary with active learner changes (pooled r = 0.387; median per-course r_delta = 0.459, 20/22 courses), identifying $β_0$ as a participation-sensitive indicator: $β_0$ and active learner counts co-respond to deadline events rather than one causing the other. Deadlines produced fragmentation in 82.6% of assessments and the full Fragment First, Converge Later (FFCL) cycle in 60.2%. 3-phase analysis confirmed structural fragmentation as the dominant long-term trajectory (90.9% of courses), moderated by curriculum structure. These findings establish $β_0$ as a participation-sensitive structural indicator with direct implications for AI-augmented learning analytics design.
cs.LG / 11 / 2610.01749
Designing for Interpretation Uncertainty: Architecture and Principles for Topological Learning Analytics Dashboards
Abstract
Topological Data Analysis (TDA) offers novel methods for understanding temporal dynamics in complex systems, yet its application in information systems design faces a fundamental challenge: how should systems present analytical outputs when interpretation frameworks are still developing? This paper reports on the development of TopoLA, a dashboard system applying Zigzag Persistent Homology to learning management system data, and proposes three early design principles for interpretation support in emerging analytics: (1) separation of objective measurement from contextual interpretation, (2) graduated disclosure from metrics through patterns to reflective prompts, and (3) explicit acknowledgment of methodological uncertainty. The system implements a modular three-stage pipeline--feature extraction, topological computation, and interpretation support--enabling extension to additional analytical methods. This work contributes to information systems research by articulating preliminary design knowledge for systems that must communicate analytical insights from methods lacking established interpretation norms--a challenge increasingly common as novel computational techniques enter applied domains.
cs.LG / 12 / 2610.01647
Towards a Cloud Fog Edge System for Smart Building
Abstract
In this article, we present our vision and recent advancements toward creating a decentralized system capable of learning from real-time data within buildings to support sustainable and privacy-preserving smart environments. Our approach promotes the concept of the building itself as the data center, aligning with the principles of edge computing to safeguard confidentiality and reduce reliance on external cloud infrastructure. This is particularly valuable in humanitarian contexts, where data sovereignty, energy efficiency, and infrastructure constraints are critical. We detail a lightweight, "Kubernetes-like" orchestration framework for deploying AI services within such environments and demonstrate our progress in implementing AI algorithms on low-power, cost-effective microcontrollers such as those in the Arduino ecosystem. By enabling in-situ learning directly on sensors or microcontrollers, our work aims to bring intelligent services to resource-limited settings, fostering autonomy, resilience, and sustainable development in vulnerable or underserved communities. The contributions in this article are related, firstly, to our project "Online Machine Learning Algorithms for Embedded Systems" and the evaluation of two new online algorithms. Secondly, we envision a cloud-fog-edge architecture based on the KOptim and FIWARE components, and we propose a methodology for coupling them. Experimental results of the online algorithms are also presented, showcasing real-world traces.
cs.LG / 13 / 2610.00493
Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts
Abstract
Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert's update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations never see the output factor. If each expert is split into a reader (down-projection) and a writer (up-projection), the usefulness of every reader--writer pair becomes an inner product in the shared rank-$r$ space, and all $N_AN_B$ pairs can be scored from $N_A+N_B$ vectors. We build VANE on this identity. A low-rank compass predicts the descent direction of each token. VANE scores every pair by the alignment between its update and the compass without forming any update, activates the top-$k$ pairs with additive gates, and gives every pair its exact first-order router gradient. On single-domain commonsense reasoning and a four-domain multi-task mixture with Llama-3.2-3B and Llama-3.1-8B, VANE attains the best average among twelve PEFT and MoE-LoRA baselines, by 0.9--1.1 and 1.3--1.5 points respectively, with less than half the trainable parameters of an 8-expert MoE-LoRA. Its router scores also track the measured usefulness of experts far more closely than token routers do.
cs.LG / 14 / 2610.00497
Gumbel Straight Flow: Distilling Autoregressive Models into One-step Flow Maps
Abstract
We present Gumbel Straight Flow (GSF), a continuous flow map language model that leverages the noise-data coupling of a pretrained autoregressive language (AR) model. We theoretically demonstrate that the coupling between Gumbel noise and one-hot token sequences induced by an autoregressive model yields non-intersecting linear paths connecting the noise to the sequence representations. To further enhance high-quality few-step path sampling, we use a flow map semigroup objective where the tangent (velocity) condition is guided directly by the AR teacher. Across various benchmarks, including pretraining and downstream tasks, GSF can outperform current few-step language generation baselines.
cs.LG / 15 / 2610.00518
One-Step Generative Modeling via Training Dynamics Action
Abstract
One-step generative models construct a static generator through iterative training-time transport. Existing transport objectives primarily assess distributional motion, although a neural generator needs to realize the requested sample displacements jointly through shared parameter updates. The training-time construction raises the question: \emph{once training becomes the iterative process that constructs the final one-step map, what to optimize: the next distributional move, or the route by which the finite generator learns the final map?} To address the question, we introduce \textbf{T}raining \textbf{D}ynamics \textbf{A}ction (\textbf{TDAction}), which selects transport targets according to local shared-parameter realization cost while retaining a prescribed level of distributional progress. We formulate the cost as a soft-terminal control problem and derive a closed-form Batch Tangent Action-to-Go value that accounts for parameter effort and terminal mismatch. The criterion captures cross-sample interactions omitted by independent pairwise costs; under isotropic mobility, the criterion agrees with quadratic Euclidean assignment for deterministic balanced couplings. Randomized tangent probes provide a low-rank implementation that constructs shared detached targets without adding an inference-time trajectory. Controlled studies examine the relationship between generator geometry, transport selection, and realized local action. On ImageNet $256\times256$, TDAction attains an FID below $1.1$ without distillation.
cs.LG / 16 / 2610.00523
SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions
Abstract
Conformal prediction guarantees marginal coverage, but a single calibration threshold can still spread that coverage unevenly, over-covering easy regions and under-covering hard ones. SimplexUQ is, to our knowledge, the first benchmark and reproducible protocol for measuring this allocation problem on simplex-valued predictions; it compares existing conformal wrappers rather than proposing a new one. Its task suite, SimplexTasks-12, combines six controlled synthetic regimes with six frozen-predictor real tasks spanning class probabilities, topic mixtures, spectral abundances, cell-type fractions, age distributions, and emotion mixtures. Each comparison fixes the predictor, score, and response-free stratification map, varies only the wrapper, and reports marginal coverage, worst-stratum coverage, max disparity, and within-task radius and compute. Global calibration can look valid while failing badly: on CIFAR-10 it attains 0.900 marginal coverage but only 0.542 in the worst entropy stratum, and Mondrian calibration raises that stratum to 0.886 while reducing max disparity from 0.358 to 0.022. No wrapper dominates, however. Under smooth synthetic heterogeneity, several repairs are competitive; fixed-map analyses show that rankings depend on the evaluation groups and protocol; and in a 12-task comparison, Mondrian has lower disparity on its single target partition for all 12 tasks, whereas BatchMVP has lower disparity over overlapping groups on five. These are empirical comparisons, not new coverage guarantees. A controlled predictor-bias sweep shows that removing predictor bias only partly reduces global-threshold disparity. We release task cards, result provenance, permitted derived arrays, and rebuild instructions, and treat wrapper selection as a diagnostic comparison rather than a universal ranking.
cs.LG / 17 / 2610.00541
Random Recursive Models
Abstract
Recursive models create computational depth through parameter reuse, offering a parameter-efficient alternative to increasing model size. However, most recursive models repeatedly apply one learned transformation or a prescribed sequence of transformations, restricting computation to a fixed layer order. We introduce the Random Recursive Model (RRM), which maintains a pool of $L$ learned layers and performs $T$ recursive steps by sampling one layer independently with replacement for each example and step. This enables flexible layer reuse while retaining the parameter efficiency of recurrence. We evaluate RRM on challenging reasoning tasks, where it matches or exceeds the baselines, often with 50-75 % fewer parameters. RRM can vary its depth at inference, including beyond that seen during training, without retraining or adding parameters, improving tasks that benefit from deeper iterative computation. RRM also supports Monte Carlo inference and probabilistic test-time scaling, both of which improve performance without retraining. These insights may open new directions in neural network architecture design.
cs.LG / 18 / 2610.00545
Geometry-Dependent Bounds for Online Non-Monotone DR-Submodular Maximization
Abstract
We study adversarial online maximization of nonnegative, non-monotone DR-submodular functions over compact convex down-closed sets. A learner commits each action before observing its objective and competes with the best fixed action in hindsight. We prove a comparator-uniform first-order inequality that gives coefficient $4/9$, improving the online $0.401$ benchmark, with one gradient query and one projection per round and $O(\sqrt T)$ expected approximate regret. If $ζ{\bf 1} \in K\subseteq[0,1]^d$, the coefficient improves to $\underlineα(ζ)=\tfrac12-(1-2ζ)_+^2/[2(3-2ζ)^2]$. The proof is a direct ordered-coordinate argument with an objective-independent rational action. Conversely, a three-group symmetry-gap construction yields an offline oracle upper bound $β_*=0.470438681380894\ldots$ at $ζ=0$, even with exact value and full-gradient responses. A parameterized extension and exact finite-instance bounds define an upper function for every $ζ$. The lower and upper bounds match at $1/2$ for $ζ\ge1/2$, and show that the optimal deficit from $1/2$ is $Θ((1/2-ζ)^2)$ as $ζ\uparrow1/2$. For coefficient-revealed polynomials we obtain $1/2$ for quadratics and a geometry-dependent cubic coefficient starting at $8/17$, including $0.49$ at $ζ=1/5$. A constant objective sequence yields an offline $(4/9-\varepsilon)$ approximation with polynomially many first-order queries on the cube and projections, without requiring a supplied positive lower bound on the optimum. We also give nonanticipating adaptive-adversary and value-feedback guarantees, including $O(T^{3/4})$ regret with one noisy value per round.
cs.LG / 19 / 2610.00554
Evaluating Hybrid Quantum-Classical Models for Reduced-Order Brain Deformation Dynamics
Abstract
We evaluate hybrid quantum-classical machine learning for the reduced-order prediction of spatiotemporal brain deformation fields. To mitigate the computational intractability of high-dimensional displacement fields, we employ Proper Orthogonal Decomposition (POD) to project the data into a compact latent space. Within this framework, we formulate two distinct learning objectives: static temporal-to-latent regression and autoregressive latent state forecasting. We systematically benchmark compact classical baselines against both minimal and enhanced hybrid quantum architectures. Our results demonstrate that classical networks provide the strongest baselines in the present setting. For static regression, a classical POD-MLP outperforms all evaluated quantum variants, although an enhanced Variational Quantum Circuit (VQC) substantially improves upon a minimal VQC baseline. For temporal forecasting, a classical POD-LSTM delivers superior predictive accuracy and statistical robustness compared to an enhanced Quantum LSTM (QLSTM) across varying history windows and random initializations. Overall, this study establishes reduced-order physical field learning as a rigorous testbed for near-term QML, highlighting that while hybrid enhancements successfully recover expressivity in weak quantum circuits, classical architectures retain a definitive advantage in both fidelity and stability.
cs.LG / 20 / 2610.00558
Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization
Abstract
While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redundancy (where experts encode overlapping representations) and synergy (where experts provide complementary error cancellation). To navigate this duality, we mathematically decouple redundancy reduction from synergy maximization by formulating the selection objective as a DS function. Furthermore, we devise a tailored majorization-minimization (MM) algorithm with provable monotonicity guarantees to efficiently identify the optimal expert subset. Extensive experiments demonstrate that DS-MoE effectively preserves indispensable expert combinations, achieving superior performance compared to the state-of-the-art baselines.
cs.LG / 21 / 2610.00563
Beyond Affine Transformations: A Soft Dominance Layer for Coordinate-Wise Neural Computation
Abstract
This paper presents a preliminary study of an alternative to the affine transformation underlying conventional neural-network layers. In the proposed Soft Dominance Layer, each output unit compares input coordinates with a learnable reference vector and aggregates smooth inequality responses. A sigmoid relaxation makes the comparisons differentiable, while a sharpness parameter $α$ controls their transition toward hard threshold decisions. The aim is to examine the trainability and direct threshold interpretation of this primitive, not to claim a replacement for affine layers. In single-run MNIST experiments, the highest observed Soft Dominance accuracy is $0.9061$ without annealing and $0.9173$ with annealing, compared with $0.9827$ for the MLP baseline. These descriptive results do not establish reliable configuration rankings or a statistically supported annealing benefit. Learned reference vectors exhibit spatial structure, providing qualitative evidence of structured learning. Repeated-seed experiments and broader datasets are required to assess robustness and practical relevance beyond this proof of concept.
cs.LG / 22 / 2610.00564
Attention Kernels for Learning Maps Between Heavy-Tailed Measures
Abstract
Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.
cs.LG / 23 / 2610.00580
From Task Mixtures to Specialized Experts
Abstract
In collaborative foundation model fine-tuning, client data is rarely homogeneous. Instead, clients typically possess unknown mixtures of distinct data distributions, or tasks. Conventional federated learning primarily addresses heterogeneity across clients without explicitly resolving latent task mixtures within each client. We study this setting as compound heterogeneity, where data is heterogeneous both across and within clients. We study adaptation over a common frozen representation and show that, when tasks share the same feature geometry, the optimal model for a client's task mixture under squared loss is a convex combination of the optimal models for its underlying tasks. Thus, a single locally trained model represents the client's overall task mixture, while individual inputs may be drawn from different underlying task distributions. This motivates routing inputs to specialized experts, and we show that, when the task optima form a simplex, task-aligned routing achieves lower risk than any single adapted model for genuinely mixed clients. With access to a small set of task-labeled public samples, we derive a convex program to recover task experts and match them to their corresponding tasks. Our routing analysis shows that effective specialization requires input-dependent expert selection aligned with each client's task mixture. Motivated by this analysis, we propose FedSEE. Across our experiments, FedSEE avoids the negative transfer observed in the evaluated baselines and improves performance by 2.9 points overall and 3.7 points for the worst-served quartile.
cs.LG / 24 / 2610.00586
Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference
Abstract
Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in $O(1)$ yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across $10{,}000$ random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1's 256 KB SRAM from six to nine.
cs.LG / 25 / 2610.00592
ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning
Abstract
In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least $0.82$ in all sixteen Endless T-Maze configurations and at least $0.99$ on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: https://quartz-admirer.github.io/ALER-Adaptive-Learnable-Experience-Rewriting/.
cs.LG / 26 / 2610.00604
MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Abstract
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference $π_{0.5}$ baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 $\pm$ 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
cs.LG / 27 / 2610.00615
Learning the identity: a case study of how SGD selects among functional decompositions
Abstract
One might think that learning the identity function with a deep linear residual network is trivial - the path along residual connections already implements the identity, and so the network need only drive its weights to zero. However, this zero-weight solution is just one point on an entire manifold of population-loss minimizers, each corresponding to a different decomposition of the identity across the network's layers. Although the population loss does not distinguish among these solutions, stochastic gradient descent (SGD) reproducibly favors particular ones. For instance, under anisotropic label noise, the learned layers exhibit a noise-dependent spectrum; even with weight decay, SGD does not generally recover the zero-weight solution. Changing only the parametrization, while leaving the set of realizable functions unchanged, yields different behavior: factoring each weight matrix as a product of two matrices causes the weights to collapse to zero, even without explicit weight decay. While perhaps mysterious and unintuitive at first, these phenomena can be understood through the lens of entropic loss, which augments the population loss with a term proportional to the expected squared norm of the minibatch gradient (Ziyin et al., 2025). On the identity manifold, the population loss is constant, while the entropic term distinguishes among these decompositions. We characterize its minimizers analytically and use them to derive predictions for the structure of solutions favored by SGD. Networks trained with SGD closely match these predictions. Overall, the identity learning task studied here serves as a clean and simple case study of how the lens of entropic loss can clarify why SGD favors particular decompositions of the same input-output function.
cs.LG / 28 / 2610.00620
Misalignment of Low-Loss Regions Causes Grokking
Abstract
Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based on mode connectivity and the geometry of low-loss regions. The framework predicts that the standard modular-arithmetic setting does not always produce grokking: under a symmetry-preserving train/validation split, we observe a stable anti-grokking case in which validation performance does not recover. This counterexample challenges several existing correlational explanations of grokking. More broadly, our analysis framework and results further suggest that grokking arises when the low-loss regions induced by the training and validation partitions are misaligned. Once these regions become well aligned, training hyperparameters alone cannot produce grokking and the observed dynamics collapse to either trainable or non-trainable behavior.
cs.LG / 29 / 2610.00637
Learning Linear Systems under Heavy-Tailed Noise: A Non-Asymptotic Analysis from A Single Trajectory
Abstract
We establish non-asymptotic sample complexity bounds for the least-squares estimation of vector autoregressive models for exponentially stable systems with heavy-tailed noise based on a single observed trajectory. By assuming i.i.d. noise, bounded noise covariance, and persistent excitation, we show that the estimation error is $\widetilde{\mathcal{O}}(r^{1/2}T^{-1/2+1/p})$ under bounded $p$th moment for $p > 2$, where $T$ is the number of samples, $r$ is the noise dimension, and $\widetilde{\mathcal{O}}(\cdot)$ hides logarithmic terms. We also introduce a unifying approach to sample complexity analysis applicable to broad classes of noise distributions and showcase this by deriving error bounds for sub-exponential and sub-Gaussian noise distributions. Finally, we specialize our analysis to autoregressive models with exogenous inputs and show that the dimension factor of the error bound is independent of the model order.
cs.LG / 30 / 2610.00665
Analysis of Quantized and Efficiently Adapted Protein Language Models
Abstract
Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.
cs.LG / 31 / 2610.00672
ORBIT-FMIB: Tracking Order-Resolved Epistatic Information Through ESM-2
Abstract
Protein foundation models support mutation-effect and structural prediction, but predictive performance alone does not reveal which forms of biological interaction information remain accessible through model depth. We ask whether ESM-2 retains higher-order epistatic information as strongly as first- and second-order information across its representation hierarchy, introducing ORBIT-FMIB, a diagnostic framework combining Walsh-based interaction decomposition with subset-conditioned neural dependence estimation. The method is validated on synthetic landscapes with known interaction structure before being applied to the dense four-site GB1 fitness landscape using frozen ESM-2 representations. An initial production run suggested ESM-2 retains higher-order epistatic information less well than lower-order information ($Δ_{\mathrm{HO-LO}}=-0.107$). An independent replication of the complete measurement grid, under matched GPU hardware and identical critic seeds, substantially reduced this contrast ($Δ_{\mathrm{HO-LO}}=-0.017$), and its sign was unstable across otherwise-defensible evaluation-pairing choices applied to the same trained critics ($-0.011$ to $+0.015$). We therefore do not currently have robust evidence that ESM-2 selectively loses higher-order epistatic information, nor that retention is equal across orders; the directional question remains open. The measurement protocol itself, including its documented removal of a positional-subset shortcut in pooled critics, remains validated and is unaffected by this finding. ORBIT-FMIB is offered as a diagnostic framework for probing interaction structure in protein foundation models; this study's own replication result illustrates why such probing requires adequately-powered reproducibility checks before its output is treated as a biological finding.
cs.LG / 32 / 2610.00675
LabBook: Harnessing Experimental History for Efficient LLM-Driven Discovery
Abstract
Evolutionary approaches to LLM-driven discovery often generate new programs from a small set of selected ancestors. This keeps contexts manageable but can omit useful evidence from other experiments, whereas including the full experimental history produces long, redundant contexts. We introduce a simple, single-agent discovery harness built around LabBook, an agent-maintained memory that serves two complementary roles: guiding retrieval of relevant evidence from a complete experimental log and informing the generation of new solutions. At each iteration, the same agent combines its memory with retrieved evidence and jointly produces the next program and an updated LabBook. This separates complete history retention from selective context construction, without requiring an explicit population or branching search structure. On 49 Frontier-CS problems, LabBook improves the observed quality-cost trade-off over the evaluated evolutionary baselines with two backbones, while remaining competitive across nine additional mathematical, systems, and heuristic-design tasks. Code will be released at https://github.com/BoYuanVisionary/LabBook.
cs.LG / 33 / 2610.00676
Learning Transferable Skills using Goal-Conditioned Bisimulation
Abstract
Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously unseen layouts. A key challenge is to learn representations that capture the temporal structure of the environment while remaining robust to variations across layouts. To address this issue, we present an objective for learning action-aware temporal representations that satisfy the functional equivariance property while preserving the local temporal structure of the environment. Building upon this embedding, we further propose unsupervised skill discovery using bisimulation, which learns transferable skills by conditioning the behavior of skills exclusively on the subset of state features that directly affect their execution. This enforces invariant behavior across different layouts, enabling skills to transfer effectively to other configurations. Finally, through comprehensive empirical evaluations, we show that skills learned in a given environment can be effectively applied to solve downstream tasks in various environment layouts, demonstrating strong out-of-distribution generalization.
cs.LG / 34 / 2610.00680
Curvature Under Attack in hZACH-ViT: Gauge Symmetry, Boundary Saturation, and Adversarial Failure
Abstract
Curvature is often treated as an intrinsic property of a representation, although its empirical effect also depends on coordinate scale, learned logit temperature, and numerical safeguards. We study this interaction in hZACH-ViT, a compact Vision Transformer with Euclidean, Poincare, and spherical prototype heads. The backbone architecture, seed-specific initialization, 50-per-class training subset, and optimization protocol are matched across three MedMNIST datasets and five seeds. At the fixed comparison curvature $c=1$, Poincare has the lowest class-macro PGD attack-success rate in all 12 dataset-budget cells and under a stronger CE+DLR multi-restart attack on all three datasets, but it also has the lowest clean MacroF1. An end-to-end curvature intervention changes the interpretation. Reducing Poincare curvature to $c=0.1$ improves clean MacroF1 in every one of the 15 paired seed-dataset comparisons and removes hard boundary clipping, yet on OrganAMNIST it increases strong attack success from $89.7\%$ to $99.3\%$ (paired difference $+9.57$ points; 95\% hierarchical bootstrap CI $[+5.52,+14.03]$). At $c=1$, $40$-$47\%$ of clean Poincare features are hard-clipped, the radial Jacobian of the inherited map is nearly zero, and dimensionless attack trajectories are unusually long and inefficient. The spherical head provides a control: its curvature change is an exact scale gauge to floating-point precision and produces much smaller attack differences. These results do not establish intrinsic hyperbolic robustness. They identify an implementation-sensitive regime in which curvature, scale, and proximity to the Poincare boundary jointly organize clean recognition and adversarial representation motion.
cs.LG / 35 / 2610.00683
Grand Canonical Generators
Abstract
We introduce Grand Canonical Generators (GCG), a generative framework that extends Boltzmann generators to the grand canonical ensemble. We present two designs. The first conditions a variable-size generative model on the chemical potential, sampling particle number and configuration jointly. The second factorizes the grand canonical distribution into a particle-number distribution and the corresponding canonical Boltzmann density. This factorized formulation can use any existing Boltzmann generator for the canonical component, encodes the known linear chemical-potential dependence analytically, and yields a tractable likelihood that supports self-normalized importance sampling (SNIS). Empirically, GCG accurately reproduces grand canonical observables on a Lennard--Jones fluid and methane adsorption in a zeolite, demonstrating generalization across chemical potentials and correction via SNIS and grand canonical Monte Carlo.
cs.LG / 36 / 2610.00708
Beyond Unimodal Bases: Pullback Geometry for Multimodal Data
Abstract
Data-driven Riemannian geometry provides nonlinear interpolation and geometric representations of high-dimensional data. For these operations to be statistically meaningful, paths between observations should preferentially traverse high-likelihood regions. Existing scalable pullback constructions typically use a unimodal Gaussian latent distribution, assuming that the data reside close to a single manifold. For multimodal data, mapping separated modes or local structures into one Gaussian region can require substantial transport deformation and compromise the resulting geometry. We introduce a pullback geometry for data supported on mixtures of manifolds. Using a latent Gaussian mixture, we define its Riemannian metric as the matrix square of the responsibility-weighted expected component precision. The metric is smooth and positive definite and recovers the existing Gaussian construction in the single-component limit. For structured overlapping mixtures, we establish conditions under which the log-density is concave along geodesics, providing a formal connection between the proposed geometry and paths through high-likelihood regions, and derive the corresponding local curvature relations. We instantiate this geometry in a normalizing flow with adaptive mixture learning, allowing the number of active components to emerge from the data and supporting component-wise reconstruction and local effective-dimension estimation. Experiments on synthetic geometric data, a controlled multi-view image setting with a known reference trajectory, and MNIST show reduced transport distortion, competitive path support, close reference-trajectory recovery, and improved interpolation realism. These results extend scalable pullback geometry beyond datasets that reside close to a single manifold while retaining tractable and interpretable local structure.
cs.LG / 37 / 2610.00713
WOMBAT: Whitebox Oracle for Molecular Benchmarking and Attribution Testing
Abstract
When a graph neural network (GNN) explainer produces an unexpected attribution on a molecule, the attribution alone cannot reveal whether the explainer has failed or the model has learned a shortcut. We introduce WOMBAT, a benchmark of 14 whitebox GNNs, each with message-passing weights set by hand to detect a specific SMARTS motif. Each model's decision rule is known by construction, providing attribution ground truth against which explainer errors can be identified and studied. We validate the models on millions of PubChem molecules and evaluate post-hoc explainers including GNNExplainer, PGExplainer, and Integrated Gradients. Guided by our qualitative analysis, we construct a model that causes Integrated Gradients to spread attribution across the graph, even though the model reliably detects the intended motif. We release the dataset, models, and evaluation code to help researchers in the development of newer XAI tools for GNNs.
cs.LG / 38 / 2610.00722
JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts
Abstract
World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time. Self-supervised updates accumulate across episodes, while the visual encoder and reward head remain fixed, preserving the pretrained representation and task objective. Planning requires neither a goal image nor online environment reward. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.
cs.LG / 39 / 2610.00729
Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation
Abstract
This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and actions, enabling zero-shot adaptation to new environments with completely different observations. We discuss the challenges and feasibility of a reward-based policy and then propose a practical algorithm for training. We demonstrate that a reward policy can be trained within three different environments, Pointmass, Cartpole, and 2D Car Racing, and transferred to completely different observations, such as different color palettes or 3D rendering, or Stretch robot navigation in Habitat-Sim, in a zero-shot manner. We also demonstrate that a reward-based policy can further guide the training of an observation-based policy in the target environment.
cs.LG / 40 / 2610.00730
Reformulation-Contrastive Learning for Mixed Integer Programs
Abstract
Mixed-integer linear programs (MILP) model many real-world decision problems, motivating machine-learning methods that exploit recurring structure to accelerate MILP solving. MILPs can admit many equivalent formulations: integrality-preserving changes of variables and the addition of redundant constraints can alter their formulations while preserving the optimization problem. We leverage these reformulations as a source of self-supervision for learning general-purpose representations of MILP variables and constraints. We characterize the affine reformulations that are valid for every input instance, and distinguish re-descriptions, which leave variables unchanged, from substitutions, which transform them predictably. Building on equivariant self-supervised learning, we introduce ReMILP (reformulation-contrastive MILP representation learning), which jointly trains a graph neural network and a hypernetwork to predict how variable embeddings transform under changes of variables. Without solver-derived labels, ReMILP learns representations that exhibit the intended invariance and equivariance on unseen problem classes. Across binary solution, constraint activity and integrality gap prediction, these representations carry task-relevant information when frozen and provide a useful initialization for fine-tuning.
cs.LG / 41 / 2610.00751
Signal-Noise Factorization Isolates Nuisance Variation into Removable Subspaces
Abstract
Recent theoretical work identified fundamental properties of representation geometry that shape inference ability of deep neural networks. These include signal-noise factorization (SNF), the ability to segregate signal from noise, and signal-signal factorization (SSF), the ability to segregate task-specific and task-irrelevant signals. Here, we built regularizers that reinforce these two properties during training. We compared networks trained with these regularizers to $L_2$-regularized baseline networks on the CIFAR-100 classification task to understand how our regularizers shape representation geometry and impact performance on a well-known computer vision baseline. Enhancing SNF via regularization improved model performance but enhancing SSF did not. Motivated by biomedical applications, we investigated how our regularizers affected performance on the BloodMNIST dataset treated with MedMNIST-C corruptions at five severity levels, and found even larger performance gains using the SNF regularizer. To understand the mechanism by which SNF-regularization produces improved performance, we analyzed the nuisance subspaces across regularization regimes, finding that the SNF-regularized models represent noise in distinct subspaces, separate from class-relevant signal. Because this geometry is explicit, the dominant corruption-induced directions can be estimated on held-out data and projected out of the representations. This manipulation led to a substantial gain in accuracy. These results show that regularizers that enforce signal-noise factorization can produce substantial improvements on computer vision tasks that contain out-of-distribution image distortions at inference time. They also highlight how shaping representations affects model performance: isolating nuisance variables from categorical ones is more important than maintaining factorized representations of categorical variables.
cs.LG / 42 / 2610.00753
Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
Abstract
End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.
cs.LG / 43 / 2610.00758
Scalable Multi-Task Inverse Reinforcement Learning
Abstract
By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect agents' state occupancy. We propose a multi-task IRL method that pools data across multiple agents with different rewards in the same environment under a low-rank assumption. In addition to alleviating coverage requirements, so each task need not visit every state as long as others do, the method offers scalable evaluation of multiple tasks under new environments as computationally intensive planning scales with rank rather than the number of tasks. We provide finite sample guarantees on reward recovery and on policy learning in new environments. Experiments show our method is robust to limited coverage, recovers rewards on and off of each task's support, transfers to target environments at lower regret than baselines, with its computational advantage over per-task methods widening as tasks grow.
cs.LG / 44 / 2610.00767
Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
Abstract
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.
cs.LG / 45 / 2610.00771
Localizing Transfer Between Memorization Tasks
Abstract
A central puzzle in transfer learning is why pre-training on one task can accelerate training or improve performance on another task, and what mechanisms underlie this transfer. In this work, we examine the transfer between memorization tasks of random input-output mappings. We find two surprising transfer patterns: equivalent transfer, where each additional pre-training epoch saves approximately one downstream fine-tuning epoch; and non-equivalent transfer, where pre-training on a mismatched task can be even more efficient than directly training on the downstream task itself. Through ablation experiments, we decompose and localize the transfer into two separate effects: a "trivial" magnitude-driven transfer in the last layer, and a "non-trivial" structure-driven transfer, partially attributable to the covariance of the other layers. These results advance our understanding of the underlying mechanisms of transfer learning and have the potential to lead to principled pre-training strategies.
cs.LG / 46 / 2610.00778
Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability
Abstract
In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction- dependent effects of control composition over a finite horizon together with environmental feasibility. We propose ReQRL, which constrains the critic's value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On OGBench, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.
cs.LG / 47 / 2610.00814
Training-Aware Target Coverage for Synthetic Data Selection
Abstract
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
cs.LG / 48 / 2610.00818
Quantifying the Impact of Ambulance Ramping: A Multi-Year Analysis of Victorian Emergency Medical Services Cases
Abstract
Ambulance ramping, the delay between hospital arrival and patient handover, is a critical operational bottleneck in Emergency Medical Services (EMS), yet its systemic magnitude and dynamics remain inadequately characterised at scale. This paper quantifies the scale, trajectory, and operational correlates of ramping across an entire statewide EMS system, analysing 2,850,575 ambulance attendances in Victoria, Australia from January 2020 to March 2024 using an Exploratory Data Analysis (EDA). After systematic preprocessing, an analytical cohort of 2,026,569 Emergency Department (ED) transports across 59 hospitals with ED and 79 Local Government Areas (LGA) was examined through interval decomposition, Pareto concentration, hourly cross-correlation, hospital arrival concurrency and priority-stratified operational comparisons. Cumulative Ambulance Hours Lost (AHL) totalled 1,491,127 hours, equivalent to approximately 96 ten-hour ambulance shift lost every day of the study window. Ten of 59 hospitals account for 57.8% of lost hours from 50.9% of cases. Annual losses rose 57% to a 2022 peak while transported demand fell 3.7%, indicating deterioration in per-case handover rather than growth in demand. Handover duration varies little with patient acuity, but rises monotonically with the number of ambulances arriving at the same hospital in the preceding hour, an effect persisting within every hour of the day. Hourly demand is moderately associated with ramping two to four hours later (r = 0.365). These findings establish the empirical preconditions for hospital-state aware ambulance routing.
cs.LG / 49 / 2610.00820
On-the-fly Weight Generation: A Hypernetwork Proof of Concept on ARC-1D
Abstract
General-purpose models can adapt to many tasks from context, while specialised models can execute individual functions with less capacity. Yet obtaining such specialists requires task-specific training or adaptation. We ask whether they can instead be generated directly from a few demonstrations. Using ARC-1D as a controlled testbed, we show that individual transformations can be represented by tiny specialist models, and that a hypernetwork can generate their parameters from context. The generated parameters form a structured weight space, while the resulting specialists show partial compositional generalisation and generalisation to transformations not seen during training. In both settings, removing explicit task identifiers improves generalisation beyond the training transformations. Together, these results provide a proof of concept that few-shot task context can be compiled on-the-fly into compact executable model parameters, and that the resulting weight space can support reuse and generalisation beyond known functions.
cs.LG / 50 / 2610.00831
AnyJev Technical Report
Abstract
A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.
cs.LG / 51 / 2610.00858
Fixing a Model That Learned Worse Cancer Means Lower Risk: Monotonic Constraints in Bladder Cancer Recurrence Prediction
Abstract
Background and Objective: Clinicians expect recurrence risk to climb with cancer severity. In a UK multicentre trial, an unconstrained XGBoost model learnt that higher tumour stage and carcinoma in situ predicted lower recurrence risk, and discrimination, calibration, and SHAP were all blind to it. We developed a counterfactual testing framework to detect this inversion and a monotonic-constraint framework to remove it without hurting performance. Methods: BOXIT enrolled 472 patients with protocol-mandated cystoscopy across 51 UK sites (2007-2012); 435 had at least two years' follow-up (153 recurrences, 35.2%). We developed a counterfactual direction test and a monotonic-constraint correction, with constraint directions drawn from the EORTC and EAU risk systems, and evaluated both against unconstrained XGBoost and logistic regression on 18 predictors (seven directed) over 50 cross-validation folds. The test worsened each patient on one directed feature at a time to check whether risk fell; SHAP direction and calibration were also assessed. Key Findings and Limitations: Tumour stage and carcinoma in situ were associated with lower recurrence, opposite to medical intuition; the unconstrained model reversed carcinoma in situ counterfactuals in 90.2% of cases and stage in 74.3%. Discrimination ($Δ$AUC 0.005, p=0.47), calibration, and SHAP magnitude were all blind to the inversion. Monotonic constraints eliminated every violation at no cost to discrimination (0.723 vs 0.718) and outperformed EORTC (p=8.9e-16). Limitations: single trial, internal-external validation only. Conclusions and Clinical Implications: A model that had learned this inversion passed every conventional check. A counterfactual direction test, run as a single refit with pre-specified monotonic constraints, catches this failure at no cost to performance and should be routine before clinical deployment.
cs.LG / 52 / 2610.00861
Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
Abstract
Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, which restrict the extent to which perturbation can be concentrated on a small number of locations. We introduce an importance-guided allocation mechanism that uses a fixed clean-gradient prior to steer perturbation toward model-sensitive regions while leaving the feasible perturbation set unchanged. A centered allocation objective encourages perturbation at above-average importance locations and discourages unnecessary expenditure elsewhere, thereby redistributing rather than enlarging the available budget. Across ten robust model--dataset configurations under a common capacity-limited threat setting, the proposed method improves attack success over matched APGD- and PMA-based baselines by $2.52$ to $17.70$ percentage points. Allocation analysis shows that these gains are accompanied by substantially greater perturbation mass in high-importance regions without increased global $\ell_1$ consumption. Mechanism ablations further show that centered non-uniform redistribution provides part of the benefit, while model-derived importance yields an additional improvement. These results identify perturbation allocation as a distinct and practically relevant dimension of adversarial optimization under shared-budget, locally constrained threat models.
cs.LG / 53 / 2610.00888
Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
Abstract
Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with $K{=}3$ chained MTP heads, we show that a post-training recipe with plain cross-entropy on $\approx\!2.5$B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes $10^3$-$10^4\times$ less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by $+12$ to $+16\%$ per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto $11$--$14\%$ loss in speedup using fixed maximum MTP draft length.
cs.LG / 54 / 2610.00890
Cross-Benchmark Transfer from RL on Agentic Coding Tasks
Abstract
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p < 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model's failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
cs.LG / 55 / 2610.00898
When Do Biological Reasoning Models Use Their Biological Inputs?
Abstract
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
cs.LG / 56 / 2610.00903
Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning
Abstract
Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman $ρ=-0.90$; $ρ=-0.99$ after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at $3.39$ nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches $0.07$ nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.
cs.LG / 57 / 2610.00921
In CEM, a World Model Is Also a Proposal Mechanism
Abstract
The cross-entropy method (CEM) uses world-model scores to select action sequences and fit the distribution sampled in its next iteration. A scoring error can therefore change both the present decision and the candidates considered later. We evaluate these two roles separately. Four types of predictive model generate CEM traces, and every model rescores every saved candidate pool. Executing the same candidates in the environment provides a reference elite set and proposal update. Across twelve independently trained task-seed units on Walker and Cheetah, the pre-specified proposal distance falls from the first to the final CEM iteration in every unit. Proposal widths contract and fitted means separate relative to the remaining search width. Pairwise ranking agreement stays near chance on Walker and declines on Cheetah; elite-set agreement does not improve. This comparison shows greater variation between scorers than between pool sources on Cheetah; Walker has variation in both and in their pairings. We use the original six units to select Random nonlinear for a one-update intervention, without inspecting intervention outcomes. Replacing its first model-ranked update with an environment-ranked update lowers final realised selected-sequence cost in those six units and in six further units held out from the selection.
cs.LG / 58 / 2610.00927
Rate-Optimal Algorithm for Adversarial Linear CMDPs
Abstract
We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves $\widetilde{\mathcal{O}}(K^{3/4})$ regret and cumulative constraint violation, leaving a gap to the optimal $\widetilde{\mathcal{O}}(\sqrt{K})$ dependence on the number of episodes $K$. We close this gap by proposing a new primal dual algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret and cumulative constraint violation without assuming Slater's condition. The main challenge is that learning linear CMDPs requires uniform concentration over a value function class with a controlled covering number, whereas standard techniques in constrained online learning, such as policy mixing, can make this class more complex. Our algorithm combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function. An adaptive dual regularizer offsets the dependence on the dual weights in the primal regret bound, removing the need for policy mixing. We further show that the normalization in the FTRL update bounds the policy parameters independently of the magnitudes of the dual weights, which explains why the resulting policy class remains compatible with uniform concentration. Under feature access, the computational complexity is independent of the size of the state space.
cs.LG / 59 / 2610.00929
Platonic Task Arithmetic
Abstract
Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so weight-space task arithmetic stays confined to a single model and cannot cross architectures without a structural correspondence. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows of one shared, model-agnostic object, which we call the platonic task vector. To make it operational for models that pair an image or audio encoder with a text encoder, we introduce Universal Task Descriptors: matrices whose shape is independent of architecture and embedding dimension, which record a task's functional effect and support addition and negation as matrix operations. Transferring a descriptor into a target means editing the target until it reproduces the descriptor on the task's unlabeled probe images and class-name prompts, requiring no per-image labels. We realize this edit in two ways. First, the descriptor factorizes into a shift field on image embeddings, so a single least-squares solve yields a linear operator that folds into the target's last layer as a weight edit; by linearity, a bank of such operators admits any composition at any strength as a signed sum. Second, a low-rank adapter trained on the same objective reaches every layer and fits compositions jointly, at the cost of one optimization per edit. Heterogeneous models share this object only partially, with a model-specific residual comparable in norm to the shared component, yet cross-model transfer still retains 74-80 percent of the gain of the target's own descriptors. Experiments across six model families, eight classification tasks, and an audio-text setting show that task knowledge transfers and composes across heterogeneous models under both realizations.
cs.LG / 60 / 2610.00948
GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
Abstract
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.
cs.LG / 61 / 2610.00968
Structure-agnostic Causal Representation Learning
Abstract
Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods require specifying the causal structure a priori, yet different structures demand fundamentally incompatible invariance constraints, and misspecification leads to representations that discard predictive information. We introduce SaCRL, a framework that jointly identifies the causal structure and learns the corresponding invariant representation without prior structural knowledge. Our approach formulates structure selection as a soft optimization over candidate invariances using HSIC-based violation metrics, with adaptive weights that automatically concentrate on the achievable structure. We provide theoretical guarantees for structure identification, including under random-feature approximation, invariance satisfaction, and out-of-distribution generalization. Empirically, SaCRL recovers the true structure on synthetic and semi-synthetic Bayesian-network benchmarks, outperforms fixed-invariance baselines on Colored MNIST, achieves state-of-the-art accuracy on three DomainBed benchmarks (PACS, VLCS, OfficeHome), and degrades gracefully under structural misspecification and limited environment diversity. Code is available at: https://github.com/ArmanBehnam/sacrl.
cs.LG / 62 / 2610.00976
Variational Streaming Flow: Probabilistic Forecasting in Physical Time
Abstract
Probabilistic forecasting is important for predicting complex dynamical systems because intrinsic randomness and incomplete observations can cause the same observed state to evolve into multiple plausible futures. While flow matching is a flexible approach for probabilistic forecasting, it is computationally expensive. Streaming flow (SF) reformulates this approach to model temporal evolution efficiently by learning a continuous velocity field directly in physical time. However, SF learns a deterministic velocity field. Thus, it provides only a single future trajectory for a given fixed initial state and observation history. To overcome this limitation, we introduce Variational Streaming Flow (VSF). Our approach learns a latent distribution that is conditioned on the dynamics of interest. In turn, this enables probabilistic forecasting. Importantly, we retain the computational efficiency of SF by generating in physical time. Across deterministic and stochastic dynamical systems, VSF demonstrates superior predictive accuracy and distributional fidelity. We demonstrate the advantage for both long-horizon rollouts exceeding 1,000 steps, and settings with bifurcating dynamics. Moreover, VSF can be integrated into existing Joint-Embedding Predictive Architecture (JEPA)-based world models as a plug-and-play predictor to improve temporal dynamics and goal-directed success rate in navigation, motion planning, and manipulation.
cs.LG / 63 / 2610.00978
Generalist Representation, Specialist Detection: TS-Router for Time-Series Anomaly Detection
Abstract
Time-series anomaly detection (TSAD) is difficult to generalize across datasets because heterogeneous temporal dynamics imply different notions of normality and favor different detection criteria. While time-series foundation models provide transferable representations, coupling them with a fixed anomaly-scoring mechanism can overlook this variation. This motivates a different perspective on foundation-model-based TSAD: using foundation models to coordinate specialized anomaly criteria rather than directly imposing a universal one. Based on this view, we propose \textbf{TS-Router}, a generalist-representation, specialist-detection framework that estimates the relative competence of heterogeneous anomaly detectors from pretrained temporal representations and selects suitable specialists for each target series. To avoid relying on specialist-performance labels from real tasks, we derive soft competence supervision from specialists' relative performance on labeled simulated tasks. At deployment, routing requires no target anomaly labels, and only the selected specialists are fitted unsupervisedly on the target series. We bound Top-\(k\) set-competence regret under representation coverage and conditional competence stability. Across 16 real-world benchmarks and four complementary evaluation metrics, TS-Router achieves the best overall average rank. Controlled ablations with multiple frozen TSFM encoders further support the use of pretrained representations for competence estimation and adaptive specialist selection. The code is available at https://anonymous.4open.science/r/TS-Router-D8FF.
cs.LG / 64 / 2610.00984
HADRec: A Hierarchy-Aware Drug Recommendation Framework by Fusing Molecular Knowledge and Electronic Health Record
Abstract
Accurate medication recommendation is central to clinical decision-making, directly determining therapeutic efficacy and patient safety. However, existing methods suffer from two key limitations: drugs are often abstracted as discrete tokens, ignoring their molecular structures and pharmacological mechanisms, and the commonly used "flat" recommendation paradigm fails to leverage the hierarchical logic of the internationally standardized Anatomical Therapeutic Chemical (ATC) classification system. To address these issues, we propose HADRec, a Hierarchy-Aware Drug Recommendation framework that integrates molecular knowledge with electronic health records (EHRs). HADRec employs LLaMA-7B to encode clinical notes for rich patient representations and ChemBERTa to encode drug Simplified Molecular Input Line Entry System strings, building a global molecular knowledge base. A cross-attention mechanism then performs deep multimodal fusion between patient states and drug features. The framework further incorporates a hierarchical predictor and a novel consistency constraint loss to enforce strict adherence to ATC logical dependencies. Extensive experiments on MIMIC-III demonstrate that HADRec achieves state-of-the-art performance across Jaccard, F1, and PR-AUC. External validation on MIMIC-IV confirms strong generalization under distribution shifts, and calibration analysis shows well-calibrated predictive confidence on MIMIC-IV with ECE = 0.04, and Brier = 0.06. Counterfactual evaluation reveals clinically aligned reasoning, disentangling disease-specific treatments from general care. Together, these results establish HADRec as a high-performance, interpretable, and clinically grounded pathway toward safe and reliable AI-driven medication recommendation.
cs.LG / 65 / 2610.00985
Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks
Abstract
Kolmogorov-Arnold Networks (KANs) represent a compelling alternative to traditional Multi-Layer Perceptron (MLP)-based neural networks. By employing activation functions as learnable elements, KANs offer superior interpretability, making them suited for scientific domains. In this work, we investigate the neural scaling laws of KANs and the structural evolution of their learnable activation functions under dataset expansion. Specifically, we evaluate the scaling behavior of three KAN variants---BSRBF-KAN, Gottlieb-KAN, and Faster-KAN---across standard image classification benchmarks (MNIST and Fashion-MNIST) and a specialized scientific regression task (magnetic parameter estimation from domain images of moiré magnetic textures). Our results demonstrate that the test loss ${\cal L}$ exhibits a broken neural scaling law (BNSL) behavior as a function of the dataset size $N_D$. After passing through a random-guess regime, the loss follows architecture- and task-dependent scaling behavior. The loss crosses from a faster- to a slower-scaling branch, ${\cal L}\propto N_D^{-α}$ and ${\cal L}\propto N_D^{-β}$ with $α>β$ for image classification tasks. The exponents $α$ and $β$ depend strongly on both the specific network architecture and the dataset-size regime, ranging from 0.4 to 1.5 and from 0.06 to 0.6, respectively. For the magnetic parameter-regression task, the loss follows a single scaling law with its exponent ranging from 1.28 to 2.59. Additionally, we provide a structural analysis of how activation functions refine their complexity as data volume increases, finding that dataset expansion drives a transition from simple linear-like approximations toward stable, interpretable symbolic forms. These findings provide a quantitative roadmap for the efficient application of KANs while managing the trade-off between model expressivity and computational overhead.
cs.LG / 66 / 2610.00988
Auditable Algebraic Counting Field for Cryptic-Pocket Detection from Apo Structures
Abstract
Cryptic ligand-binding pockets are not apparent in experimentally determined apo structures, making them difficult to identify from unbound receptor geometry. A complementary challenge is to make the structural measurements and learned evidence behind each prediction directly inspectable. We introduce a supervised algebraic counting field (ACF) for predicting cryptic-pocket residues from apo structures. ACF compiles explicit geometric, physicochemical, and topological features into compact, integer-weighted lookup tables. Each prediction score can be reconstructed from feature values, training counts, table weights, and spatial aggregation, without sequence search, structural-template transfer, or a protein language model at inference. We evaluate ACF on CryptoBench and two locked external collections, separating ranking performance from the effects of residue-calling budgets. On an external set of 57 post-CryptoBench apo-holo units, ACF exceeded P2Rank by +0.044 in mean paired ROC-AUC (multiplicity-adjusted 95% CI [+0.010, +0.079]). The advantage was dataset-dependent: official-fold ROC-AUC and matched-budget F1 differences against P2Rank remained unresolved, and a second external evaluation did not confirm gains from added structural features. ACF thus provides a compact predictor with externally validated signal and an inspectable path from structural measurements and training counts to residue scores.
cs.LG / 67 / 2610.00991
Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
Abstract
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
cs.LG / 68 / 2610.00996
Reliability-aware short-term roll prediction for unmanned surface vehicles via multi-task learning and adaptive centralization
Abstract
Reliable roll prediction of unmanned surface vehicles (USVs) is essential for ensuring navi?gational safety and enhancing autonomous decision-making. While existing studies primarily focus on improving prediction accuracy, the quantification of prediction reliability remains insufficiently addressed. To bridge this gap, this paper proposes a reliability-aware prediction paradigm that integrates confidence assessment into the predictive pipeline. The architecture utilizes a multi-task learning structure where a shared feature extraction backbone feeds into dual heads: a regression head for precise roll prediction and a quantification head for confidence scoring. This configuration provides accurate prediction and corresponding confidence for risk?sensitive downstream tasks. In addition, an adaptive centralization strategy tailored for short?term real-time roll prediction is introduced to improve model generalization under varying operational conditions. Experiments conducted on a real-sea dataset demonstrate that the proposed method effectively quantifies the reliability of prediction results and maintains superior generalization under varying conditions, offering significant potential for practical engineering applications.
cs.LG / 69 / 2610.01028
Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise
Abstract
Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.
cs.LG / 70 / 2610.01062
Kernelized Activation Steering
Abstract
Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local geometry of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space. KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, while richer kernels enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS outperforms or is on par with the existing methods.
cs.LG / 71 / 2610.01076
GLoC-EHR: Evidence-Cited Clinical Reasoning over Global Context and Local EHR Events
Abstract
Structured electronic health records (EHRs) contain a patient's clinical trajectory as a sequence of clinical codes. Answering clinical questions from such records requires both the context of the whole trajectory and the specific events that support the answer. We introduce GLoC-EHR, a multimodal language model that reads a contextual encoding of the record through a fixed-size global memory of the trajectory and a local memory of selected events. The model learns to generate hospital-course summaries from the global memory and descriptions of masked concepts from the local memory, aligning both with clinical text. It is then trained to cite evidence before answering, through rationale fine-tuning followed by group relative policy optimization (GRPO) with rewards for correct answers and record-supported evidence. On three MIMIC-IV outcome tasks, GLoC-EHR attains the highest macro AUROC among the compared models when it answers directly, whereas zero-shot LLMs reading the serialized record fall far behind. With evidence-cited reasoning, it stays close to its direct multi-task counterpart in macro AUROC, and the evidence terms of the objective reduce unsupported evidence at a similar macro AUROC. The local memory adds distinct supported findings, particularly under strict matching, without a detectable change in macro AUROC. Without retraining, GLoC-EHR transfers to EHRSHOT on par with EHR-BERT and answers two unseen laboratory questions better than zero-shot prompting of its own backbone.
cs.LG / 72 / 2610.01096
Dataset Identity, Not Novelty: The Source of an Inflated OOD Detection Gain
Abstract
A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and returns a score. It fits that score on in-distribution data, and the benchmarks that evaluate it supply a second piece of OOD data for the fitting itself. Some detectors tune a constant on it. Others fit a direction in feature space or train a flexible combiner and report the number that it reaches as the gain that is still available. Every such fit is validated on held-out samples of the same OOD dataset. That check rules out memorizing individual images. It says nothing about a fit that has instead learned which dataset it is looking at, and a direction that recognizes one OOD dataset rather than novelty passes it perfectly. The detector that a practitioner installs meets OOD data from a source that nobody fitted it on, so the difference decides what the reported number is worth. We measure it by holding out the whole OOD dataset rather than a sample of it, and we call that gap the inflation. We read it across a range of combiners on ImageNet and CIFAR-100 backbones. Most of the gain that the usual protocol reports turns out to be dataset identity rather than novelty. The size of the fit does not move what survives, so the effect is not ordinary overfitting. The share depends instead on whether the input exposes class identity, and two controls that vary that property alone separate the inflation on every backbone of both benchmarks. A closed form accounts for the effect and computes it from the fitting rows, so a practitioner can tell which fits will inflate without running the hold-out protocol. One of these fits survives, namely the single constant that the field already picks on a designated validation dataset. Anything above it reports a gain that the hold-out protocol does not return, and on one benchmark what survives falls while what is reported climbs.
cs.LG / 73 / 2610.01110
How Much Can Language Models Gain from Test-Time Computation?
Abstract
How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.
cs.LG / 74 / 2610.01126
Latent Information Sharing for Accelerating Federated Learning
Abstract
Federated learning (FL) is a communication-efficient distributed learning paradigm. However, client drift remains one of the most critical challenges, hindering the efficient training of a global model. In this study, we propose a novel latent information sharing scheme that directly mitigates data heterogeneity across clients. Our theoretical and empirical results show that sharing a small amount of hidden-layer activations significantly improves training efficiency while preserving convergence guarantees and data privacy. Furthermore, we compare our method with existing FL approaches designed to address client drift, including FedProx, SCAFFOLD, FedPVR, FedProto, and SplitFed, and demonstrate superior model accuracy under a fixed round budget without incurring excessive communication overhead. Overall, this work presents a promising new knowledge aggregation scheme and provides a comprehensive analysis of the impact of activation sharing on federated optimization.
cs.LG / 75 / 2610.01133
Does Scaling Reinforcement Learning Really Require More Training?
Abstract
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.
cs.LG / 76 / 2610.01143
Parameter-Efficient Distributionally Robust Adaptation of Tabular Foundation Models under Subpopulation Shift
Abstract
Despite strong mean accuracy, tabular foundation models (TFMs) can perform poorly on underrepresented groups under subpopulation shift, where group proportions change between training and deployment. We propose DR-TFM, a parameter-efficient distributionally robust adaptation framework that requires no true group annotations. DR-TFM adjusts attention to labeled context examples by fine-tuning an existing query scaling network or adding and training one, while keeping all other parameters fixed. We instantiate the framework with two robust objectives using estimated groups or source conditional distributions derived from training data. For TabPFN-3, adaptation updates only 0.016% of the pretrained model's parameters. Across five tabular benchmarks, DR-TFM achieves substantially higher average worst-group accuracy than pretrained TFMs and the compared robust baselines without true group annotations, while maintaining competitive mean group accuracy. DR-TFM also improves average worst-group accuracy on ACS Income and across four additional TFMs.
cs.LG / 77 / 2610.01153
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
Abstract
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.
cs.LG / 78 / 2610.01165
Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining
Abstract
Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-forced contexts across five released trajectories and 11 model-domain combinations. We find that block-bypass responses retain recognizable depth ordering while their magnitudes redistribute: nearby checkpoints preserve stronger rank correspondence than distant ones, and large changes concentrate at positions that recur across text samples and transfer across evaluation domains. Controlled experiments further show that changes in the natural bypass effect cannot be reduced to a single downstream sensitivity: in replicated Pythia runs, local missing-update magnitude grows while the pooled matched downstream response decreases, whereas OLMo-2 7B exhibits a different balance. These matched responses also depend on perturbation strength and direction, without identifying targeted compensation. Together, our results show that longitudinal layer sensitivity is structured but not static, and that single-checkpoint intervention responses should be interpreted in the context of how the underlying perturbation pathway evolves during training.
cs.LG / 79 / 2610.01168
Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability
Abstract
Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.
cs.LG / 80 / 2610.01172
Learning Rate Transfer for Hybrid Transformer-SSM Architectures
Abstract
We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 up to billion-parameter scale using only the original $μ$P prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of $μ$P correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by $μ$P's initialization and LR scaling; and a local per-component balance, provided by AdamW's per-parameter normalization. Our observations show that the optimal LR is invariant to width up to 8$\times$, that this width invariance holds across depth, sequence length, batch size, and Transformer-to-SSM ratio, and that it transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.
cs.LG / 81 / 2610.01173
CAGE-NAS: Certified Functional Descent for Efficient Model Growth
Abstract
The progressive growth of neural networks requires deciding when the current representation remains sufficient for optimization and when it should be expanded. CAGE-NAS formulates this decision in function space through an admissibility criterion on approximations of the functional gradient. As long as a representation enables a certified Functional Gradient Descent step, the architecture remains fixed; when the criterion fails, a function-preserving expansion is applied and the resulting representation is evaluated again. As the main instance, we study the family induced by the tangent space, using a regularized projection of the functional gradient. In a controlled setting with exact certification, CAGE-NAS produces architectures positioned above the 99.8th performance percentile by held-out RMSE among all admissible alternatives within the same parameter budget, without enumerating them during the growth trajectory.
cs.LG / 82 / 2610.01175
Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions
Abstract
Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z;X), implicitlytreating all information as homogeneous. However, a single global compression control couples label-relevant structurewith residual within-condition variation, rather than regulating their allocation independently, allowing nuisanceinformation to persist in learned representations. For example, in medical imaging applications, residual variation oftenstems from acquisition conditions, background factors, or subject-specific appearance. This issue becomes particularlypronounced in data-limited settings, where models tend to overfit such variation, hindering generalization. While existingregularization methods can stabilize training, control capacity, or shape representation geometry, they do not explicitlyseparate nuisance-like variation from task-supporting structure. To address this limitation, we revisit IB from a structuredperspective based on a label-induced partition, where condition-level structure and within-condition information playdistinct roles. This leads to a dual-bottleneck formulation: a standard KL term controls global information capacity, while aconditional KL term targets within-condition information. We show that the conditional KL admits an exact decompositioninto a within-condition information term and a prior-mismatch term, explaining its alignment with the design objective.With a simplex-structured conditional prior, the method provides controllable latent geometry and integrates seamlesslyinto existing pipelines. Experiments on classification and segmentation show the clearest gains in low-data classificationand consistent improvements across dense prediction benchmarks.
cs.LG / 83 / 2610.01181
Fully Online Decentralized Learning in Stochastic Games with Unknown Independent Chains
Abstract
We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs. We develop a fully online, decentralized, and uncoordinated mirror-descent algorithm that operates in the dual space of occupancy measures for approximating stationary Nash equilibrium (NE) policies. The algorithm uses a single transition/reward sample at every primitive time step, relies only on local information, and requires neither coverage of the joint state space nor synchronized episodes. Under uniform-ergodicity and finite-coverage assumptions, we show that, with high probability, the time-averaged fixed-comparator regret decays at the canonical $O(T^{-1/2})$ rate, up to logarithmic factors and polynomial dependence on the game parameters. In particular, the complexity depends on the cover times of the individual local state spaces rather than the product state space, avoiding exponential dependence on the number of players and the sizes of the joint state and action spaces. The resulting finite-time regret bound further yields an approximate coarse-correlated-equilibrium guarantee, which is natural for arbitrary reward functions since computing a stationary $ε$-NE is PPAD-hard in this setting. Under an additional global variational-stability condition, we show that the same fully online algorithm converges asymptotically in the last iterate to a stationary $ε$-NE. Our results provide a fully online and scalable learning framework for stochastic games with unknown independent chains. The algorithm can also be viewed as a primal-dual framework for Markov games that exploits the independence and local structure of the players' controlled transition chains.
cs.LG / 84 / 2610.01193
Counterfactual Generation via Flow Matching: Coupling-Sensitive End-to-End Rates
Abstract
Counterfactual generation seeks to sample outcomes under a hypothetical intervention or decision using observational data collected under the factual assignment mechanism. We develop a flow-matching approach that combines a sample-split, doubly robust training objective with a learned coupling between observed source outcomes and target outcomes drawn from a fitted conditional outcome model. To enable finite-step generation, we leverage a score-corrected stochastic sampler based on a Gaussian-smoothed interpolation. Our main theoretical contribution is a coupling-sensitive KL bound for constant-step Euler discretization: the error is controlled by moments of the source--target displacement under the chosen coupling, rather than by global uniform regularity of the velocity field, and has near-linear dependence on the ambient dimension. We also establish finite-sample non-parametric guarantees for the learned velocity and score fields when both the conditional outcome model and the source-target coupling are estimated from data. These bounds separate approximation, coupling-replacement, nuisance-estimation, generalization, and Monte Carlo errors and, combined with the sampler analysis, yield an end-to-end guarantee for counterfactual generation. Experiments on synthetic and semi-synthetic image benchmarks support the coupling-dependent theory and show that, at finite discretization budgets, the stochastic sampler can outperform the corresponding deterministic ODE sampler.
cs.LG / 85 / 2610.01199
Low-Budget Active Learning through Entropic Optimal Transport
Abstract
We consider low-budget active learning, which consists of selecting a limited number of points, the coreset, such that a model can be trained to high accuracy on the selection only. This problem is particularly relevant in contexts where labeling requires costly expert intervention, as in medical applications. We leverage features extracted from a pretrained self-supervised model to represent the data, and perform coreset selection directly in this feature space. In this paper, we use entropic optimal transport, specifically the Sinkhorn divergence, as the coreset selection criterion, which first allows us to get dimension-free sample complexity results, and second admits computationally efficient gradient evaluations. This opens the way to using gradient-based algorithms to rapidly compute solution candidates, further improved by a swap-based local search, with guarantees on the solution quality. Experiments on image benchmarks and medical datasets show that our method outperforms state-of-the-art heuristics in low-budget settings.
cs.LG / 86 / 2610.01204
Autoregressive Drillhole Modelling Under Distribution Shift
Abstract
Autoregressive modelling has achieved remarkable success in language and sequence tasks by learning to predict future states from previous observation. Mineral-exploration drillholes provide a natural but largely unexplored setting for this paradigm: as drilling proceeds, lithology is revealed sequentially from shallow to deep, making prediction of deeper strata inherently autoregressive. Existing drillhole modelling, however, is dominated by spatial interpolation and reconstruction, or largely rely on masked modelling, leaving strictly autoregressive prediction largely underexplored. We introduce DrillBench, a benchmark of 49,671 Western Australian drillholes for next-layer prediction and autoregressive stratigraphic generation across a graded transfer spectrum, from local prediction through spatial shift to cross geological province transfer. Benchmarking classical, geostatistical, and neural models reveals a clear \emph{transfer boundary}: spatial and geochemical conditioning provides large local gains but deteriorates sharply under stronger shift, whereas lithology-sequence autoregressive models transfer more robustly. Guided by this finding, we develop a backbone-agnostic recipe combining large-scale pretraining on historical drillholes with spatial retrieval of neighbouring lithology. Retrieval is most effective in weathered cover, when local spatial continuity remains informative, whereas pretraining contributes more strongly in bedrock and under broader geological shift. Together, they retain strong local performance while improving generalisation under spatial and cross-province shift, most markedly on the most distant splits. The benchmark and code are available at https://github.com/yihaoding/drillbench.
cs.LG / 87 / 2610.01224
Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning
Abstract
Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.
cs.LG / 88 / 2610.01238
Mixture-Trained Merging for Unified Multi-Objective Models
Abstract
Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on single objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.
cs.LG / 89 / 2610.01253
Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
Abstract
Quantum Reinforcement Learning (QRL) integrates reinforcement learning with parameterized quantum circuits and is a promising approach to combinatorial optimization. On Noisy Intermediate-Scale Quantum (NISQ) devices, however, decoherence, gate imperfections, and measurement errors reduce policy quality and make learning less reliable. Existing error mitigation techniques are generally applied as fixed corrections that do not adapt to changing noise conditions or to the evolving state of training. This work presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer of the hybrid quantum-classical training loop that dynamically selects the most suitable mitigation strategy during QRL training. APGEM evaluates Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) using policy-level indicators, including quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio, and integrates the selected strategy directly into the reinforcement learning loop. The framework is evaluated on the Capacitated Vehicle Routing Problem (CVRP), a representative NP-hard problem in urban logistics, under a range of NISQ noise models and noise levels. APGEM consistently outperforms conventional static mitigation methods, reaches approximately 94% of the utility of an oracle strategy, maintains higher quantum-state fidelity as noise increases, and produces more stable learning behaviour throughout training. Ablation studies show that the framework learns context-aware mitigation policies that adapt to different noise environments and circuit execution conditions. These findings demonstrate that integrating adaptive error mitigation into the learning process substantially improves the robustness and reliability of QRL on NISQ hardware.
cs.LG / 90 / 2610.01269
IQS-BO: In-Context Query Selection for Bayesian Optimisation
Abstract
Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogate and maximising an acquisition function at every evaluation step. In-context approaches based on Prior-data Fitted Networks (PFNs) amortise part of this cost by pre-training transformers on functions drawn from synthetic priors. PFNs4BO amortises the surrogate but still relies on a numerically maximised acquisition function, while FIBO performs BO fully in-context by sampling optimiser locations from a learned density, which fixes the decision rule and admits no surrogate. Learned acquisition functions score a finite candidate set with a trained network, but, lacking a label for the query, learn the score by reinforcement learning on previously solved tasks. We propose IQS-BO, a PFN that learns the query decision by supervised learning on synthetic priors. In a single forward pass, IQS-BO predicts the probability that each candidate maximises the objective over the set, and we show that the minimiser of its objective is the posterior probability of this event. The model can be pre-trained without a surrogate for fully in-context BO, or take the predictions of a fixed probabilistic surrogate as additional input, amortising only the decision step. Our method proposes queries at a fraction of the cost of acquisition-based methods, while either matching or outperforming standard BO with Gaussian processes (GPs) and available in-context methods on synthetic and real-world benchmarks. Finally, we propose a mixture prior for pre-training PFNs which combines samples from GPs with functions exhibiting warped inputs, isolated narrow optima, or plateaus that are poorly modeled by stationary kernels common in GP surrogates. We show that pre-training on this prior can lead to improved optimisation performance.
cs.LG / 91 / 2610.01270
Not All Is Lost: Repairing Lossy User Preference States of Personalization Encoders
Abstract
Personalization encoders compress evolving interaction histories into preference states used to rank items or condition text generation. A task head operating only on this state can miss useful evidence that remains in the frozen encoder's cached representations for individual timesteps. We study this recoverability gap and propose REPAIR, which compares cached representations with the current preference state in a compact learned coordinate space. It resolves corrective evidence over extended history, recent interactions, and localized bursts. It then selects which patterns at which timesteps contribute and adds their aggregate correction to the state before the task head. Encoder-host repair reuses representations from the existing forward computation without re-encoding the history. Across MovieLens, PENS, MIND, and Amazon Reviews 2023, training only REPAIR improves MRR and nDCG@10 for all twelve representative recommendation hosts while both encoder and task head remain frozen. Head-only finetuning of the same hosts yields smaller gains. For example, Mamba4Rec on MovieLens gains 3.96 MRR points, compared with 0.19 from head-only finetuning. Rank and temporal diagnostics support a compact, host-dependent corrective structure. In personalized generation, IMPerSumm improves the two reported weighted PerSEval variants, which assess responsiveness to user preference, by up to 25.23%. These results support post-compression state correction and distinguish the availability of preference evidence from its downstream use.
cs.LG / 92 / 2610.01284
Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research
Abstract
Model validation estimates the performance of a complete learning procedure on new data. However, an invalid split can produce an optimistic and stable result. This tutorial reviews hold-out validation, train/validation/test designs, repeated random subsampling, k-fold and repeated stratified cross-validation, leave-one-out and leave-p-out schemes, group-aware validation, and nested group cross-validation. General machine-learning principles are linked to EEG epochs, paired-eye OCT images, repeated clinical measurements, and multicenter data. Eight controlled scenarios compare flawed and leakage-safe designs: seven use locked confusion matrices with auditable metrics, and one uses a reproducible repeated-study simulation. The scenarios cover global feature selection, normalization leakage, dependent records, center mixing, repeated test-set use, and estimator instability. Bias, variance, metric aggregation, uncertainty, and computational cost are also examined. A data-size matrix, a decision tree, and reporting checklists are provided. Reproducible MATLAB templates and scikit-learn counterparts are included. The results show that no validation method is universally best. The independent unit must match the intended deployment target. Every data-dependent operation must also exclude the observations used for performance estimation.
cs.LG / 93 / 2610.01315
EP-Flow: Disordered Crystal Structure Prediction without Site-Level Annotations
Abstract
Generative models have made rapid progress in ordered crystal structure prediction, yet many functional materials are intrinsically disordered, with substitutional mixing, vacancies, or interstitial species controlling their properties. Existing crystal generators either assume deterministic site occupations or require site-level disorder annotations, which are often unavailable when the chemical formula is the primary input. We formulate disordered crystal structure prediction through an Occupancy Distribution Matrix (ODM), a continuous site-by-species representation that unifies ordered crystals, solid solutions, vacancy disorder, and interstitial occupancy. A valid ODM must satisfy coupled site-wise occupancy, mass-conservation, and non-negativity constraints, placing each sample on a formula-dependent transportation polytope. We propose Entropic Polytope Flow (EP-Flow), a marginal-constrained flow matching framework that canonicalizes heterogeneous polytopes into a shared double-centered space, learns a marginal-preserving flow, and recovers feasible occupancies through a Sinkhorn inverse map. By jointly generating occupancies, fractional coordinates, and lattice parameters, EP-Flow achieves state-of-the-art performance on formula-conditioned disordered CSP benchmarks derived from COD and MPDS, substantially outperforming adapted ordered-crystal generators. Analyses further show that EP-Flow recovers sparse and chemically meaningful local disorder patterns rather than merely matching global composition statistics.
cs.LG / 94 / 2610.01317
Prediction-powered Neural Architecture Search
Abstract
Evaluating candidate architectures in neural architecture search (NAS) faces an inherent trade-off: on the one hand, reliable performance labels are limited because training and evaluating architectures is costly; on the other hand, zero-cost proxies (ZCPs) are cheap to compute at large scale but can be noisy. Yet, how to effectively combine these two sources of supervision remains unclear. In this paper, we propose PPNAS, a novel prediction-powered inference (PPI) approach for NAS. PPNAS fuses (1) a small set of architectures with observed performance labels and (2) a large set of architectures with ZCP information. To combine these two sources of supervision, PPNAS exploits the ordinal information provided by ZCPs to construct additional pairwise ranking supervision, while PPI debiases systematic discrepancies between ZCP-based and true performance rankings. We evaluate PPNAS in end-to-end predictor-based NAS, where it achieves state-of-the-art under limited evaluation budgets. To the best of our knowledge, PPNAS is the first prediction-powered approach for label-efficient NAS.
cs.LG / 95 / 2610.01322
Clifford Sheaf Neural Networks
Abstract
We introduce the Clifford Sheaf Neural Network (CSNN), an equivariant sheaf neural network for geometric graphs that places a Clifford algebra on each stalk of a cellular sheaf and transports multivector features along edges. The canonical choice of restriction map for sheaves with algebra-valued stalks is algebra homomorphism. Adding the constraint of equivariance, the naive choice becomes versor conjugation. However, versor conjugation is expressively weak, so we drop algebra homomorphism and arrive at the K-term sandwich. The resulting sheaf Laplacian is positive semidefinite by construction, needs no versor constraint, and still mixes grades. Our main contribution characterizes the resulting family of restriction maps along three axes: which grades a map couples, how much of the endomorphism space it reaches, and how well it is conditioned. The K-term sandwich spans half of the endomorphism space, and in Cl(3, 0, 0) it corresponds to the maps that commute with the central pseudoscalar. The number of terms controls expressivity. CSNN is the reversion member, a first-order model by construction and the grade-mixing corner of this family, developed as a sheaf construction for graph-level equivariant regression.
cs.LG / 96 / 2610.01343
Robust Non-Clairvoyant Scheduling with Classification Models
Abstract
We study the classical single-machine scheduling problem of minimizing the sum of completion times of jobs in a non-clairvoyant setting, where the processing time of each job remains unknown until its completion. This is a hard problem for which no constant competitive algorithm is possible. Inspired by robust optimization and learning-augmented algorithms, we introduce a novel robustness framework that leverages structural information provided by a classification model to overcome this limitation. Specifically, we assume that jobs are partitioned into classes and we have access to the confusion matrix of the classifier, whose entry $(k,\ell)$ indicates the number of jobs predicted to belong to class~$k$ but that actually belong to class~$\ell$. In this manner, we are able to characterize uncertainty as a set of permutations within each predicted class, rather than as a collection of discrete numerical scenarios, avoiding the computational difficulty of classical robust metrics, such as Min-Max and Min-Max Regret. In addition to these worst-case metrics, we also consider the expected objective over all scenarios. We first propose an optimal non-adaptive strategy that is oblivious with respect to all three robust criteria. We then investigate adaptive and randomized algorithms, showing that they can outperform the optimal non-adaptive strategy when the matrix exhibits particular structural properties.
cs.LG / 97 / 2610.01355
Discrete Wasserstein Flows for One-Step Generative Modeling
Abstract
We introduce a new framework for one-step generative modelling on finite state spaces. To extend drifting beyond continuous domains, we use discrete Wasserstein geometry to define a target-relative KL gradient flow over the transitions of a reversible Markov kernel. We realize this probability flow at the particle level through Markov jumps and amortize the resulting transport updates into a latent-conditioned generator, so that the iterative dynamics are required only during training while inference remains one-step. In a controlled setting where the underlying distributions and transport dynamics can be computed exactly, we verify KL dissipation, consistency between the particle dynamics and the probability flow, and the predicted numerical scaling. We further show that a finite-capacity neural generator can track these exact transport targets while retaining one-step generation. These results validate the basic construction and provide a foundation for scaling Discrete Drifting to structured discrete data.
cs.LG / 98 / 2610.01356
Port-Hamiltonian Neural Networks for Systems with Multiple Asymptotically Stable Equilibria
Abstract
Stable port-Hamiltonian neural networks certify asymptotic stability by construction. Yet, their Hamiltonian is a global Lyapunov function with a single global minimum, so they can represent only dynamic systems with {one} attractor. We demonstrate that this excludes even simple systems with energy landscapes forming a double well, and we overcome the restriction by parametrising the Hamiltonian as a {product} of Bregman divergences generated by one input-convex network. We prove that the resulting model is locally Lyapunov stable, that the coexistence of stable equilibria forces additional non-asymptotically-stable equilibria to exist, that all equilibria lie in a bounded region, and under a hyperbolicity assumption that almost-everywhere stability holds. On three systems our approach is able to recover the energy surface characteristics and improve the convergence speed by 1.8$\times$-8.5$\times$.
cs.LG / 99 / 2610.01369
From Redundancy to Minimality: Fixed-Point-Guided Hierarchical Reduction of Learned Piecewise-Linear Dynamics
Abstract
Understanding a nonlinear dynamical system from time series requires not only reproducing its trajectories, but also identifying a simple representation that preserves its essential dynamical structure. Almost-linear recurrent neural networks (AL-RNNs) are piecewise-linear RNNs in which only a subset of units use ReLU nonlinearities, so that nonlinear capacity is explicitly controlled by the number of ReLU units. Their activation patterns define linear regions, represented as symbols, whose observed transitions form a symbolic transition graph. However, directly training AL-RNNs with few ReLU units to realize minimal dynamical representations can be unreliable. We ask whether an AL-RNN with more ReLU units can instead be trained first and systematically reduced to a minimal dynamical representation. We introduce a fixed-point-guided hierarchical reduction procedure that progressively linearizes selected ReLU units, merging neighboring linear regions and graph nodes while preserving distinct symbols containing fixed points (FPs). The resulting reduction tree defines a hierarchy of progressively simpler candidates. Each reduced candidate is initialized from the parent parameters and retrained under guidance from the parent dynamics. We also prove that reproducing $Q$ distinct fixed points requires at least $Q$ FP-containing symbols, providing a certificate of symbol-level minimality when this bound is attained. On the 3-scroll Chua system, direct training with the theoretical minimum of three ReLU units achieves high-fidelity minimal realizations in only 20% of seeds, whereas our learn-reduce-retrain strategy increases the seed-macro success rate to approximately 71% at the same final nonlinear capacity. These results show that redundant nonlinear capacity can serve as a scaffold for discovering and realizing minimal dynamical representations.
cs.LG / 100 / 2610.01373
Learning Commute-Time-Preserving World Models for Planning
Abstract
World models allow agents to plan in latent space by choosing a sequence of actions that most reduces the distance to a given goal state. Thus, planning can benefit from latent representations whose distances mirror commute-times in the environment. The spectral embedding space of the graph Laplacian provides such a representation, if it obeys a specific eigenvalue-dependent scaling. Unfortunately, instantiating the graph Laplacian is intractable in large, continuous environments. Self-supervised learning offers a natural route to such commute-time-preserving embeddings at scale. However, here we show that existing methods, which commonly encourage isotropic representations to prevent representational collapse, tend to degrade the "correct" eigenvalue-dependent scaling, leading to an inaccurate representation of commute times. To address this problem, we introduce Commute-Time-Preserving World Models (CTWMs), combining a latent displacement predictor and a log-determinant regularizer that prevents collapse, which provably recover the correctly scaled Laplacian representation under reversible deterministic dynamics and at the predictor's fixed point. In numerical simulations, CTWM matches or outperforms LeWM, a task-agnostic baseline, on several complex, continuous goal-reaching benchmarks, while using half the parameters.
cs.LG / 101 / 2610.01375
Action-On-Item Preference Flow: A Shared Event Schema for Predictive and Generative Personalization
Abstract
A user's movie, news, and dialogue histories differ in their native actions and outputs, yet each interaction supplies evidence that can update user memory. We study whether these histories can train one reusable update mechanism. An action-on-item schema pairs a mapped interaction role with a content embedding, allowing shared update parameters to operate on separate user states. We establish invariance to native relabeling, bounded state changes under item-embedding perturbations, and a pooled-training bound under explicit compatibility conditions. The Multi-Timescale State Hypothesis (MTSH) specifies how this evidence enters, persists, and is consumed; PerTIDE implements it with action gating, three state-space traces, fusion, and command-conditioned readout. On PENS, the same history encoder supports both next-news prediction and personalized headline generation. In a controlled PENS-to-MovieLens experiment, a frozen source-trained core exceeds an identically structured random core by 15.23 MRR points after fitting the same target consumer. On MIND, PerTIDE retains a 4.12-point MRR advantage over a same-input three-branch state-space control. Action, readout, and trace interventions identify complementary contributions to these gains. Together, the theory and experiments support learning history updates across compatible sources and reusing them through predictive and generative consumers.
cs.LG / 102 / 2610.01377
Minimax Optimal Regret for Causal Logistic Bandits with Counterfactual Fairness
Abstract
We study causal logistic bandits with counterfactual fairness constraints. The causal structure is given through known factual and counterfactual feature maps that share an unknown logistic reward parameter, but the learner observes only factual rewards. Consequently, the directions determining counterfactual feasibility need not be identifiable from the available feedback. The closest prior analyses either omit a coverage condition or impose a comparatively strong one, and do not establish matching lower bounds. We first show that some coverage condition is necessary: without a coverage-type restriction, factually indistinguishable environments with different optimal fair actions force $Ω(T)$ expected joint loss. Under a weaker full-rank condition on the factual covariance pooled across actions, we identify a target-specific information scale $V_\star$ that measures the difficulty of estimating rewards and counterfactual effects from factual feedback. We construct worst-case families satisfying this condition on which every policy incurs expected joint loss $Ω\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}\right)$. We also give an explore--then--exploit procedure tuned using $V_\star$ and an adaptive algorithm that does not require its value. Both algorithms achieve $\max\{R_T,V_T\}=\widetilde{O}\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}+κd/σ_0^2\right)$, where $R_T$ is regret relative to the best fair action and $V_T$ denotes the cumulative stage-wise positive violations. Thus the upper and lower bounds match in their leading dependence on $T$, $V_\star$, and $\min\{\log K,d\}$, up to logarithmic factors.
cs.LG / 103 / 2610.01384
Robust Evidential Learning Through Latent Consistency
Abstract
Reliable uncertainty quantification is essential for deploying deep learning models in high-stakes settings, where out-of-distribution and adversarial inputs can induce confident but unreliable predictions. Evidential Deep Learning provides efficient uncertainty estimates in a single forward pass, but can still assign high evidential strength to inputs that are poorly supported by the learned representation, such as adversarial inputs. We introduce CLEAR, a lightweight, task-agnostic post-hoc method that improves evidential robustness without retraining or altering the base prediction. Using held-out calibration data, CLEAR characterises the group-conditioned geometry of the model's latent space. At inference, it efficiently generates perturbation views directly in the latent space and measures their conflict relative to the calibrated geometry of the predicted group. High latent conflict indicates unsupported evidence, which CLEAR uses to selectively reduce evidential strength while retaining evidence for latent-consistent inputs. On ImageNet$\rightarrow$CUB, CLEAR improves OOD and adversarial AUROC by $+8.29$ and $+5.01$ while running 17.4$\times$ faster than competing post-hoc methods while preserving predictive performance across classification, regression, and object detection benchmarks.
cs.LG / 104 / 2610.01395
AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models
Abstract
Muon improves large-scale training by applying a spectral-norm steepest-descent update to matrix parameters, but practical models also contain parameter blocks that do not fit dense-matrix geometry. One important case is the tied vocabulary table, which appears in language models and other token generators and can receive multiple structurally different gradient sources, from sparse input lookups to dense output-classifier updates. In the reference recipe these blocks are handed to an auxiliary AdamW optimizer, which restores second-moment state and updates the aliased table as a generic tensor. We propose AF-Muon, an AdamW-free extension of Muon that keeps the Muon matrix update for hidden weight matrices while using a support-aware finite-cap linear minimization oracle for tied vocabulary tables and an RMS-normalized update for one-dimensional auxiliary parameters. AF-Muon therefore trains every parameter class with a single first-moment buffer and no second-moment state, saving around 20% optimizer-state memory relative to Hybrid Muon in our benchmark. Across nine tied-token settings - decoder-only language models from 124M to 1B parameters, a fully shared T5-style encoder-decoder, and ImageGPT-style image-token, protein, and sparse-MoE variants, spanning text, image, and protein-sequence data - AF-Muon improves mean validation loss and perplexity over both Hybrid Muon and a SCION-style Sign endpoint. Long-horizon runs and hyperparameter sensitivity studies confirm the gain is robust, and identical-momentum diagnostics attribute it to the finite cap, which preserves more within-row magnitude than Sign while bounding the coordinate concentration of row-RMS. These results identify tied vocabulary tables as a distinct optimizer geometry and yield a robust AdamW-free Muon variant across models, modalities, and architectures, with about 1% step-time overhead in matched training.
cs.LG / 105 / 2610.01399
Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?
Abstract
Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic study of when reuse helps, in which scenarios, and to what extent remains missing. In this work, we study the effectiveness of sample reuse in PPO by instantiating two variants within a multiple importance weighting framework. Both retain the core PPO mechanics, reusing only samples from a window of recent iterations, thereby isolating the effect of data reuse from other factors. The variants, termed wPPO-U and wPPO-BH, employ vanilla importance weights or balance-heuristic-corrected ones, respectively. For both, we derive policy improvement lower bounds providing theoretical grounding for their respective losses. We use them to empirically study when and how data reuse improves sample efficiency or final performance of PPO across continuous control tasks.
cs.LG / 106 / 2610.01425
Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones
Abstract
Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from $p=.3$ to $.5$ to $.7$, while preserving the relative mixture $1{:}2{:}3{:}4$ among the four shared feature levels, increases the final mean accuracy gap from $.185$ to $.331$ to $.527$ across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.
cs.LG / 107 / 2610.01426
Least-time Gradient Flow
Abstract
Prescribing the speed of gradient flow on the risk itself, by the dynamics $\dot w=-u(E(w))\nabla E(w)/\abs{\nabla E(w)}^{2}$, makes the risk $e(t)=E(w(t))$ obey $\dot e=-u(e)$ exactly, whatever the landscape~$E$; the time needed to reach zero risk from $e_0$ is $\int_0^{e_0}\dd e/u(e)$. Minimizing this time alone is ill posed, and we study the regularized problem $\inf\{\int_0^{e_0}(\tfrac\lambda2\abs{u'}^{2}+1/u)\,\dd e:\ u\in H^{1}(0,e_0),\ u\ge0,\ u(0)=0\}$, $λ>0$. We prove that the minimizer exists, is unique, and is a linearly scaled cycloid, and we show that the optimal rate behaves like $u^{*}(e)\sim(9/(2λ))^{1/3}e^{2/3}$ near zero risk: the exponent $2/3$ is the one found in \cite{betti2026holder} by a power-law ansatz, and it lies in the Hölder window $(\tfrac12,1)$ where the arrival is in finite time with vanishing weight speed. The proof follows the classical route: existence by the direct method, uniqueness by strict convexity, positivity of the minimizer away from the origin, and the explicit integration of the Euler-Lagrange equation.
cs.LG / 108 / 2610.01435
Distillation of Tabular Foundation Models into Efficient Predictors
Abstract
Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .
cs.LG / 109 / 2610.01445
ibUMAP: Coherent and Scalable Field Evaluation for UMAP Optimization
Abstract
UMAP achieves scalable layout optimization through stochastic negative sampling. However, this stochasticity can lead to unstable embeddings across reruns and downstream reuse, as the estimated repulsive forces depend on the ordering of sampling events. We present ibUMAP, a coherent field-based alternative that evaluates attraction and repulsion from a shared embedding snapshot and applies them synchronously. Its degree-weighted repulsive field is motivated by the conditional expectation of negative sampling for a fixed embedding and represented by three scalar moments, which are evaluated efficiently on CPUs and GPUs using an interpolation-based FFT scheme. This formulation avoids explicit all-pairs computations while inducing optimization dynamics that differ from those of standard online UMAP. Controlled experiments show that synchrony and kernel capping alter the local-global fidelity trade-off, whereas FFT evaluation produces small average changes in final quality. End-to-end benchmarks show median speedups of 3.29x unseeded and 5.79x seeded over umap-learn on CPU, and 1.44x over cuML on million-scale datasets under unseeded GPU execution. These gains accompany greater run-to-run stability and measurable fidelity trade-offs.
cs.LG / 110 / 2610.01453
Repurposing Obsolete Representations for Post-Deployment Adaptation
Abstract
Deep neural networks are increasingly deployed in long-lived systems, where task requirements may change after training. In such settings, part of the original output space may become obsolete: a class, prediction region, or learned behaviour may no longer be valid. Existing approaches either leave the obsolete behaviour intact or require fine-tuning, which can be expensive. We propose Deep Repurposing (DR), a post-hoc framework for adapting models under task obsolescence. DR estimates the latent geometry of obsolete and retained regions, removes obsolete-supporting components, and reallocates retained-compatible evidence through an analytic repair map without gradient updates. This yields repaired predictions and representations in which obsolete regions no longer act as valid outputs, while useful obsolete structure can support the retained task. Across multiple task settings, DR removes obsolete behaviour while preserving retained utility. More importantly, across classification benchmarks, DR matches or exceeds competing unlearning and editing baselines in retained accuracy, eliminates obsolete predictions, and adapts up to $60\times$ faster than competing unlearning methods.
cs.LG / 111 / 2610.01456
Streaming algorithms for robust max-min diversification
Abstract
Given a set of $n$ points $X$ in a metric space and an integer $k$, max-min diversification aims to select $k$ points of $X$ maximizing their minimum pairwise distance. This objective function is however highly vulnerable to noisy points. In[Amagata, AAAI23], a robust formulation is proposed which addresses this vulnerability by excluding solutions containing any of $z$ outliers, defined as the $z$ points in $X$ with the largest nearest-neighbor distances. That paper also presents a coreset-based streaming algorithm for the new formulation, based on a suitable inlier-outlier separation assumption. However, we identify three shortcomings in the algorithm by [Amagata, AAAI23]: its coreset construction requires an offline computation over $X$, which needs memory linear in $n$, in stark contrast with the typical goals of stream processing; the one-pass procedure used to extract the solution from the coreset may return fewer than $k$ points (hence, an unfeasible solution) because it permanently discards points too far from the current solution; and its outlier-exclusion guarantee is only probabilistic and weakens as the coreset size shrinks. In contrast, we present a deterministic coreset-based algorithm that, under a natural inlier-outlier separation assumption (similar to the one used in [Amagata, AAAI23]), returns exactly $k$ inliers which are a $(2+\varepsilon)$-approximate solution, for any $\varepsilon>0$, thus only $\varepsilon$ above the best polynomial-time sequential approximation, even without outliers. Its one-pass streaming implementation adapts obliviously to the dataset's doubling dimension $D$ and, for wide ranges of $k$, $z$, $\varepsilon$, and $D$, it uses memory independent of $n$. For sufficiently long streams, its amortized update time is proportional to the coreset size, thus also independent of $n$.
cs.LG / 112 / 2610.01459
Tight Transition Time Bounds for Separable Logistic Regression at the Edge of Stability
Abstract
We study logistic regression on linearly separable data under gradient descent with a large constant stepsize $η$. Such dynamics may exhibit a characteristic Edge of Stability phenomenon, in which the loss initially oscillates before transitioning to a stable phase of monotone decrease. Existing work provides a tight $Θ(1)$ bound in dimension $d=2$ as $η\to \infty$ and conjectures a bound independent of $η$ in arbitrary dimensions $d\geq 2$. In this paper, we disprove this conjecture by showing that, for every fixed sample size $n\geq 2$ and sufficiently small margin $γ$, the worst-case transition time is $$Θ\!\left((\logη)^{\min\{n-2,d-2\}}\right)$$ uniformly over $d\geq2$. The key challenge in establishing a tight bound is that the sample contributing most strongly to the gradient can change repeatedly across iterations. To address this issue, we control such changes by induction on dimension and sample size, and construct matching hard instances.
cs.LG / 113 / 2610.01515
FedMIX-P: Mixing Local and Global Preconditioners for Federated Vision and Language Model Training
Abstract
Adaptive preconditioners accelerate model training, but heterogeneous client geometries can bias federated updates even when gradients are evaluated at the same model. Round-start synchronization alone cannot prevent this mismatch from reappearing during local training. We propose \texttt{FedMIX-P}, which mixes shared and local preconditioners at every local step, retaining local adaptation while reducing mean-squared operator mismatch by a factor of $λ^2$. For smooth nonconvex objectives with stochastic gradients and partial participation, we establish an $O(R^{-1/2})$ stationarity bound using suitable stepsizes and a horizon-dependent mixing weight, without requiring local preconditioners to converge to one another. A two-client counterexample shows that fixed positive mixing can preserve a nonstationary fixed point. The theory covers bounded linear symmetric positive-definite preconditioners. Experiments with SOAP, Sophia, and Muon variants across vision and language tasks show improvements over corresponding local optimizers, including accuracy gains of up to $19.47$ percentage points and lower validation loss for 60M--350M language models. Full nonlinear and momentum-based updates require separate analysis.
cs.LG / 114 / 2610.01519
Auto-Formalizing Neuro-Symbolic Predictors
Abstract
Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified constraints, making them particularly suitable for high-stakes applications where compliance with domain knowledge is essential. A key bottleneck in this paradigm is the acquisition of symbolic constraints: encoding domain knowledge into logical formulas remains a manual and expert-intensive process. In this work, we investigate the extent to which auto-formalization via LLMs can systematically translate textual knowledge into symbolic knowledge that can be plugged into NeSy predictors. To this end, we introduce auto-nesy-bench, a new benchmark for evaluating constraint formalization and its impact on downstream accuracy of NeSy predictors. Through an extensive evaluation across several domains, we find that LLMs can formalize constraints to a meaningful extent, generating formulas that are often similar to those provided by human experts. Moreover, when the generated formulas are syntactically valid, they can lead to high-quality downstream predictions. The code and benchmark are available at https://unitn-sml.github.io/auto-nesy-bench/.
cs.LG / 115 / 2610.01527
Exact Distinguishability in Non-Markovian Decision Processes
Abstract
Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.
cs.LG / 116 / 2610.01530
Calibrating Prediction Timeliness Through Multi-Objective Hyperparameter Optimization for Remaining Useful Life Prediction
Abstract
In predictive maintenance, early and late RUL prediction errors carry asymmetric consequences, yet hyperparameter optimization typically targets a single accuracy metric that treats both directions equally. This study treats the optimization objective itself as a design variable. Five architectures (MLP, LSTM, XGBoost, TCN, and Transformer) are evaluated under three regimes: single-objective maximization of $R^2$, single-objective minimization of the NASA scoring function, and a multi-objective formulation that jointly optimizes both criteria. The multi-objective search employs NSGA-II with Entropy-CRITIC weighting for Pareto selection. Seventy-five model-dataset-strategy combinations are assessed on the NASA C-MAPSS turbofan and BackBlaze hard-disk drive benchmarks. On C-MAPSS, all strategies achieve comparable accuracy ($R^2 \approx 0.89$), yet multi-objective optimization reduces directional imbalance by approximately 33%, improving calibration of early versus late predictions. Model rankings prove configuration-dependent, with simpler architectures frequently outperforming deeper temporal models. On BackBlaze, the objectives shift from complementary to conflicting, producing divergent Entropy-CRITIC weights and a substantial generalization gap (best $R^2 \approx 0.34$). These results demonstrate that the optimization objective materially shapes prognostic behavior and that multi-objective search provides a practical mechanism for calibrating prediction timeliness in RUL modeling.
cs.LG / 117 / 2610.01566
Towards Optimal Policy Improvement
Abstract
Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced MDP, characterizing planning with an explicit or implicit model as a path towards optimal policy improvement. Because practical methods commonly solve such induced problems through iterative improvement in the form of greedification, we take steps towards optimal greedification under the central practical constraint of approximate evaluation. We formulate greedification under this constraint as probabilistic decision-making under uncertainty and derive a novel operator that is optimal with respect to the resulting objective. Empirically, the operator and its practical gradient-based approximations improve aggregate performance across GumbelAlphaZero, SAC, ReBRAC and Generalized Policy Iteration, in experiments spanning discrete and continuous actions, model-based and model-free, online and offline RL.
cs.LG / 118 / 2610.01579
Beyond Pointwise Error: A Multi-Metric Evaluation of Spatial Climate Downscaling
Abstract
Climate downscaling aims to reconstruct fine scale spatial fields from coarse resolution inputs. Evaluating the quality of these reconstructions is challenging: low pointwise error can come at the cost of fine scale variability, while realistic spatial variability can be achieved with inaccurate local structures. The evaluation metric can therefore change which method appears to perform best. This work presents a multi metric benchmark comparing five spatial downscaling methods on ERA5 temperature, wind, and precipitation fields. Five criteria assess complementary properties: pointwise error, structural similarity, distribution error, spectral error, and gradient error. The results reveal a systematic trade off between spatial fidelity and fine scale variability. Some methods perform best on pointwise and spatially aligned metrics, but lose high frequency content, while others preserve substantially more spectral variability at the cost of less accurately positioned local structures. Consequently, method rankings change across metrics and variables. These results show that there is no single best downscaling method. Multi metric evaluation is therefore essential for assessing which properties of a climate field are preserved.
cs.LG / 119 / 2610.01590
Two Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize
Abstract
Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.
cs.LG / 120 / 2610.01601
Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
Abstract
Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate's score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}, and the \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}} is available on Hugging Face.
cs.LG / 121 / 2610.01619
Exposing the Cost of Deep Learning Audio Development
Abstract
The environmental impact of deep learning has attracted increasing attention over the past decade. Existing studies mainly focus on the energy and carbon emissions of model training and inference, while the whole development phase is often overlooked. Yet, architecture prototyping and intensive experiments are conducted during this stage, which is highly energy-demanding. In this article, we propose a methodology to estimate these costs, based on activity logs from the Grid5000 shared computing platform used by the LORIA laboratory. As a case-study, we focus on audio projects developed in the Multispeech research team. We evaluate the overall energy cost of four projects, and we compare them to those of training the reported models. Our results show that the energy required for the development phase is 3 to 256 times greater than that required to train the best-performing model alone. These results advocate for a more systematic reporting of energy consumption across the entire life cycle of deep learning-based audio projects.
cs.LG / 122 / 2610.01633
Generalization in Neural Networks Through the Lens of Magnitude Potential
Abstract
Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of {\em magnitude potential} -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.
cs.LG / 123 / 2610.01638
FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices
Abstract
Federated learning (FL) on heterogeneous edge devices must jointly accommodate unequal resource budgets and domain-shifted local data. Existing resource-adaptive methods decide how much of a model each client trains but not where retained capacity should reside or how it should be shared, whereas federated domain-generalization methods usually assume a shared full architecture. Uniform compression can therefore discard high-utility channels, and a single aggregation path can mix transferable features with domain-sensitive updates. We propose FedSAP, a domain-aware heterogeneous FL framework that casts structured pruning as budget-constrained tri-state channel allocation. FedSAP converts each keep ratio into non-uniform layer budgets, assigns stable channels to a Global pool, useful domain-sensitive channels to pseudo-domain-specific Private pools, and low-utility channels to a Dropped state. This partition lets broadly useful features benefit from cross-client pooling while isolating domain-sensitive updates from incompatible clients. Domain-Guided Assignment infers pseudo-domains from shallow-gradient similarity, while Type-Matched Aggregation restricts each channel to its intended sharing scope. Across three random seeds, FedSAP reaches 76.00% and 72.67% mean global accuracy on Digits and Office-Caltech, exceeding the strongest baseline by 1.70 and 4.92 percentage points while supporting client pruning ratios of up to 80% across heterogeneous clients.
cs.LG / 124 / 2610.01641
MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
Abstract
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.
cs.LG / 125 / 2610.01645
Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction
Abstract
Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.
cs.LG / 126 / 2610.01649
CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations
Abstract
Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.
cs.LG / 127 / 2610.01652
Iterative Policy Refinement through Semantic Rollout Analysis
Abstract
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
cs.LG / 128 / 2610.01663
pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows
Abstract
Biomolecular therapeutics often start from known sequences and require targeted editing to improve multiple properties while satisfying hard biochemical and manufacturability constraints. However, existing generative methods do not jointly support multi-objective optimization, hard feasibility, and sequence editing in discrete, variable-length biological spaces. In this work, we introduce Pareto-Constrained Molecule Editing (pCoMole), a framework built on discrete flow matching that steers a pre-trained Edit Flow toward user-specified preferences while enforcing terminal feasibility. pCoMole defines a feasibility-gated terminal distribution using an augmented Tchebycheff utility and realizes the resulting preference tilt through a Doob-h transform of the underlying edit process. To make this construction practical, we approximate the required harmonic function using short Monte Carlo rollouts over candidate edits, yielding an efficient guided editor with provable preference consistency. We validate pCoMole by shrinking GFP while retaining fluorescence-related properties, shortening diverse Cas9 orthologs while preserving PAM specificity, and compressing peptide binders into short peptidomimetics that optimize seven drug-related properties under hard constraints. In wet lab testing, two 229-residue pCoMole-designed eGFP variants retained clear green fluorescence in BL21 cells after 10 deletions, with either one or two substitutions. Together, pCoMole enables constraint-aware, Pareto-aligned editing of biomolecular sequences in discrete, variable-length spaces.
cs.LG / 129 / 2610.01674
Invent a Dataset: Measuring dataset generation abilities with zero seed
Abstract
Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.
cs.LG / 130 / 2610.01685
MiLoop: Selective Memory Propagation for Neural Combinatorial Optimization
Abstract
Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.
cs.LG / 131 / 2610.01692
Artifact Annotations Partially Substitute for Per-User Calibration: SAFE-EDA and a Normalization-Controlled Evaluation of Wrist-EDA Affect Recognition
Abstract
Wrist electrodermal activity (EDA) differs in amplitude from one person to the next, so affect-recognition models normalize their input before classification. Studies that test such models on held-out subjects seldom report where the normalization statistics come from, yet statistics computed from the held-out subject's own recording give the model information that a device does not have when it is first worn. We asked how this choice alters the measured benefit of pretraining. A compact convolutional network, SAFE-EDA, was pretrained on expert artifact annotations from 43 subjects and compared with the same network trained from scratch on the Wearable Stress and Affect Detection (WESAD) dataset (15 subjects, leave-one-subject-out), with two normalization sources crossed with four window hops. When the statistics came only from training subjects, pretraining raised macro-F1 by 0.078 to 0.227; when they came from the held-out user's full recording, the gain fell to between 0.020 and 0.050 and was no longer significant. Artifact supervision was far more useful than self-supervised pretraining on the same recordings (0.078 versus 0.008). Across 13 configurations in two datasets, the pretrained network was better in 12, but on the second dataset (26 subjects) per-user normalization increased the gain instead of reducing it, so the interaction depends on the data. Only five of 50 published WESAD studies state which data were used for normalization. Reporting this choice is necessary to separate first-use performance from performance after calibration.
cs.LG / 132 / 2610.01712
In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners
Abstract
In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.
cs.LG / 133 / 2610.01721
Anomaly Detection and Localization for the Pantograph-Catenary System
Abstract
Monitoring the Pantograph-Catenary System (PCS) provides insight into the health conditions of the pantograph and the railway infrastructure. Recent industrial solutions trace the pantograph's contact wire height and stagger (PCS height/stagger) using video monitoring through convolutional neural networks. However, these solutions do not account for the train route's geographic location. Therefore, in this paper we propose a novel framework for 1) localization of the PCS height/stagger by alignment with the nominal GPS coordinates of the reference route, and 2) collective anomaly detection to evaluate the health conditions of the PCS. We apply and assess the localization and detection performance of the methodology to a case-study based on a real-world industrial dataset provided by a railway transportation company, which includes the PCS height/stagger of several train journeys across Italian railway routes.
cs.LG / 134 / 2610.01725
RelICL: Training-free Relational Learning with Tabular Foundation Models
Abstract
Tabular foundation models achieve state-of-the-art performance on single-table tasks without any training. Recent work suggests that they are also well-suited for relational learning via deep feature synthesis (DFS), which flattens a relational schema into a single table by adding aggregates of the other tables' columns as features. This approach is appealing because it directly benefits from improvements to or customization of the underlying tabular foundation model. In this paper, we identify two key problems with DFS: feature explosion and interaction blindness. The first problem arises because the number of DFS features grows quickly as the schema becomes more complex, limiting scalability and performance. The second problem arises because column-wise aggregates do not account for feature interactions, limiting performance. We propose and explore an alternative method termed RelICL, which keeps the benefits of DFS but alleviates these two problems. At its heart, RelICL propagates and fuses information step by step through the schema graph, using the same tabular foundation model that is eventually used for prediction to do so. In our experimental study using RelBench tasks, RelICL was on par with the strongest approach based on deep feature synthesis.
cs.LG / 135 / 2610.01728
Removing spurious minima for planar features by skip connections
Abstract
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.
cs.LG / 136 / 2610.01729
Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning
Abstract
Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at https://github.com/ZihanLiummyycc/FSG-RL.
cs.LG / 137 / 2610.01751
Evidence-Gated Research: Statistically Controlled Model Adoption in Adaptive Search
Abstract
Adaptive model search is path dependent: once a challenger is adopted, it becomes the reference from which later candidates are generated. A statistically unsupported replacement can therefore alter hypotheses that have not yet been proposed. We introduce Evidence-Gated Research (EGR), a statistical adoption layer for moving-incumbent search. EGR freezes each challenger before decision evidence is revealed, builds anytime-valid evidence across a predeclared set of environments, routes evidence predictably toward unresolved components, composes a persistent candidate e-value, and passes that e-value to an online controller. Under explicit conditional-validity and predictability conditions, the resulting procedure controls false discovery rate for the declared all-environment adoption target even though earlier adoptions change later challengers. In a 5,000-trajectory closed-loop benchmark, development-only e-LOND attains persistent FDR 0.621, whereas no persistent false-adoption path is observed for the audited EGR variants in that finite run. In matched replay over 600 challenger--incumbent pairs, Stagewise EGR preserves fixed-anytime alternative crossing decisions while using 56.1% less decision evidence at the representative threshold. A three-environment public-data study and a 40,000-sample controlled neural benchmark reproduce the evidence-efficiency pattern. These results identify model replacement as a distinct statistical control point in adaptive model development.
cs.LG / 138 / 2610.01765
Physics-Refined Spatiotemporal Forecasting on Open-Boundary Hydrologic Graphs
Abstract
Spatiotemporal forecasting on hydrologic graphs is especially prone to instability in open-boundary systems, where the forecast domain exchanges fluxes with an unobserved exterior. In such systems, boundary nodes receive external forcing, e.g., upstream inflows in rivers or tidal signals in coastal regions, that is typically unavailable at prediction time. The absence of this information can compound errors as forecasts unfold in an autoregressive fashion, leading to inferior long-horizon performance. This paper dissects this instability issue by exploring two questions. 1) What boundary forcing enters the forecast domain when information beyond the boundary is missing? 2) How should this forcing propagate through the domain without incurring error amplification under autoregressive rollout? To address both, we propose a new computing framework comprising two key components. First, to compensate for the boundary forcing, our framework learns ghost node proxies from the boundary and interior nodes, striving to approximate unobserved external inputs. Second, to control error accumulation from these learned proxies, we leverage two physics refiners. In particular, one refiner enforces local consistency by aligning ghost proxies with their two-hop neighbors (i.e., boundary nodes and their immediate interiors). The other refiner enhances global stability by correcting the model forecasts through a physics-guided graph neural operator, reducing long-horizon numerical drift. Two real-world hydrologic graphs are employed for empirical evaluation. Comparative results show that our proposal enjoys higher prediction accuracy and long-horizon stability over both learning-based and physics-informed model competitors.
cs.LG / 139 / 2610.01786
Inferring Multi-Timescale Neural Dynamics with Switching Linear Dynamical Systems
Abstract
Neural activity often exhibits multiple timescales that can vary with behavioral states and task conditions. Identifying these timescales from neural recordings is important for better understanding neural computation and function. However, traditional approaches based on autocorrelation fitting are difficult to scale to high-dimensional population recordings and can become unreliable when neural dynamics change with behavior. State-space models have been a powerful framework for modeling high-dimensional neural population activity through latent dynamical systems, but standard formulations and inference methods do not explicitly account for multiple timescales and therefore do not guarantee accurate recovery of the underlying temporal structure. Motivated by these questions, we introduce the Multi-Timescale Switching Linear Dynamical System (MTS-SLDS), a framework for identifying regime-specific latent timescales from continuous or spiking neural observations. MTS-SLDS combines a multi-lag moment initialization, which captures temporal structure across multiple observation lags, with \textit{regime-conditioned} Laplace-EM inference, which reduces mixing of dynamical statistics across uncertain regimes. Characteristic timescales can then be extracted directly from the eigenvalues of the learned latent transition matrices. In synthetic and neural experiments with Gaussian and Poisson spike observations, MTS-SLDS accurately recovers timescales and switching structure over multiple datasets.
cs.LG / 140 / 2610.01788
SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
Abstract
As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
cs.LG / 141 / 2610.01801
A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings
Abstract
Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
cs.LG / 142 / 2610.01819
MECHVAR: Variance-Guided Mechanism Discrimination for Autonomous Machine Learning Experiment Selection
Abstract
Benchmark gains are often mechanism-ambiguous: reproducing an improvement does not by itself identify why it occurs. We study finite-library mechanism discrimination, where posterior-weighted candidate mechanisms, executable probes, and a limited experimental budget define a sequential experiment-selection problem. MECHVAR selects the next probe by maximizing the posterior-weighted variance of its predicted responses. Under a shared-Gaussian predictive model, this score is exactly proportional to the classical Box--Hill posterior-weighted pairwise-KL criterion, yet it admits O(KE) vectorized rescoring and a transparent additive audit over mechanism pairs. A local expansion further links the score to expected information gain (EIG) when predicted response separations are small. In a 25-block stress audit, MECHVAR outperforms confirmation-first in several moderate misspecification regimes, while its primary comparisons with EIG remain statistically unresolved. In a held-out Digits loop, normalized mechanism-identification AUC is 0.8975 for MECHVAR, 0.7825 for a score-greedy policy, and 0.9092 for EIG. At K = 100, E = 200, median single-thread full-library scoring is 10.36 microseconds for MECHVAR versus 57.69 ms for six-node quadrature EIG in the recorded environment. MECHVAR therefore provides a lightweight, auditable acquisition rule for finite-library experiment selection when a shared predictive scale is a defensible approximation.
cs.LG / 143 / 2610.01827
Scientific Discovery under Validation Congestion via Multi-Fidelity Pairwise Rankings
Abstract
Modern computational methods can now propose candidate molecules, materials, and other scientific designs at an unprecedented scale, creating a validation congestion where candidates are abundant, but experimental capacity to physically evaluate them remains scarce. Discovering novel scientific designs has therefore become increasingly dependent on curation: selecting a small set of promising designs for slow and costly experiments. Existing curation methods typically rely on data-driven regression models that predict absolute scores, but training these models requires substantial experimental data to begin with. Yet, useful curation signals do not have to take the form of absolute measurements, as scientific design discovery is often comparative in nature. Here, we propose that curation can instead be primarily driven by expert pairwise rankings, which are substantially easier to gather. The expertise can come from computational tools or human input of multiple levels of fidelity, ranging from empirical rules of thumb to agentic workflows and experienced scientists. We introduce PRISMS, a framework that uses pairwise rankings from one or more experts, potentially spanning multiple levels of expertise, to identify the most promising candidates without relying on data-hungry regressors. When experts differ in fidelity and cost, PRISMS escalates pairwise queries from lower- to higher-fidelity rankers based on a Fisher-information criterion. In iterative screening that selects designs from fixed drug discovery libraries, PRISMS achieves 50% top-10 discovery recall in ~42% fewer rounds than regression-only active learning, and in ~15% fewer rounds than the ranking-based method with no selective escalation. In optimization that generates new designs without restriction to a predefined library, PRISMS achieves ~18.8% higher hypervolume than the Bayesian optimization baseline.
cs.LG / 144 / 2610.01831
Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention
Abstract
Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix $α$ at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming $α$ in the first block alone improves every ICL configuration we test.
cs.LG / 145 / 2610.01889
Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
Abstract
Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.
cs.LG / 146 / 2610.01892
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
Abstract
Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.
cs.LG / 147 / 2610.01903
Higher-Order Positional Encodings for Graph Representation Learning
Abstract
Many real-world systems exhibit higher-order interactions among groups of entities that cannot be captured by pairwise relationships alone. Graph Transformers and Graph Neural Networks increasingly rely on positional encodings to enrich graph representations, yet existing positional encodings are computed solely from the original graph and therefore cannot directly capture observed higher-order interactions. Topological Deep Learning addresses this limitation by lifting graphs to simplicial complexes, but typically requires performing message passing or attention on higher-order neural network representations. We introduce a representation learning paradigm that enriches graph representations with higher-order topology through positional encodings, enabling standard graph learning models to exploit lifted incidence structure without modifying the backbone. We derive a theoretical characterization of the expressivity of higher-order positional encodings, proving that node-level operators induced by higher-order lifts can mix graph Laplacian frequencies in ways that scalar graph spectral filters cannot. Guided by this theory, we instantiate higher-order positional encodings using Hodge Laplacians derived from clique complexes. Experiments with Graph Transformers on ZINC and controlled synthetic benchmarks demonstrate improvements in predictive performance, while a fixed-1-skeleton experiment shows that the pipeline can transmit higher-order information when cells are supplied independently of the graph. Together, our results establish higher-order positional encodings as a principled bridge between graph positional encodings and topological deep learning.
cs.LG / 148 / 2610.01908
Same Reward, Different Skills: When Multimodal RL Learns to Look
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
cs.LG / 149 / 2610.01951
Sharp Non-Asymptotic Analysis of the Penalized Challenger in $β$-EB-TCI for Bernoulli Bandits
Abstract
Top-two algorithms are simple and effective for fixed-confidence best-arm identification, but their sharp non-asymptotic behavior is still not well understood. We study this problem for Bernoulli bandits through $β$-EB-TCI, the empirical-best top-two rule of Jourdan et al., whose challenger is chosen using a Bernoulli transportation cost with a logarithmic count penalty. We prove that, after the empirical leader has become the true best arm and its sampling fraction stays close to $β$, the stopping time is $T_β^{\star}(μ)\log(1/δ)$ up to lower-order concentration terms. We also show that, in this regime, every challenger is sampled linearly often. Thus, for the original algorithm without forced exploration, the main remaining difficulty is to control when the empirical leader becomes permanently correct. These results imply a non-asymptotic high-probability bound for all Bernoulli instances with a unique best arm. If the algorithm satisfies a finite-mean sufficient-exploration condition, the bound further yields the sharp expected sample complexity. In particular, this gives the sharp expectation result for the unguarded Bernoulli rule when all arm means are pairwise distinct, using the sufficient-exploration result of Jourdan et al. Finally, if we add a mild forced-exploration rule that contributes only $O(\sqrt{Kt})$ pulls up to time $t$, we obtain a self-contained expected sample-complexity theorem for any number of arms under the unique-best-arm assumption. We also identify a limitation of proof strategies that try to handle equal suboptimal means through a single index-comparison argument.
cs.LG / 150 / 2610.01955
Do Your Own Research: Learning to Forecast by Learning to Search
Abstract
Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.
cs.LG / 151 / 2610.01962
SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning
Abstract
The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.
cs.LG / 152 / 2610.01974
Sim+Real: Joint Simulation - Experiment Training Improves Balanced Prediction in Physical Systems
Abstract
Simulation and experimental measurements provide complementary data for learning spatiotemporal physical systems, but standard simulation-to-experiment fine-tuning optimizes only the experimental objective after transfer and can degrade simulation performance. We formulate simulation--experiment prediction as a multi-objective learning problem with domain-specific simulation and experimental risks. On four fluid systems from RealPDEBench and two model capacities, we compare Simulation only, Experiment only, Sim$\rightarrow$Exp, and Joint training, evaluating every final model on both held-out domains. Sim$\rightarrow$Exp tends to specialize more strongly to experimental data at the cost of simulation-domain forgetting. Joint training consistently achieves the best balanced performance over a broad range of simulation--experiment evaluation weightings, while substantially improving simulation retention over Sim$\rightarrow$Exp. Joint also better preserves simulation-only fields absent from experimental measurements. Project page: https://mahindrautela.github.io/morph.
cs.LG / 153 / 2610.01980
The Curvature of Regret in Contextual Linear Optimization
Abstract
Decision-focused learning for linear optimization is complicated by the discontinuity of the optimizer, where small cost errors may leave the decision unchanged or move it to a different vertex. We show that this non-smooth pointwise behavior becomes locally quadratic after averaging over the data distribution, and we derive the curvature in closed form, specifically, a matrix-valued measure supported on the walls of the normal fan. This measure depends only on the feasible set, with the data distribution entering only as a weight. We then offer a tractable approximation for this curvature, computable with just one projection to the feasible set. We prove that the approximation weakly converges to the true population curvature. We offer one application of our findings, a decision-aware scenario generation method for expected-cost linear optimization. Our experiments test the quadratic and weak convergence laws and show a 30.8% regret improvement over uniform allocation on battery arbitrage.
cs.LG / 154 / 2610.01981
Universal interpolation for deep residual self-attention networks
Abstract
Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of $N$ sequences of $n$ tokens to any other collection of $N$ sequences of $n$ tokens. Crucially, these transformations are \emph{fixed independently of the input and output} collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.
cs.LG / 155 / 2610.02012
Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos
Abstract
Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.
cs.LG / 156 / 2610.02013
BranchIP: Learning Adaptive Equivariant Computation for Interatomic Potentials
Abstract
Equivariant machine learning interatomic potentials (MLIPs) have revolutionized atomistic modeling, but accurate treatment of complex materials and molecular systems demands expensive models. This limits simulation length- and time-scales, with tensor products a key computational bottleneck. The recent emergence of foundation-scale MLIPs further exacerbates this challenge. We present Branch Interatomic Potential (BranchIP), a single-model framework for learned adaptive tensor product computation, trained with a novel distillation loss. In our experiments on two systems of physical interest, a heterogeneous catalysis system and a proton-conducting solid acid electrolyte, BranchIP accelerates MLIPs across model sizes by up to $2.4\times$ while reducing memory usage by up to $2.6\times$. This is achieved while maintaining physical fidelity. Furthermore, the learned adaptive computation provides model interpretability by revealing which interactions demand deeper computation and showing how computational depth relates to chemical complexity and dynamics.
cs.LG / 157 / 2610.02015
On Language Drift during RLVR Post-Training
Abstract
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
cs.LG / 158 / 2610.02033
Relative Transitions, Not Absolute Destinations: A Transfer-and-Ground Framework for Target-Trajectory-Free Human Mobility Generation
Abstract
Individual mobility trajectories support urban analysis and location-based services, yet most trajectory generators require observations from their deployment city. This assumption excludes precisely the cities where trajectories are unavailable even though points of interest (POIs) and their attributes can be obtained from public maps. We study target-trajectory-free generation: learning from POIs and trajectories in source cities while utilizing only POI coordinates and categories in a target city, with no target trajectory or trajectory-derived statistic available for training, model selection, or generation. Existing trajectory generators typically predict absolute destinations, entangling reusable movement behavior with city-specific POI identities and spatial layouts. Our core insight is to replace this city-bound output with context-conditioned relative transitions. We propose Nomad, a transfer-and-ground framework that separates learning how people move from determining where those movements are realized. Specifically, a history-conditioned flow-matching model learns from source trajectories a transition prior over semantic displacement between POI contexts, geographic displacement, and elapsed time; at inference, a behavior graph and an exploration--return walk ground sampled transitions onto the target POI map. This factorization enables a direct test of representation level transferability without assuming invariance of the full mobility distribution. Extensive experiments across ten cities and 14 transfers show that Nomad outperforms adaptation baselines in trajectory fidelity and downstream utility, lowering the average error over the best baseline of each metric by about 15% in distributional fidelity and about 3% in downstream utility.
cs.LG / 159 / 2610.02043
Distributionally Robust Schrödinger Bridge
Abstract
Schrödinger bridge (SB) learns stochastic transport between prescribed initial and target distributions. When the initial distribution shifts at test time, the learned dynamics can fail to recover the target distribution. We introduce the Distributionally Robust Schrödinger Bridge (DRSB), which learns a single controller that accounts for uncertainty in the initial distribution. The DRSB objective consists of control energy and a KL penalty between the resulting terminal distribution and the target distribution. DRSB seeks a single controller that minimizes the worst-case value of this objective as the initial distribution varies within an ambiguity set around the nominal distribution. We derive an exact variational formulation of this objective and connect its fixed-terminal-cost subproblem to stochastic optimal control and distributionally robust optimization. This formulation motivates an alternating algorithm that updates the adversarial initial distribution, estimates the terminal log-density ratio, and trains the controller. We develop Wasserstein and Sinkhorn variants using stochastic control optimality conditions to approximate the gradients required for adversarial updates. Experiments on two-dimensional transport tasks and image-to-image translation show improved robustness to input perturbations relative to standard SB, with a tradeoff in nominal performance. On Gaussian mixture transport, Sinkhorn DRSB also achieves lower mean sliced Wasserstein distance than fixed-level noise augmentation at both tested unseen noise levels.
cs.LG / 160 / 2610.02058
Foundations without Fundamentals: Zero-Shot Blind Spots in Time Series FMs
Abstract
Despite the success of Time Series Foundation Models (TSFMs) on broad benchmarks, their ability to internalize basic temporal logic, especially in settings supported by exogenous covariates, remains under-examined. We introduce SimpleTimeBench, a diagnostic univariate and multivariate "unit test" suite for primitives such as monotonic trends, periodic signals and leading indicator covariates, scenarios where near-perfect forecasts should be trivial. Surprisingly, prominent multivariate TSFMs (Chronos-2, Moirai and Toto) frequently produce suboptimal zero-shot forecasts for these inputs. While fine-tuning Chronos-2 improves its behaviour on specific tasks, we show that this adaptation degrades performance on other fundamental patterns rather than enhancing its generalizable foundational capabilities. This reveals a gap between pre-training scale and basic temporal reasoning, suggesting that current TSFMs could potentially lack the inductive biases needed to capture simple predictable functions. We further demonstrate that these failures are not merely synthetic curiosities: they persist in real-world sensor forecasting, where TSFMs consistently underutilize leading indicators available in observed covariates. This inability to capture simple relationships limits the practical utility and reliability of current multivariate models.
cs.LG / 161 / 2610.02067
Learn the Directions, Normalize the Gains: Post-Training Normalization for LoRA
Abstract
While Low-Rank Adaptation (LoRA) enables efficient task specialization, its learned updates can compromise capabilities beyond the target task. We identify \textbf{adaptation imbalance}: a few singular directions dominate the trained update, leaving its performance sensitive to how gains are allocated. We argue that \textbf{learning where to adapt does not ensure that adaptation gains are well balanced}. This motivates \textbf{LoRA-Norm}, a post-training normalization method that retains learned directions while rebalancing their gains. LoRA-Norm combines spectral rebalancing, a fixed nonlinear transformation of singular values, with nuclear-norm restoration, which preserves the original total spectral mass. It requires no calibration data or additional training and introduces no inference overhead. Across two backbones and three adaptation tasks, LoRA-Norm improves average specialization and capability retention, outperforming the evaluated post-hoc spectral pruning and gradient-guided editing configurations on both measures. Stronger functional equalization brings no consistent additional gains, revealing that balancing adapter gains and equalizing their responses are distinct objectives.
cs.LG / 162 / 2610.02098
Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
Abstract
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
cs.LG / 163 / 2610.02126
Local Support Learning
Abstract
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.
cs.LG / 164 / 2610.02131
Linear Programming Representations and Strongly Polynomial Algorithms for Robust Markov Decision Processes
Abstract
We study linear programming (LP) representations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, we construct a single LP whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. At fixed discount, the LP has polynomial dimension and encoding length and can be constructed in strongly polynomial time. We also develop a general complexity analysis of robust policy iteration that combines the cost of minimizing over uncertainty sets with the number of iterations needed to evaluate a policy. For a fixed discount factor, we use this analysis to improve the known complexity bounds for $\ell_1$ and $\ell_\infty$ RMDPs and establish new strongly polynomial bounds for general interval, weighted $\ell_1$, and Wasserstein RMDPs, as well as turn-based stochastic games with these uncertainty sets.
cs.LG / 165 / 2610.02140
Finetuning with Sampling: SFT Learns Better Than You Think
Abstract
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
cs.LG / 166 / 2610.02144
Faynt: Scaling and Optimizing Policies for Competitive Melee
Abstract
We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. Post-training combines rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches. On the initial 152-game benchmark, the supervised 10M wins 69.7% of games, compared with 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies. After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. We open-source the weights, both benchmark suites, and a platform for automated model tournaments.
cs.LG / 167 / 2610.02158
Muon meets Tamed Langevin: Momentum Preconditioning beyond Convex and gradient-Lipschitz Potentials
Abstract
We consider the problem of sampling from Gibbs distributions on matrix spaces whose potential energies are neither convex nor globally gradient-Lipschitz. We introduce a family of non-quadratic kinetic energies that lead to a new underdamped Langevin system with momentum preconditioning, in which the gradient of the kinetic energy acts as a smooth spectral taming of the momentum. We prove that, under these relaxed assumptions on the potential, the resulting dynamics leaves the target Gibbs measure invariant, and we establish exponential convergence to equilibrium in a weighted total variation distance. Finally, we show that the corresponding Euler-Maruyama discretization admits moment bounds that are uniform in time, without any modification of the potential gradient, which ensures the stability of the resulting sampling algorithm.
cs.LG / 168 / 2610.02159
When Do Intrinsic Rewards Lead to Exploration?
Abstract
Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.
cs.LG / 169 / 2610.02173
Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
Abstract
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.
cs.LG / 170 / 2610.02175
Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes
Abstract
Protein function annotation needs to know which predictions to distrust, not only what a model predicts. We ask whether tissue-specific interaction structure carries that information. Our candidate signal is effective resistance, used previously to relieve over-squashing by rewiring. Across 24 tissue-specific interactomes it is dominated by inverse degree, and the degeneration deepens as the co-expression filtered network grows, with a Spearman correlation of -0.955. The residual departure from that limit exceeds degree-preserving null graphs in all 24 networks. Controlling for predictive entropy, degree, annotation cardinality, local structure and feature-only difficulty, the residual explains additional per-node loss in 19 of 24 held-out networks once a permutation floor is subtracted, at every depth, and the effect strengthens monotonically with depth. The increment reaches 0.37% of the variance the controls leave unexplained, 5.6 times a permutation floor, against 1.5 times when the model is retrained in a degree-preserving null world. Selective prediction improves negligibly. The signal is reproducible; degree degeneration bounds it.
cs.LG / 171 / 2610.02179
From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
Abstract
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
cs.LG / 172 / 2610.02185
Decoding Looped Transformers Better for (Almost) Free
Abstract
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.
cs.LG / 173 / 2610.02186
Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry
Abstract
Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.
cs.LG / 174 / 2610.02189
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Abstract
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.
cs.LG / 175 / 2610.02190
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Abstract
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.
cs.LG / 176 / 2610.02195
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
Abstract
The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous-time Markov chain, with a temporal-difference penalty that restores the cost. A state cost folds into the reference process as a Feynman-Kac tilt. The cost-augmented bridge is then a plain bridge against the tilted reference, and the penalty is unnecessary. The bridge is computed exactly by alternating two endpoint rescalings, each one sparse matrix-exponential application; nothing is discretized in time or learned. The alternation converges at a rate set by the endpoint coupling alone. For a quadratic congestion cost on time-averaged occupancies, damped best response around the exact bridge is gradient descent on a strongly convex function, and its residual bounds its error. On a protein-folding model, a free-energy cost lowers the expected barrier of the folding paths. On the learned approach's road network, roll-outs of the exact bridge match the target within sampling error, and on networks with millions of intersections its memory grows linearly.
cs.LG / 177 / 2610.02198
FERPO: Forward Entropy-Regularized Policy Optimization
Abstract
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
cs.LG / 178 / 2610.00524
Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs
Abstract
Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: https://taegeunyang.github.io/craft/
cs.LG / 179 / 2610.00727
CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization
Abstract
Controlling an agent with vision requires being able to separate useful information from irrelevant background information. JEPA-style latent world models seem like a natural approach for this, as they do not perform pixel-level reconstruction; however, they are still sensitive to these distractor signals and experience latent collapse. In this work, we introduce Controllability Factorized JEPA (CF-JEPA), a JEPA-style world model which splits the latent space into controllable and uncontrollable subspaces. This factorization allows us to capture all the distractor information into the uncontrollable region, while we use the control-relevant latent information for our task. With this, we show comparable performance across 2D and 3D control tasks under nominal conditions and improved performance under distracted conditions, where CF-JEPA is the only model that does not experience latent collapse. We also validate our model under distracted conditions for a simulated robot task, highlighting the practical application of such a scheme.
cs.LG / 180 / 2610.00926
A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform
Abstract
Autonomous driving is a cornerstone technology for the future of intelligent transportation, where end-to-end learning has emerged as a transformative paradigm that directly maps multimodal sensory inputs to driving actions through unified differentiable models. While offering advantages, the effectiveness of end-to-end autonomous driving (E2E-AD) is ultimately determined by the quality of its training ecosystem. This paper provides a comprehensive review of training methods and ecosystem for E2E-AD. We introduce a Data-Strategy-Platform taxonomy that conceptualizes training as an interdependent system. The data layer defines what can be learned, the strategy layer governs how learning aligns with driving objectives, and the platform layer supports scalability and continuous evolution. Within this framework, we survey recent advances across data-centric pipelines, learning paradigms, and training infrastructures, and analyze their interplay in shaping model performance, robustness, and deployability. Finally, we reflect on current limitations and articulate a forward-looking vision that emphasizes a shift from data quantity to data value, from isolated optimization to foundation-driven generalization, and from static training to integrated training-testing loops, aiming toward robust, scalable, and trustworthy autonomous driving systems. We maintain a continuously updated repository tracking cutting-edge literature and works at \href{https://github.com/Jiaaqiliu/Awesome-Training-Ecosystem-for-E2E-AD}{Our Project Page}.
cs.LG / 181 / 2610.01102
MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending
Abstract
Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.
cs.LG / 182 / 2610.01301
Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations
Abstract
Most current grasp synthesis systems are trained offline and remain fixed during deployment. While this works well when deployment conditions resemble the training data, performance can degrade when robots encounter conditions they have not seen before, such as unfamiliar objects. In this work, we present a continual-learning framework for single-view 6-DoF grasp synthesis for a parallel-jaw gripper in cluttered scenes. Rather than finetuning a large parametric model, our method adapts through memory in a learned embedding space: grasp outcomes update future grasp scores, while optional user demonstrations are recalled and transferred to new scenes as additional candidate grasps. We evaluate our method in simulation and in extensive real-world experiments comprising over 1500 grasp trials. We show that our method matches the performance of existing 6-DoF grasping baselines even before adaptation, improves online on unseen objects from categories absent or underrepresented during training, and supports long-horizon continual learning with limited forgetting. In real-world experiments, our method reaches over 90\% success rates on several challenging object categories after only 50 online grasp attempts. Videos and code at https://giuschio.github.io/cl_grasping/.
cs.LG / 183 / 2610.01746
End-to-End Learning vs. Modular Architectures: Comparative Insights into Autonomous Driving Systems
Abstract
Autonomous driving systems have become a central focus of intelligent transportation research, with End-to-End Learning and Modular Architectures offering two prominent design paradigms for their implementation. E2E Learning uses deep learning algorithms to map raw sensory inputs directly to driving actuators, providing a streamlined and adaptable solution. while Modular Architectures employ a pipeline-based approach, dividing the system into distinct subsystems for perception, cognition, planning, and control. This paper presents a comprehensive comparative analysis of these paradigms, focusing on their strengths, limitations, and trade-offs to provide insights into their suitability for various autonomous driving applications. The study evaluates key factors such as interpretability, scalability, robustness, and real-world applicability. While End-to-End Learning emphasizes simplicity and adaptability in dynamic environments, it lacks transparency and is highly dependent on large datasets. Conversely, Modular Architectures offer superior interpretability and task-specific optimization, but face challenges related to integration complexity and scalability. To address these limitations, hybrid approaches that combine the strengths of both paradigms have emerged, offering a promising direction for overcoming these challenges. Beyond this comparative synthesis, following work proposes a Four-Dimensional Architecture Selection Framework, comprising twelve binary criteria across safety, operating environment, data/computational resources, and deployment context, and validate it against ten published autonomous driving systems, correctly recommending 7/10 deployed architectures. This work synthesizes existing literature to highlight key trade-offs between the paradigms and identifies hybrid architectures as a promising direction for future research.
cs.LG / 184 / 2610.00649
On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection
Abstract
Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited training data. We compare a Quantum Support Vector Machine (QSVM), a classical support vector machine (SVM), and a multilayer perceptron (MLP), all trained on frozen wav2vec 2.0 embeddings using a strict budget of 200 training samples. To match the qubit budget of near-term quantum hardware, embeddings are reduced to four dimensions using principal component analysis, and all models use the same reduced features. Experiments on ASVspoof 2019, ASVspoof 5, the ADD 2023 Challenge, and the In-the-Wild dataset show that under severe domain shift from ASVspoof 2019 to ADD 2023, the MLP degrades to near-random performance, with an area under the curve of approximately 50% and an equal error rate of 50.0%. In contrast, the QSVM maintains meaningful discrimination, achieving an area under the curve of 76.0% and an equal error rate of 27.0%. This advantage is not consistent across transfer directions. When trained on ADD 2023, the QSVM falls below chance on two of three transfers, while the MLP performs better. These results suggest that quantum kernel methods can be competitive under severe cross-corpus shifts and strict low-resource constraints, but do not provide a consistent advantage under near-domain transfer. We interpret these findings as an empirical characterization of quantum kernel inductive bias under distribution shift, rather than evidence of quantum advantage, since the four-qubit kernel can be simulated exactly on classical hardware.
cs.LG / 185 / 2610.00706
AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models
Abstract
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
cs.LG / 186 / 2610.01926
LAST: Looped Audio Spectrogram Transformer
Abstract
Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
cs.LG / 187 / 2610.01361
Degree-Corrected Joint Matrix Factorization for Multilayer Community Detection
Abstract
Multilayer networks allow the modeling of interactions between the same entities across different contexts, such as temporal observations, varying settings, or interactions of different types. The goal of community detection in multilayer networks is to identify groups of nodes exhibiting similar connectivity patterns, which may vary across layers. We propose a method based on a joint nonnegative symmetric matrix trifactorization for community detection in multilayer networks, where each graph is approximated by a nonnegative symmetric matrix trifactorization. Our approach enforces constraints on the factor matrices so that communities are disjoint and shared across layers, while allowing each layer to have its own connectivity patterns and node degrees. This flexibility enables the model to capture both local and global structural variations across layers. We also develop an algorithm to efficiently solve this problem. We evaluate multilayer community detection methods using the multilayer degree-corrected stochastic block model (MDCBM), a flexible framework for generating realistic multilayer graphs with heterogeneous degrees and varying connectivity patterns. Experiments show that our method reliably detects communities across diverse regimes, whereas existing state-of-the-art approaches are often limited by restrictive structural assumptions.
cs.LG / 188 / 2610.00607
End-to-End Historical Music Restoration in Latent Space
Abstract
Historical music restoration (HMR) has almost exclusively focused on constrained problems such as Super-Resolution or the restoration of solo pieces, under-exploring the general task of restoring orchestral historical music, which has multiple instruments. This under-exploration is largely because the HMR domain, early-20th-century recordings, has no pre-degradation ground-truth pairs, making the restoration task unsupervised and more challenging. This paper presents a supervised end-to-end orchestral HMR benchmark by exploring both the synthetic degradation functions and the end-to-end generative deep-learning restoration methods. We simulate the historical recording degradation chain more faithfully than prior work, which makes orchestral restoration into a tractable supervised problem. A latent flow-matching model trained on the resulting synthetic pairs outperforms existing HMR baselines on intrusive, non-intrusive, and subjective evaluations. We also curate and release a 9.3-hour license-free, unpaired, historical classical-music test set, along with code and audio demos.
cs.LG / 189 / 2610.00678
Nonparametric Distribution Matching for Self-Supervised Whole-Slide Image Condensation
Abstract
Histological whole-slide images (WSIs) are central to computational pathology but pose severe computational challenges due to their extremely high resolution, often spanning several gigabytes per slide. To enable scalable learning, existing methods apply self-supervised data condensation to reduce computational cost, but typically rely on heuristic prototype learning and do not explicitly preserve learning-relevant feature distributions for downstream tasks. In response, we introduce a principled reformulation of WSI condensation as a distribution-matching problem under a fixed representational lens, and develop NICER, a tractable approximation framework based on a nonparametric prior with slide-adaptive capacity. Experiments on five histopathology datasets, together with clinical evaluation from a board-certified pathologist, show that NICER consistently outperforms prior methods, achieving an average accuracy improvement of 7.44% while offering improved efficiency-accuracy trade-offs, highlighting the benefits of principled, distribution-aware condensation for scalable histological representation learning. Source codes are available in https://github.com/nmduonggg/NICER.
cs.LG / 190 / 2610.00596
Mean Spatial Frequency Decoupling for Learning-Based Uplink-to-Downlink Covariance Conversion in FDD Massive MIMO
Abstract
In frequency division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems, the uplink (UL)-to-downlink (DL) channel covariance matrix (CCM) conversion problem is studied to relieve the heavy burden of DL training and feedback required for channel estimation. Learning- based methods perform well up to a certain array size, but for a fixed dataset size their accuracy deteriorates with the number of antennas, to the point where simple model-based methods outperform them. This paper identifies a key cause of this behavior and addresses it. The mean angle of arrival (AoA) induces a phase ramp along the lags of the CCM. Since the oscillation rate of this ramp grows with the number of antennas, a dataset of fixed size becomes increasingly sparse relative to the variation that must be captured. We propose estimating the slope of this ramp from the UL CCM separately and mapping it to the DL band in closed form, leaving the learner with a residual that is largely insensitive to the mean AoA, which substantially reduces the performance degradation with an increasing number of antennas. The proposed scheme, termed deramping, is a combination of pre- and post-processing steps that applies to learning-based conversion methods without altering their internal structure, as demonstrated on three structurally different learners. Simulation results show that deramping reduces the covariance estimation error of all three learners under uniform, Laplacian, and Gaussian angular power spectra,keeps the interpolation-based learners ahead of a model-based benchmark at large array sizes, and improves downlink channel estimation.
cs.LG / 191 / 2610.00819
PI-AMFM: Permutation-Invariant Learning for Variable-Cardinality AM-FM Mode Decomposition in Biomedical Signal Analysis
Abstract
Physiological recordings often contain nonstationary oscillatory components whose number and dynamics vary across signals. Amplitude- and frequency-modulated (AM-FM) representations are well suited to characterizing such dynamics and have shown broad utility in biomedical signal analysis. Recent approaches have incorporated neural networks to learn mode decomposition patterns from data, but component cardinality is often predefined or determined through separate stopping or selection mechanisms. We propose a permutation-invariant neural framework for variable-cardinality AM-FM mode decomposition (PI-AMFM). PI-AMFM combines a multiscale temporal encoder, Mamba backbone, and component-presence estimation, with permutation-invariant Hungarian matching during training. On synthetic AM-FM signals, PI-AMFM achieved lower decomposition, instantaneous-frequency, reconstruction, and mode-count errors than the compared methods while preserving the overall trajectory pattern in a crossing-chirp example. On photoplethysmographic recordings, recovered modes captured cardiac and respiratory dynamics despite training only on synthetic signals. These results support the feasibility of PI-AMFM for variable-cardinality decomposition of nonstationary biomedical signals.
cs.LG / 192 / 2610.00569
Scaling Collider Event Generation with Residual-Quantized Tokens
Abstract
Full detector simulation and reconstruction of collider events are projected to become major bottlenecks at the High-Luminosity Large Hadron Collider, motivating the development of fast, ML-based surrogates. At the same time, LLMs have driven fast progress in generative discrete modeling: autoregressive transformers trained on tokenized data now represent the state of the art across a range of generative tasks. We extend the discrete modeling paradigm by introducing a particle-level generative model trained on residual-quantized full-event data. We demonstrate the ability of this model family to perform conditional generation from detector-stable particles; we study its scaling behavior across a range of dataset and model sizes, characterize the effects of repeated data exposure and demonstrate that token-level loss systematically predicts downstream physical fidelity. These results provide an empirical framework for scalable collider full-event generation based on residual-quantized representations.
cs.LG / 193 / 2610.02084
Kolmogorov-Arnold Networks for Free-Boundary Partial Differential Equations
Abstract
We study free-boundary problems within a physics-informed framework using Kolmogorov-Arnold network (KAN) approximations. The proposed approach incorporates obstacle constraints, partial differential equation (PDE) inequalities, complementarity conditions, and boundary conditions through residual-based loss functions. We consider a linear elliptic obstacle problem, a nonlinear $p$-Laplacian obstacle problem, and a time-dependent one-phase Stefan problem. The proposed KAN solver is compared with physics-informed neural network (PINN) and residual-network baselines. Numerical experiments show that KANs achieve low relative $L^2$ and $L^\infty$ errors while accurately resolving contact regions and moving interfaces. The results indicate that KAN representations provide an effective alternative for solving free-boundary PDEs.
cs.LG / 194 / 2610.01015
Initial condition recovery in nonlinear damped viscous photoacoustic tomography using a convolutional neural network-guided gradient-free optimization framework
Abstract
Photoacoustic tomography (PAT) is a hybrid imaging modality that combines high optical contrast with high ultrasonic resolution for biomedical imaging applications. In this work, we investigate the inverse problem of recovering the initial pressure distribution from boundary measurements in the presence of nonlinear acoustic propagation and viscous attenuation effects. To model these phenomena more accurately, we consider a nonlinear damped viscoelastic wave equation incorporating spatially varying sound speed, temporal attenuation, and nonlinear propagation mechanisms. We first establish the well-posedness of the corresponding forward problem using a Galerkin approximation combined with energy estimates and a fixed-point argument. For the inverse problem, we derive existence, uniqueness, and local uniqueness results under suitable assumptions through a harmonic extension reduction, spectral Laplace transform techniques, and observability estimates. To numerically reconstruct the initial pressure field, we develop a hybrid reconstruction framework that combines a convolutional neural network (CNN) with a gradient-free optimization strategy based on the sequential quadratic Hamiltonian (SQH) method derived from Pontryagin's maximum principle. The CNN is used to generate an informative initial guess, while the SQH framework enforces the governing PDE dynamics during the reconstruction process. Numerical experiments demonstrate that the proposed hybrid strategy significantly improves reconstruction quality, contrast, and robustness compared to standalone time-reversal and CNN-based approaches.
cs.LG / 195 / 2610.01344
Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers
Abstract
Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
cs.LG / 196 / 2610.01572
Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization
Abstract
This paper investigates stochastic multi-level optimization where the objective is a nested composition of several smooth non-convex functions. We assume that only stochastic estimates of the gradient and function values for each level are accessible. Consequently, obtaining an accurate estimate of the overall gradient is challenging due to the nested structure. To address this, we employ a momentum-based estimator with mini-batches to track the function values of each level, which are subsequently used to construct momentum gradient estimators. We establish an optimal sample complexity of $\mathcal{O}(ε^{-4})$ for finding an $ε$-stationary point, avoiding the stronger average smoothness assumption commonly relied upon in prior literature. Furthermore, by employing a normalization technique, we attain the same rate without requiring problem-dependent constants to set hyperparameters. To achieve the optimal rate without mini-batches, we further develop a batch-free method that incorporates a first-order approximation and a clipping technique for function value estimation. Finally, we validate the effectiveness of our proposed methods through experiments on risk-averse portfolio optimization and hierarchical tilted empirical risk minimization.
cs.LG / 197 / 2610.01599
Convergence Analysis of STORM Under Different Geometries
Abstract
Stochastic recursive momentum (STORM) achieves fast convergence for nonconvex optimization via the variance reduction effect, but existing analyses rely on the strong average smoothness assumption. In this paper, we study the convergence of STORM for different objectives without average smoothness. We first revisit the results under average smoothness, obtaining the $O(T^{-1/3})$ bound for nonconvex objectives and the $O(σ^2/(μT))$ bound for last-iterate output under the $μ$-Polyak--Łojasiewicz~(PL) condition. Without average smoothness, we design an auxiliary sequence and compare the STORM update with it in the analysis. With the help of this sequence, we prove that STORM still attains an $O(T^{-1/4})$ rate for nonconvex objectives, which is optimal under standard smoothness. For convex and $λ$-strongly convex objectives, we further prove averaged and last-iterate bounds with optimal rates of $O(σR/\sqrt T)$ and $O(σ^2/(λT))$, respectively. All the obtained results use the same STORM recursion with different hyperparameter choices.
cs.LG / 198 / 2610.01662
Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex--Concave Minimax Optimization
Abstract
We establish complexity lower bounds for stochastic first-order algorithms in nonconvex--concave minimax optimization, allowing algorithms to use variance reduction. Our main contribution is a lower bound for a zero-respecting algorithm class that permits variance reduction, extending beyond the algorithmic restrictions imposed by some existing lower bounds. We consider objectives with an $L$-Lipschitz continuous joint gradient, a compact convex dual domain of Euclidean radius at most $D_Y$, and a primal value function, defined by maximizing the objective over the dual variable, with initial suboptimality at most $Δ$. The target accuracy $\varepsilon$ is measured by the gradient norm of the Moreau envelope of the constrained primal value function with parameter $1/(2L)$. Under an unbiased stochastic first-order oracle with variance at most $σ^2$ and mean-square smoothness, we prove the lower bound $Ω\!\left(L^2D_YΔ\varepsilon^{-3}+L^3D_Y^2Δσ^2\varepsilon^{-6}\right)$. This result quantifies the dependence on accuracy, dual-domain radius, and oracle noise even when variance reduction is allowed. We also establish complementary lower bounds for nonconvex--strongly-concave minimax optimization. With dual strong-concavity parameter $μ>0$ and condition number $κ:=L/μ$, we obtain $Ω\!\left(LΔ\sqrtκ\,\varepsilon^{-2}+LΔκσ^2\varepsilon^{-4}\right)$ under the bounded-variance oracle model. Under the additional mean-square smoothness condition with constant $\bar L$, we obtain $Ω\!\left(LΔ\sqrtκ\,\varepsilon^{-2}+Δ\bar Lσκ^{3/2}\varepsilon^{-3}\right)$. Together, these results identify complexity barriers across the concave and strongly concave regimes, with the main nonconvex--concave bound remaining valid for algorithms that use variance reduction.
cs.LG / 199 / 2610.01843
Optimal Stochastic Bilevel Optimization with First-Order Oracles
Abstract
We study nonconvex--strongly-convex bilevel optimization under a stochastic first-order oracle. We introduce MRT-FD, a single-loop first-order method that simultaneously tracks the upper-level variable, the lower-level solution, and the auxiliary response arising from implicit differentiation of the hyperobjective. MRT-FD performs one update of each variable per iteration and approximates the second-order derivative actions using order-$p$ finite differences. For any fixed finite smoothness order $p\ge1$ in the lower-level variable, MRT-FD finds an $\varepsilon$-stationary point using $\mathcal{O}(\varepsilon^{-4-2/p})$ stochastic gradient queries. We also prove a matching $Ω(\varepsilon^{-4-2/p})$ oracle lower bound. The lower-bound construction starts from a hard nonconvex minimization chain with a stronger stochastic oracle, and lifts it to a bilevel problem through a sinusoidal coupling with a scalar lower-level variable. Consequently, the dependence on $\varepsilon$ is optimal for every fixed finite $p$, closing the upper--lower complexity gap in this stochastic first-order oracle setting.
cs.LG / 200 / 2610.01835
Varda-single-1.0: deterministic data-driven weather forecasting at 1 km resolution over Switzerland's complex topography
Abstract
We present Varda-single-1.0, a medium-range data-driven weather prediction system built for the Alpine domain. It provides hourly deterministic regional forecasts on a mesh of 1 km resolution and global forecasts on a 31 km mesh. The system comprises two independently trained stretched-grid Graph Transformer models with encoder-processor-decoder architecture, developed in the Anemoi framework: a 6-hourly autoregressive forecaster and a temporal downscaler reconstructing hourly forecasts between the forecaster's steps. Its training curriculum includes pre-training on ERA5 reanalysis data, followed by training on a 20-year kilometre-scale regional reanalysis, and finally fine-tuning on operational kilometre-scale analyses. Verified over one year against operational analyses and surface station observations, Varda-single is competitive with or improves on MeteoSwiss' operational numerical weather prediction baselines for most headline scores and variables. It broadly matches the skill of the high-resolution 1 km ICON-CH1-EPS control at lead times up to +33 h and generally outperforms the 2 km ICON-CH2-EPS control at lead times up to +120 h. Despite competitive aggregate scores, Varda-single underestimates some local wind maxima and produces overly smooth convective precipitation fields, consistent with the smoothing associated with squared-error training. To gain insight into the model's behaviour, we investigate three case studies beyond the aggregated headline scores, and find particular weaknesses in Varda-single's representation of local winds over complex terrain. Varda-single represents an important step in the development of high-resolution ML forecasting over complex terrain, in complementing the operational regional numerical weather prediction models of MeteoSwiss with data-driven models and in providing a pretrained model for researchers and user-specific applications.
cs.LG / 201 / 2610.02069
AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure
Abstract
Rare weather regime transitions pose a challenge for data-driven modeling due to class imbalance. In this study, we develop a probabilistic deep learning emulator for a prototypical system with regime transitions, the stochastic Holton--Mass model of stratospheric variability, and analyze the structure of its learned latent space. The Holton--Mass model exhibits two metastable regimes, a strong and a weak polar vortex, maintained by nonlinear wave--mean flow interactions, with weak stochastic forcing intermittently triggering rare transitions between these regimes that qualitatively represent SSW events. We employ a ResNet-inspired Conditional Variational Autoencoder with six-layer encoder and decoder layers and explicit current-state conditioning to model the distribution of the system's state at the next time step (one day). The emulator accurately reproduces short-term dynamics, steady-state probability distributions, regime persistence statistics, rare transition rates, the transition committor function, and the transition expected lead time of the physical model. Beyond emulation fidelity, we interrogate the learned latent representation to understand how the model internalizes the underlying metastable structure of the dynamics. Principal Component Analysis of the 32-dimensional latent space reveals a clear and unsupervised separation into four physically interpretable clusters corresponding to strong versus weak vortex regimes and stable versus transition-prone configurations. Such emergent regime separation in latent space is hard to identify for deep generative models applied to high-dimensional stochastic systems. Our results show that carefully designed probabilistic emulators can uncover physically meaningful manifolds governing extreme-event dynamics, potentially aiding the development of improved operational advanced warning systems.
cs.LG / 202 / 2610.01381
SupraTITO: Transferable Generative Molecular Dynamics for Supramolecular Systems
Abstract
Peptide sequence governs both the structures formed through supramolecular assembly and the dynamics by which they emerge, but predicting either requires resolving slow collective processes among many interacting molecules. Molecular dynamics (MD) provides microscopic insight into these processes, yet the long timescales of assembly and the vast peptide sequence space make systematic exploration computationally demanding. We introduce SupraTITO, a transferable generative molecular dynamics (GenMD) framework for supramolecular systems, demonstrated through peptide self-assembly. SupraTITO learns transferable implicit transfer operators (TITO) conditioned on peptide sequence, molecular topology, and periodic geometry, allowing configurations to be propagated over physical intervals much longer than an MD integration step. On a comprehensive dipeptide benchmark, SupraTITO generalizes to held-out sequences and reproduces sequence-dependent structures and dynamics while maintaining molecular integrity over long rollouts. Compared with direct ensemble prediction trained on the same trajectory data, SupraTITO more accurately reproduces assembly structures while also resolving their temporal evolution. The learned dynamics generalize across peptide concentrations, including dilute conditions not represented during training. These results extend transferable GenMD to collective dynamics in periodic supramolecular systems and provide a foundation for modeling related processes beyond peptide assembly.
cs.LG / 203 / 2610.00742
StabilityArc: Decoding Protein Sequence Embeddings into Generalizable Stability Landscapes
Abstract
Every protein has a unique stability landscape, but the physical consequences of mutation are governed by recurring biochemical constraints. We test whether a shared decoder, trained on measurements from diverse proteins, can interpret these constraints in an unseen target, enabling cross-protein transfer for initial experimental round prescreening. We present StabilityArc , which maps frozen ESMC-600M residue representations through a shared RoPE transformer to an Lx20 matrix of substitution effects; a symmetric, contact-aware residual aids in predicting epistasis in simultaneous substitutions. In 66 strict leave-one-protein-out evaluations covering 134,794 ProteinGym variants, StabilityArc achieves 0.7134 Spearman correlation, exceeding the strongest zero-shot baseline, ProSST-2048 (0.6526), by 0.0608. We further explore the utility of this method by providing the score as a prior for Kermut, achieving Spearman correlation of 0.8280 across three supervised split schemes, improving on Kermut's reported 0.8167.
cs.LG / 204 / 2610.00841
Neural Fourier Surrogates for Data Reuploading Quantum Neural Networks
Abstract
For quantum machine learning, the exact boundary between classical and quantum advantage is still poorly understood. Direct comparison between quantum neural networks (QNNs) and existing classical models, which encompass fundamentally different function classes, often fails to provide broader insight into the difference between the two. Inspired by the techniques of Neural Quantum States and Random Fourier Features, this work introduces Neural Fourier Surrogates (NFS), a stochastic classical neural network architecture for efficiently learning coefficients over the same finite Fourier series support as quantum neural networks. Testing on a selection of tabular benchmark datasets, we find that NFS is an effective classifier architecture broadly competitive with established classical baselines, including a comparable Random Fourier Features model, and possessing comparable performance to data-reuploading QNNs; combined with additional analysis comparing the learned Fourier spectra of QNNs and NFS on synthetic data, these results establish NFS as a natural classical baseline for evaluating QNN performance.
cs.LG / 205 / 2610.01141
Classical Hardness of Learning Functions of Hamiltonians
Abstract
Morohoshi, Nakayama, Manabe, and Mitarai proposed a physically motivated quantum machine learning problem in which the goal is to predict quantities of the form $\operatorname{Tr}[f(H)ρ]$ from classical descriptions of a Hamiltonian $H$ and a quantum state $ρ$, where $f$ is an unknown function. We call this problem Hamiltonian function learning in this paper. They constructed an efficient quantum learning algorithm under suitable conditions, while leaving a rigorous proof of average-case classical hardness open. In this paper, we rigorously prove the average-case classical hardness for two distribution-specific Hamiltonian function learning problems for $f_{\cos,π}(λ)=\cos(πλ)$ and $f_{\exp,β}(λ)=e^{-βλ}$ discussed in the paper of Morohoshi et al. under the assumption of the average-case hardness of factoring random RSA moduli. More specifically, we show that an efficient classical randomized learner under squared loss whose output hypotheses are evaluable in classical polynomial time for either problem would yield a classical randomized polynomial-time algorithm for factoring random RSA moduli.
cs.LG / 206 / 2610.02068
Sequential Capacity of Quantum Processes with Finite Memory
Abstract
How complex can the responses of a quantum device become as it runs longer with a fixed internal memory? We quantify this complexity through sequential response capacity: how many adaptive testing stages, each using a fresh run, can continue to separate possible processes by a prescribed gap in response probabilities. For fixed system and memory sizes, we establish a tight law relating this capacity to run length and probability resolution. At fixed resolution, the capacity grows on the order of $K\log K$, where $K$ is the number of time steps in each run. Our construction attains this growth using time-dependent phase rotations on a single visible qubit with no additional internal memory; its tests give response probabilities exactly zero or one. Under the same tests, classical stochastic processes that measure in a fixed basis at every step have only linear capacity at fixed sizes and resolution. For phase sequences selected by a stored classical label, we then quantify how known independent Pauli noise changes this logarithmic enhancement. With ideal controls and weak residual phase noise after correction, we prove matching capacity bounds at a fixed small probability gap. These bounds identify the inverse residual phase-flip probability as the coherence timescale that limits the extra logarithmic growth.
cs.LG / 207 / 2610.00498
Heteroskedastic Canonical Polyadic Tensor Decomposition
Abstract
When minimizing the squared-error loss, the popular CP decomposition can be interpreted as parameter inference in a Gaussian model with a low-rank mean tensor and constant variance across the tensor entries. We introduce heteroskedastic-CP (HCP), which models entrywise variability with a non-constant, low-rank precision tensor, and develop an alternating block-coordinate ascent method to recover both the low-rank mean and precision tensors from noisy observations. Our procedure is computationally competitive, with the same leading-order factor-update complexity as CP-ALS. We demonstrate HCP on synthetic experiments and an EEG application.
cs.LG / 208 / 2610.01088
Polylogarithmic Sparsity of Randomly Reweighted NPMLEs for Gaussian Mixtures
Abstract
The nonparametric maximum likelihood estimator (NPMLE) of a Gaussian location mixture maximizes the likelihood over the infinite-dimensional space of mixing distributions. The maximizing mixing distribution can be nonunique, and the classical bound on its number of atoms grows linearly with the sample size $n$. We show that a vanishingly small random perturbation of the likelihood yields exact polylogarithmic sparsity. The resulting randomly reweighted NPMLE maximizes a weighted likelihood whose independent weights, taken to be Gamma in our analysis, concentrate around one as $n$ grows. With high probability, it is unique, has $O\{(\log n/\log\log n)^d+\log n\}$ atoms in dimension $d$, nearly maximizes the ordinary likelihood, and estimates the mixture density at a Hellinger rate that is parametric up to logarithmic factors. This sparsity holds for the estimator itself, not for an approximation of it, and requires no support penalty. The proof rests on an effective-dimension principle for positive kernel mixtures: low-dimensional variation of the fitted values controls the support of every extreme point of the set of maximizers. Numerical illustrations verify that the reweighted NPMLE has Hellinger risk and support size comparable to those of the ordinary NPMLE.
cs.LG / 209 / 2610.01823
Generalized Engression Models
Abstract
We consider estimating the conditional distribution of a multivariate outcome given covariates when its coordinates may be continuous, binary, categorical, ordinal or rankings, and are conditionally dependent on one another. Different statistical methods have been developed for each outcome type, and most of them target a summary of the conditional distribution, such as the mean of each coordinate, rather than the joint distribution of the outcome vector. We develop generalized engression models, a unified nonparametric distributional regression framework for outcomes of any type. The proposed method builds upon engression, a scoring-rule-based deep generative model, and introduces a data-type-specific link function and a stochastic perturbation that smooths the loss, enabling gradient-based training even with discontinuous links. We establish universal representation results for continuous, discrete and mixed outcomes. In simulations and in two applications, 242 species in a community ecology benchmark and a 17-dimensional mixed-type health outcome, the method matches type-specific models on marginal scores, improves on them on the joint distribution, and matches or exceeds purpose-built state-of-the-art joint species distribution models. Software is available in Python.
cs.LG / 210 / 2610.00515
Fractional Laplace Neural Operators: Exact Architectures, an Expressivity Frontier at Criticality, and Certified Stability for Memory-Driven Network Dynamics
Abstract
Neural operators learn maps between function spaces, while hereditary network dynamics are described by Volterra resolvents with non-rational Laplace symbols. We introduce a fractional Laplace neural operator (fLNO) that embeds this structure in the learned map. For commuting excitation--Laplacian pairs, one block graph-spectral layer represents the full linear Volterra solution operator exactly. We establish an expressivity frontier for finite rational realizations: they approximate fractional memory geometrically on compact frequency windows, but cannot reproduce the non-integer critical asymptotics generated by a branch point, and on the half-line the best rational rate is root-exponential. The same theory yields trainable parametrizations that enforce a prescribed stability margin by construction, and a graphon-transfer theorem separates genuine operator consistency from parameter sharing. In a common-data benchmark, positive rational operators can match or exceed fLNO accuracy on finite horizons, whereas in controlled near-critical experiments fLNO recovers the branching coordinate more faithfully with far fewer parameters; unconstrained rational fits can cross the stability boundary, while certified parametrizations cannot. A four-parameter spectral law transfers without retraining from graphs of size 48 to 192 with 0.51--0.62% relative error. Applications to Chilean aftershock sequences and to renewal models for Chile and 21 Italian regions illustrate structured inference with explicit uncertainty. The contribution is an operator-learning architecture in which exact memory structure, physical coordinates and stability guarantees coexist with competitive accuracy.
cs.LG / 211 / 2610.00755
Learning to Price Electricity for Optimal Demand Response
Abstract
There is considerable interest in using time-varying electricity prices to shape consumer demand response, and better align energy demand with renewable production. However, optimal prices generally vary over time in response to complex signals such as weather forecasts, sunrise/sunset times, and day-of-week patterns; and existing methods are not able to make efficient use of such rich contextual information. Here, we propose a neural-network-based algorithm for contextual energy pricing, modeling pricing as a Stackelberg game and leveraging a mean-field solution representation from Mehrabi et al.~(2024). The approach learns constrained mappings from contextual features to feasible price signals. We validate our approach by simulating the energy grid in several US cities, and show that incorporating contextual information can considerably increase the value of the demand response programs.
cs.LG / 212 / 2610.00911
Block Optimism for Nonstationary Bandits with Latent Linear Dynamics
Abstract
We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach achieves $\tilde{O}(T^{2/3})$ regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order $\tilde{O}(\sqrt T)$, significantly improving over the previous $\tilde{O}(T^{2/3})$ guarantee for the same model. To the best of our knowledge, this is the first $\tilde{O}(\sqrt T)$ regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.
cs.LG / 213 / 2610.00993
The Price of Correlated Tests: How Strict Should a Model Release Gate Be?
Abstract
Before a machine learning model ships, it often has to pass a suite of automated tests. Requiring every test to pass looks safe, yet it can reject many models that would have served users well, and it does not say how trustworthy a passing model actually is. We treat the release gate as a design problem: choose how many tests a model must pass so that cleared models meet a stated reliability target, while keeping as many good models as possible. A two-class latent-factor model makes both costs explicit and reduces each calculation to a one-dimensional integral. We prove that when both classes share the same latent correlation, a stricter gate always raises reliability, so the gate that keeps the most good models is the most lenient one that still meets the target. Under pass-all gating, any reliability target short of perfection is attainable within the model, but the share of good models kept tends to zero as the suite grows. Correlation between tests sets the price. In one configuration, a 99 percent target needs 8 independent tests, but 74 tests at a latent correlation of 0.3 and 5,182 at 0.5, where the gate keeps fewer than one good model in ten. We also give a validation procedure, built on exact binomial bounds, that certifies a gate from labelled data even when the gate is chosen from a fixed shortlist.
cs.LG / 214 / 2610.01005
Tolerance-Based Fairness Auditing: Violation Certification and Sensitivity Screening
Abstract
As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairness auditing. In practice, the tolerable degree of algorithmic unfairness depends on the specific legal, ethical, or application context. Given a prespecified tolerance threshold, an important statistical question is how to determine whether a group disparity exceeds the allowable tolerance across different auditing objectives. To address this problem, we develop a unified tolerance-based fairness auditing framework for two complementary auditing objectives: violation certification, which prioritizes control of false violation declarations, and sensitivity screening, which prioritizes reducing missed violations. For the first objective, we develop a constrained empirical likelihood test for formal settings that uses least-favorable-point calibration and can be combined with false flagging rate control for simultaneous subgroup auditing. For the second objective, we develop split empirical likelihood and adjusted split empirical likelihood tests using an adaptive boundary-proxy principle for early-warning settings. Numerical experiments show the distinct error-control--sensitivity trade-offs of these procedures. A COMPAS analysis illustrates the framework in predictive fairness auditing.
cs.LG / 215 / 2610.01472
Zero Flux: Flow-Based Comparison of High-Dimensional Discrete Distributions
Abstract
Comparing two high-dimensional discrete distributions has always been a challenging task due to the exponentially growing state space and complex changes in interactions. A recent work suggests comparing distributions through a vector field trained using flow matching between two continuous distributions. The resulting vector field at mid-point vanishes if and only if two distributions identical. However, such a flow-based criterion does not naturally apply to discrete distributions. We extend this principle to the discrete domain and introduce the \emph{Zero Flux} criterion, a discrepancy based on local probability fluxes. Under independent coupling, we show that all local probability fluxes vanish at the midpoint if and only if two distributions are the same. This discrepancy decomposes the joint distributional difference into smaller, local contributions and can be efficiently estimated from samples. We establish finite sample error bounds for our estimator. Experiments on synthetic and real categorical data demonstrate reliable recovery of sparse dependence signals and stable tracking of distribution shifts in high dimensions.
cs.LG / 216 / 2610.01578
The hidden advantage of mask resampling: a theory of masked autoencoders
Abstract
Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of $K$ masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.
神经与进化计算 (cs.NE)
7
cs.NE / 1 / 2610.00719
Neuromorphic Pseudo-Random Number Generators with a Low Power Hardware Implementation
Abstract
Pseudo-random number generation often requires trade-offs among quality, power consumption, and bandwidth to produce unpredictable sequences of numbers. The brain, on the other hand, efficiently generates unpredictable output complex network dynamics occurring in a high-dimensional state. This state, which is hypothesized to be chaotic, relies on the balance between excitation and inhibition. Here, we investigated if computational models of these chaotic balanced states can be harnessed for Neuromorphic Pseudo-Random Number Generators (NPRNGs) in low power hardware. We successfully constructed a balanced spiking neural network model consisting of leaky-integrate-and-fire neurons that could be readily implemented in low power FPGAs and used as a NPRNG. The prototyped NPRNG consumed 3.24 mW during operation and produced pseudo-random numbers at 120kbps. In both hardware and software instantiations, NPRNGs produce high-quality random numbers as validated by standard metrics for testing RNG quality.
cs.NE / 2 / 2610.01232
Inherited Learning in an Artificial Ecology: How Controls and Update Allocation Shape Benefits
Abstract
Learning can improve an individual's behavior, yet a population risks losing that experience whenever individuals die and are replaced. Inheriting learned preferences offers a way to preserve useful behavior across generations, raising a question for artificial populations: when does inheritance improve collective performance, and how can its benefits be measured fairly? The challenge is that inheritance changes not only offspring behavior but also survival, reproduction, and opportunities for further hereditary updates. Random controls with equal update magnitudes may therefore yield misleading comparisons if they alter different states or obey different stability constraints. We investigate this problem in a resource-limited artificial ecology, combining structured random controls with interventions on newborn preferences and the allocation of hereditary updates. Preserving the state structure of random updates substantially narrows the apparent inheritance advantage, while a conditional establishment-speed benefit remains. Preference erasure and faster-learning compensation support a contribution from reduced offspring relearning. Update allocation also changes the comparison: event quotas and common time cutoffs can reverse rankings, although they also change realized update amounts. With update count and cumulative magnitudes matched, staged release improves occupancy but does not achieve the prespecified establishment criterion. These findings provide a framework for distinguishing the value of inherited preferences from the effects of control design and update allocation, clarifying how inherited learning should be evaluated in artificial populations.
cs.NE / 3 / 2610.01468
LESS: Lightweight Evolutionary Supernet Search in Minutes
Abstract
Low-cost NAS must both explore high-performing architectures and identify them reliably, yet reducing evaluation cost often weakens the fidelity of candidate comparisons. Training-free methods reduce evaluation cost by replacing learned task feedback with proxy signals measured at initialization. We introduce LESS (Lightweight Evolutionary Supernet Search), a data-driven method that combines a brief fair hard-path warm-up with discrete search under a single CMA-ES distribution. Each proposal is evaluated as its decoded hard genotype after six candidate-conditioned supernet updates. On NAS-Bench-201, LESS achieves \(93.189\pm0.467\%\) CIFAR-10 test accuracy in 409.1 seconds, coming within 0.04 percentage points of FairNAS using approximately \(1/24\) of its source-reported search time. Matched controls show that calibration improves selected validation accuracy by \(0.577\) percentage points while changing best-visited accuracy by only \(0.054\) points, indicating that its primary effect is to reduce selection regret. The frozen configuration transfers without tuning to CIFAR-100 and ImageNet16-120 with \(69.615\pm1.139\%\) and \(43.720\pm1.697\%\) accuracy. Applied without tuning to the larger DARTS space, LESS achieves \(96.95\pm0.14\%\) on CIFAR-10 and \(82.43\pm0.80\%\) on CIFAR-100, with each search completing in approximately 43.5 minutes on a single GPU. Together, these results show that short, balanced, data-dependent updates enable competitive neural architecture search across datasets and search spaces within minutes.
cs.NE / 4 / 2610.01558
Controllable Stochastic Quantization Encoding for Adversarially Robust Spiking Neural Networks
Abstract
Spiking Neural Networks (SNNs) have attracted increasing attention due to their impressive temporal dynamics, energy efficiency, and brain-inspired mechanisms. Although SNNs have demonstrated promising performance in image classification tasks, recent studies have shown that they remain vulnerable to adversarial attacks, where imperceptible perturbations are added to input images to mislead model predictions. Existing defense methods mainly focus on training strategies, while the role of input encoding remains less explored. An observation is that the robustness advantage of Poisson encoding over direct encoding may benefit from its inherent randomness. Motivated by this, we propose a stochastic quantization encoding method that encodes the input image with controllable randomness adjusted by the quantization scale, thereby improving the adversarial robustness of SNNs. We further show that this method constitutes a general framework that reduces to both Poisson encoding and direct encoding under different choices of the quantization scale. Since it enhances robustness at the input encoding stage, it can be combined with existing training-based defenses for further gains. Experimental results on CIFAR-10 and CIFAR-100 demonstrate the effectiveness of the proposed stochastic quantization encoding method. To sum up, this work highlights the importance of input encoding for the adversarial robustness of SNNs, providing a new perspective for understanding and improving it.
cs.NE / 5 / 2610.01583
Continual Reinforcement Learning with Neuroevolution
Abstract
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
cs.NE / 6 / 2610.01887
TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design
Abstract
Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators can read, audit, and execute within tight latency budgets. LLM-based Automatic Heuristic Design (AHD) promises to automate writing such rules. However, existing AHD frameworks were developed for combinatorial problems fully specified to the LLM, and they learn only from a scalar fitness score. In real systems, the behaviour that determines a good heuristic, such as processor speeds or power consumption, is unknown a priori: the score reveals which heuristic performs better, but not why. This missing information is recorded in the system logs that every evaluation produces. Exploiting it is non-trivial: logs are massive and noisy, the relevant signals depend on the objective, and their content and format vary across hardware and software stacks, so they can neither be fed to an LLM as is nor processed by a fixed parser. We propose TRACE, which couples an evolutionary AHD loop with an agentic knowledge-extraction workflow. A Reasoner agent analyzes the log schema in light of the objective and formulates hypotheses about the system dynamics; a Coder agent writes and executes schema-specific code to test them, producing insights or executable tools for the evolved heuristics. We evaluate TRACE on a synthetic cloud benchmark and a 5G vRAN scenario built from industrial testbed measurements and operational traffic traces. TRACE consistently outperforms state-of-the-art AHD methods in resource assignment problems and yields more auditable heuristics at under 2% overhead.
cs.NE / 7 / 2610.02129
Spiking neural networks for streaming qubit readout
Abstract
Fast and accurate qubit-state assignment is essential for feedback, calibration, and error correction in quantum processors. In superconducting platforms, frequency-multiplexed readout makes this task intrinsically multivariate as measured traces can encode crosstalk, qubit-state relaxation events, and other transient nonidealities that are not fully captured by conventional matched filtering. Here, we introduce spiking neural network (SNN) discriminators for superconducting qubit readout. By processing the measurement window in successive time chunks, the networks exploit temporal structure and update classification scores as data arrive, rather than waiting until the end of the readout window. The spiking networks outperform matched-filter discrimination and approach the accuracy of a full-trace artificial neural network. Beyond reaching the performance of artificial neural networks, the key advantage of SNNs is that they provide a streaming, time-resolved estimate of the qubit state that evolves as the readout signal is acquired. Using quantisation-aware training and hls4ml synthesis, we further demonstrate that each FPGA inference update can be completed before the next readout chunk arrives. These results establish spiking neural networks as a promising route to low-latency, real-time qubit readout on FPGA hardware, with broader implications for time-critical quantum-control and scientific-inference applications.
计算语言学 (cs.CL)
58
cs.CL / 1 / 2610.00492
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Abstract
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
cs.CL / 2 / 2610.00526
Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning
Abstract
In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections. We ask a linguistic version of this question: which linguistic operations can be amortized out of the prompt? We train a 2.6M-parameter network that reads the geometry of a few-shot support set (centroid, principal subspace, spectrum, computed once and cached) and produces an input-conditioned additive update to the query's residual stream at a mid-depth layer of a frozen GPT-2-large/XL. Across eight inflectional directions and one lexical relation, under a canonical split that bars inverted-pair leakage between directions, three regimes emerge. On forward inflection, where 10-shot ICL is strong (0.67-0.89) and extracted task vectors collapse (<=0.06), the transform matches ICL at strictly zero-shot per-query cost. On lemmatization directions, which frozen GPT-2 can execute but 10 demonstrations systematically fail to convey (ICL 0.13-0.48 at 1.5B), the transform is not capped by ICL at all: it reaches 0.78-0.92, up to +72 points over ICL (past to present: 0.85 vs. 0.13). On arbitrary pairings (antonymy) every amortizer plateaus near half of ICL at every scale, capacity, and seed tested. Controls show the support manifold acts as a causally necessary task fingerprint: wrong-task manifolds collapse accuracy to <=0.06, query-only variants cannot disambiguate tasks sharing an input space, and leave-one-task-out transfer is zero. Productive rules amortize into latent task representations, sometimes better than prompting can convey them; memorized pairings do not.
cs.CL / 3 / 2610.00610
Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA
Abstract
Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three tasks: risk-level classification, evidence phrase extraction, and multi-label factor identification. Our approach adapts Qwen2.5-Instruct models using quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. We jointly train across all three tasks for risk classification, jointly train on Tasks~1a and 1b for evidence extraction, and adapt Task~2 separately for factor identification. We also tailor aggregation to each output: we average risk-level probabilities from the 32B and 72B models, combine evidence phrases through cross-fold consensus, and calibrate factor-specific decisions through rate matching based on out-of-fold operating points. On the official leaderboard, the final system achieved a composite score of 0.7738, with 0.8089 on Task~1 and 0.6919 on Task~2. Across the evaluated configurations, three-task training performed best for Task~1a, joint training on Tasks~1a and 1b performed best for Task~1b, and task-specific training performed best for Task~2. Probability averaging further improved Task~1a when component models had complementary errors. These findings highlight the value of tailoring both training objectives and aggregation strategies to the output structure of each task within a unified language-model framework.
cs.CL / 4 / 2610.00621
Mixture of Decoders for Diverse Dialog Response Generation
Abstract
Mixture modeling is a long established machine learning technique for learning large sets of multi-modal data. While it is known that sequence-to-sequence models for dialog response generation suffer from the problem of low diversity, we hypothesize that it is because sequence-to-sequence models tend to learn a degenerate uni-modal distribution of responses. We then propose to incorporate a mixture of decoders into sequence-to-sequence models and try to make each decoder learn specialized topics in order to improve the diversity of generated responses. Our model is developed under the framework of conditional variational autoencoder (CVAE). We evaluate our approach on an open domain chat corpus and show improvement over strong baselines in quantitative measures and human evaluation.
cs.CL / 5 / 2610.00650
Self-Evolving Coding Rules for AI Coding Agents
Abstract
The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, we propose RuleEvolve, a self-evolving framework for coding rules. RuleEvolve maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleEvolve outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).
cs.CL / 6 / 2610.00664
PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation
Abstract
Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.
cs.CL / 7 / 2610.00673
Closing the Loop: Practical Training Recipes for Looped Language Models
Abstract
Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-stage schedules. (2) Under controlled comparisons, our 1.4B LoopLM outperforms a parameter-matched dense model trained on the same data and token budget on all 12 evaluated benchmarks, including +14 points on GSM8K, +10 on MATH, and +22 on DROP. At matched inference compute, it approaches a 3.9B dense model on mathematical reasoning and reading comprehension while using only 36\% as many parameters. (3) We introduce a minimal recipe for converting pretrained dense models into looped ones: a single learned input-mixing scalar and a smoothed exit loss, with no step-specific parameters. Applied to Qwen3-1.7B-Base, Looped Qwen improves over an identically continued dense baseline on every evaluated benchmark across two data regimes, with statistically clear gains on GSM8K, MATH, and MMLU-Pro on the curated mixture. Together, these results make looped language models substantially cheaper to train from scratch and practical to introduce into existing pretrained checkpoints, while isolating the gains due to recurrence itself.
cs.CL / 8 / 2610.00679
Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow
Abstract
Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us $\textit{why}$ tuning on a $\textit{Bayesian}$ or an $\textit{oracle}$ (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.
cs.CL / 9 / 2610.00694
How Divergence Becomes Decision Flips in Compressed Language Models
Abstract
Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed copies of 19 open models on five corpora and nine mechanically unrelated perturbation families, the rate at which the arg-max token changes (the \emph{flip rate}) tracks total variation at a ratio with median $1.05$, with no fitted constant. KL converts into flips only through its square root and a factor that varies fourfold across models and corpora, because KL averages over tokens before the root is taken; first-order statistics averaged per token, such as Hellinger distance, avoid this, but reports rarely give them. As a result, of two compressors reported on different models and corpora whose flip rates differ by at least $10%$, KL assigns the smaller divergence to the one that changes more decisions in $11%$ of cases, total variation in $1%$. Two pre-registered tests mark the limits: on a held-out code corpus the ratio held for all eight models while three predictions about KL each failed for half of them or more, and on three new models with real kernels it stayed in its band for 37 of 38 checkpoints but fell below one on code for two models. In vLLM speculative decoding, total variation measured under teacher forcing predicts greedy draft acceptance with a mean relative error of $1.1$--$2.4%$, without the task-specific calibration that KL needs.
cs.CL / 10 / 2610.00724
Reason in Style: Discovering and Controlling Style in Language Models
Abstract
Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@$k$ over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.
cs.CL / 11 / 2610.00779
Effective Synthetic Data Curation Requires Group-Level Signals
Abstract
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
cs.CL / 12 / 2610.00833
VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence
Abstract
Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify every fact in the prose. At r=0 and r=1, we test 900 instances per setting (450 grounded-ungrounded pairs) with GPT-4o-mini, Llama-3.3-70B, and Claude Sonnet 4.6. Under this schema-level contract and before repair, 80.3% of mini claims and 47.9% of Sonnet claims fail. These are verifier rejection rates, not prose-hallucination rates. One repair pass raises claim survival from 19.7% to 28.0% for mini and from 52.1% to 54.3% for Sonnet. Verified claims per example change by +0.14 for mini, -0.71 for Llama, and -0.47 for Sonnet, so survival and output volume must be reported together. A second Sonnet pass gives no clear gain. At r=1, Gate 4 covers 97.0%, 98.7%, and 100% of failing claims for mini, Llama, and Sonnet. Small human studies support the rules but show gaps between schema checks and correct prose. A domain-specific GPT-4o judge test shows an order effect, so it is only a usefulness check. We release the code and data.
cs.CL / 13 / 2610.00852
Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis
Abstract
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) across two initialization strategies (Adult PhonoQ and scratch). Generalization is evaluated against manual child-speech annotations. On 1,352 consonant targets from 58 typically developing children, child-speech adaptation improves voicing recognition across all supervision conditions, from 0.922 macro-F1 for Adult PhonoQ to 0.972--0.987 after adaptation. Manner is more sensitive to alignment supervision: Adult+Child MFA reaches 0.804 and 0.796, compared to approximately 0.70 under Adult MFA supervision. Place remains comparatively strong across systems (0.871--0.902), although per-class performance varies substantially. The velar-fronting contrast is preserved across all seven model variants. Longitudinal UltraPhonix analysis further reveals speaker-specific velar and post-alveolar changes that are largely preserved across models and broadly consistent with reported clinical progress.
cs.CL / 14 / 2610.00883
DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
Abstract
Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.
cs.CL / 15 / 2610.00910
The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention
Abstract
Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes when the question switches from what Alice eats to what Bob eats. Averaged over many lists, this change is a steering vector, which we call the \emph{ordinal vector}. It points to a fact by its \emph{order of mention}, the order in which the facts were stated in the context. We find that LLMs represent the fact a question asks about by its order of mention, not by the name the question contains. We state this as the \textit{ordinal addressing hypothesis}: each order of mention has a \emph{fact address} in the model's state, shared by all contexts, and a question moves the state to the fact address of the fact it asks about, while the context supplies what that fact says. Across Qwen, Gemma, and Llama, fact addresses are (1) \emph{ordered by mention}: query states are organized by the order of facts, not of names, even when one fact has multiple subjects; (2) \emph{steerable}: added to a question about the first fact of a new list, the ordinal vector makes the model answer with the second fact of that list; (3) \emph{low-rank}: they span a low-rank subspace in which the first-mentioned fact is the easiest to reach, surprisingly similar to human recall; and (4) \emph{emergent}: they are shared in late-middle layers, hold from 1.5B to 32B parameters, and form early in pretraining. Language models reach a stated fact by where it was mentioned, deepening our understanding of LLM reasoning.
cs.CL / 16 / 2610.00983
The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization
Abstract
Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss. A common practice is to use mean squared error (MSE) as the reconstruction loss function, yet its induced optimization behavior remains largely unexplored. In this work, we take a holistic view of sequential quantization and systematically investigate how optimization evolves from the first quantization stage to the last, aiming for a deep understanding of optimization in learning-based PTQ schemes. Through extensive empirical studies spanning representative learning-based PTQ methods, LLM families, model scales, architectures, quantization settings and various tasks, we consistently uncover Optimization Imbalance: reconstruction loss magnitudes vary dramatically across stages, accompanied by highly uneven gradient magnitudes and parameter updates under MSE. We term the cross-stage range of loss magnitudes the reconstruction loss scale, and reveal that MSE translates the unexpectedly large reconstruction loss scale into highly uneven gradient magnitudes, which in turn lead to uneven optimization strength across quantization stages. This finding suggests a general principle for improving learning-based PTQ: optimization strength across stages should be decoupled from the reconstruction loss scale. Theoretically, we show that root mean squared error (RMSE) variants defined at the sample, channel, token, and element levels naturally realize this principle through implicit gradient normalization, outperforming MSE significantly as a drop-in replacement.
cs.CL / 17 / 2610.01046
Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study
Abstract
Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicated. This work audits sentence-specificity scoring artifacts on technical documentation and tests whether scores applied only after generation help choose among fixed LLM-generated revisions. Across Wikipedia and three technical-documentation corpora, the fixed general-domain predictor SpeciTeller and the pinned post-publication author-repository implementation of Ko et al.'s target-adapted predictor produce different corpus orders and same-sentence rank agreement from -0.066 to 0.510. Strict filtering and token-length adjustment change these patterns without reconciling them. In the Gemma set, SpeciTeller ranking raises direction-valid selection from 71.7% to 83.3% (+11.7 points; 95% source-case bootstrap interval +1.7 to +21.7); in the GPT-OSS-120B set, SpeciTeller ranking raises direction-valid selection from 51.7% to 56.7% (+5.0 points; 95% source-case bootstrap interval -6.7 to +16.7), and every primary single-score GPT-OSS-120B interval includes zero. These findings tie score interpretation and decision value to the predictor and candidate set.
cs.CL / 18 / 2610.01054
Capturing In-Context Learning Dynamics with Task Operators
Abstract
In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at https://github.com/gzxiong/task_operator.
cs.CL / 19 / 2610.01064
JoinGR: Learning to Traverse Join Graphs for Table Retrieval
Abstract
Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, but this ignores a key source of evidence: some required tables are not mentioned in the question and become identifiable only through their join relationships to already relevant tables. We introduce JOINGR, a join-aware table retrieval method that treats the database join graph as the retrieval space. Columns are represented as graph nodes, while intra-table and foreign-key relationships are represented as typed edges. Given a question, JOINGR selects semantically similar anchor tables, traverses join edges with a query-conditioned scorer, and aggregates the resulting edge deposits into table scores. The scorer is a lightweight MLP on top of frozen query, node, and edge embeddings, trained with a pairwise margin loss over gold tables. On BIRD and Spider datasets, JOINGR is competitive with the strongest retrieval baselines. On BEAVER, a challenging enterprise benchmark with multi-hop table requirements, JOINGR substantially improves recall over dense retrieval and re-ranking baselines. Cross-domain experiments show that the learned scorer transfers across benchmarks, indicating that the method captures reusable joingraph traversal behavior.
cs.CL / 20 / 2610.01082
Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models
Abstract
Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testset, alongside established BOUQuET and FLORES benchmarks. Results show our rule-based system is consistently strongest on SiLTT and BOUQuET datasets and that TranslateGemma fine-tuned on a curated dataset improves over strong neural baselines but does not surpass the rule-based system in dialectal settings. We release SiLTT and our best neural model to support further research.
cs.CL / 21 / 2610.01108
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Abstract
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37$\times$ at batch size 1 and 4.76$\times$ at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
cs.CL / 22 / 2610.01118
Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
Abstract
A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand LLM calls per memory bank and up to several thousand context tokens per query. We argue that association is a learnable relevance: the pointwise mutual information of memories under how human lives unfold. We introduce Madeleine, which learns amortized association: offline, an LLM life simulator writes simulated lives, whose cue-trigger pairs teach a query encoder a residual association on top of frozen similarity; online, it calls no LLM and plugs into any vector memory by replacing only the query encoder. On LoCoMo-Plus under the official protocol, Madeleine (I) reaches 66.6 when plugged into HyperMem, the highest among all systems evaluated under this protocol; (II) used alone, reaches the score of HyperMem as released (52.4 vs. 52.9) with zero LLM calls and about 1/21 of its answer context; and (III) lifts T-Mem by 26.2 points, significantly outperforms the same untrained backbone inside both systems, and leaves ordinary QA intact on the 4B backbone.
cs.CL / 23 / 2610.01150
BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text
Abstract
Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class. The resource was constructed using a corpus-grounded synthetic procedure incorporating regional variation in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, vocabulary, and Bengali-script spelling conventions while preserving the underlying hostile or abusive meaning. Descriptive analysis shows broadly comparable sentence-length distributions but partially distinct lexical spaces across the four classes. Pairwise Jaccard vocabulary similarity ranges from 0.37 to 0.56. The primary task is four-class regional dialect identification rather than binary abusive-text detection. The dataset is publicly available through Zenodo under a Creative Commons Attribution 4.0 license. The current version is intended as a research and prototyping corpus rather than a native-speaker-validated gold-standard linguistic resource. Keywords: Bangla, Bengali, dialect identification, regional dialect, abusive language, low-resource NLP, Chattagram, Sylhet, Barishal, dataset
cs.CL / 24 / 2610.01170
HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix
Abstract
Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.
cs.CL / 25 / 2610.01218
AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation
Abstract
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining deterministic checks, local guardrails, and structured LLM evaluation; a stratified beta-binomial gate that quantifies regression risk probabilistically; and a mandatory meta-evaluation protocol to validate the LLM judge before it influences decisions. Since engagement data is proprietary, we evaluate the judge layer on RAGBench, a public benchmark of 100k annotated RAG traces across 12 datasets. On identical stratified test samples (N=1200 per judge), a low-cost judge (gpt-4.1-nano) detects non-adherent answers barely above chance (AUROC 0.603 [0.570, 0.634]), despite producing flawless protocol output, while gpt-4o reaches 0.783 [0.756, 0.807] -- yet its per-domain performance still ranges from 0.62 to 0.88. A fixed-seed gate study spanning regression, no change, and improvement quantifies unsafe promotion, false-alarm cost, and improvement throughput. Under regression, the decision-grade profile reduces unsafe promotion to 22.2%-35.1%, against 29.3%-41.8% for a naive gate. These results support the design choices that judge quality must be measured per engagement and that point estimates alone are not a release decision.
cs.CL / 26 / 2610.01234
ASCRIBE: Atomic and Significance-Based Reasoning for Thai Clinical SOAP Note Generation
Abstract
Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physician-inspired reasoning framework that ascribes a clinical-significance level to each extracted atomic fact in the conversation before summarization, making a general-purpose LLM a more reliable scribe. We also release ThaiClinicBench, the first de-identified Thai clinical summarization benchmark of real encounters, together with a synthetic training corpus derived from real clinical notes. As a prompt, ASCRIBE outperforms chain-of-thought prompting on GPT-5.4 and Gemini 3.1 Pro across the physician-aligned LLM-judge metrics and improves on standard prompting by up to 10.3 points on the completeness LLM-judge metric. As a GRPO reward, it enables a Gemma-4-E4B model trained solely on synthetic data to match Gemini 3.1 Pro in factual precision and surpass it in completeness. Code and data can be found at https://github.com/loolootech/ascribe.
cs.CL / 27 / 2610.01235
Harness Annealing: Learning to Act with Less External Control
Abstract
Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories can improve task performance while leaving these decisions dependent on runtime intervention. We ask whether harness-supported experience can also teach the model to make these decisions, allowing the division of control to change as the model learns. We call this objective harness internalization: learning to assume specified control responsibilities while retaining task performance after the corresponding support is withdrawn. We introduce HARNESS ANNEALING TRAINING (HAT), which combines explicit control supervision with a curriculum over teacher trajectories collected under progressively weaker harnesses. Experiments with 9B and 35B models on SWE-QA and SWE-QA-Pro evaluate every checkpoint under four deployment harnesses. Selected annealed checkpoints operating with tools alone achieve scores close to those of their respective starting checkpoints deployed with the full harness. The benefits vary with model scale and deployment configuration, and further annealing does not uniformly improve performance. These findings suggest that harness-supported experience can help reduce the runtime control required by a trained agent.
cs.CL / 28 / 2610.01244
Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration
Abstract
Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introduce OffQuery, which separately evaluates evidence verification (T1), shared-state reconstruction (T2), and task resolution (T3) in two representative high-stakes settings: healthcare and disaster response. Across GPT, Gemini, and Qwen models, standard collaboration shows much stronger task performance than state reliability. Averaged over 21 model--setting combinations, task resolution reaches 64.7%, while evidence verification and state reconstruction reach only 14.3% and 43.1%. We trace this gap to selective information use: current queries often bypass corrupted facts, which become consequential when later tasks require them. We further introduce ReGround, which resolves conflicting evidence, verifies shared facts, reconstructs a trusted state, and reasons over that state. Across seven models from three families, ReGround improves all three capabilities in every evaluated setting, with average relative gains of 309.0%, 82.9%, and 17.6% on T1, T2, and T3. Reliable collaboration therefore requires both a correct decision and a reliable shared state for future reasoning.
cs.CL / 29 / 2610.01257
Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems
Abstract
Scientific progress emerges from a longitudinal ecosystem in which researchers, institutions, funding agencies, collaboration networks, and the scientific literature co-evolve. As AI becomes increasingly involved throughout the scientific research cycle, understanding these interconnected and evolving processes becomes increasingly important. We introduce SciUtopia, a persistent, closed-loop LLM-agent simulation framework for studying academic research ecosystems. SciUtopia models interconnected scientific processes such as research-direction choice, collaboration, submission, peer review, resubmission, citation, funding, and researcher attrition, while maintaining evolving states across simulated years. Its configurable institutional mechanisms and information channels provide a controlled testbed for matched counterfactual experiments and targeted interventions. Across 61 simulation worlds, SciUtopia simulates over 40,000 researchers from 8,000 institutions, producing around 400,000 publication decisions and 1.2 million LLM-generated peer reviews. Using these longitudinal simulations, we find that rejection-driven resubmission substantially amplifies reviewer burden beyond population growth alone, cautious exploration balances citation impact with career success and long-term topic diversity, and resource inequality can emerge even without detectable cumulative advantage from narrowly winning early funding. Code is available at https://github.com/Ahren09/ScienceUtopia.
cs.CL / 30 / 2610.01316
What Wins a Vote? Formatting, Length, and Lexical Diversity in the French Compar:IA LLM Arena
Abstract
LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 137,113 battles across 116 models, and the joint estimates use the 127,092 battles with all required measurements. For each battle, we reconstruct the response visible when the user voted. We then compare the raw ranking with rankings adjusted for formatting, length, readability, vocabulary variety, and sentence structure. Presentation is associated with winning, but length, bold text, and lists tend to occur together, making their individual contributions hard to separate. Across the measured features, two associations change least across specifications: bold usage (+11.0% win odds per standard deviation in the joint model) and moving-average type-token ratio (MATTR), a measure of vocabulary variety that is less sensitive to answer length (+16.8%). The bold association is substantially smaller in observed multi-turn conversations, whereas the MATTR association changes little; because users choose whether to continue, this difference is descriptive rather than causal. The full adjustment moves 36 of 116 models by at least ten ranks. Yet comparisons with external benchmarks do not show that adjusted rankings better measure capability. We therefore recommend publishing raw and adjusted rankings side by side as a transparent sensitivity analysis.
cs.CL / 31 / 2610.01324
Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes
Abstract
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.
cs.CL / 32 / 2610.01345
ARCCS: An Automated Regulatory Compliance Checking System
Abstract
Regulatory compliance checking - deciding whether a target document satisfies the obligations of a regulation - requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications. This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. We evaluate ARCCS in two complementary settings. First, in a GDPR policy-document evaluation, LLM-based judges find its decisions and justifications legally and evidentially consistent in up to 96.67% of the assessed cases. Second, on an EU public-procurement benchmark comprising more than 1,200 individual rule checks, the system attains 98.8% accuracy in violation detection. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.
cs.CL / 33 / 2610.01393
LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction
Abstract
Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for ontology network construction that automates the discovery and generation of both intra-domain and inter-domain relationships. Our approach combines domain-adapted DistilBERT embeddings for dense contextual representation, clustering-based pre-filtering to reduce the candidate search space, and GPT-4o-driven relationship generation via iterative prompt engineering to produce semantically rich, interpretable links. Applied to ReproduceMeON - a network of 33 ontologies spanning machine learning, microscopy, computational science, and experimental workflow - the pipeline reduces approximately 800k raw concept pairs to 95k high-quality candidates. Human expert validation of 429 generated relationships by two independent annotators yields an overall precision of 80.19% (91.49% on high-certainty annotations) and an F1 of 0.890, with substantial inter-annotator agreement. Comparative experiments against five similarity-based baselines, including Sentence-BERT, show a substantial performance gap (best baseline F1 = 0.581), while an ablation study demonstrates that similarity-based methods alone fail to discriminate valid from invalid relationships (AUC approx 0.5) on the filtered candidate set. These findings highlight the necessity of LLM-based reasoning over concept roles and domain semantics for accurate relationship construction.
cs.CL / 34 / 2610.01427
SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic
Abstract
Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotated Multi-dialect Speech), a benchmark comprising 1,300 utterances drawn from open audio corpora, balanced across five LA varieties (Urban and Rural Palestinian, and Urban Jordanian, Lebanese, and Syrian). Each utterance is represented across four aligned tiers: audio, unvocalized orthography, diacritized text, and phonetic transcription. This structure supports evaluation of various downstream tasks such as diacritization, grapheme-to-phoneme conversion, automatic speech recognition, and audio-to-phoneme, grounded in audio and stratified by variety. We benchmark open and proprietary models across these tasks to demonstrate the utility of this benchmark for measuring progress across LA. We release SHAMS at https://shams-nlp.github.io .
cs.CL / 35 / 2610.01490
The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents
Abstract
In February 2026, an always-on personal agent (``Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated ``heartbeat'' checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to ``Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emph{represented} from \emph{enacted} identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to ``I.''
cs.CL / 36 / 2610.01493
No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse
Abstract
Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $h_k$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric ($p > 0.23$), whereas $h_k$-filtering yields $+42\%$ unique trigrams, $+30\%$ vocabulary, and $-19\%$ repetition (all $p < 0.001$). We validate $h_k$ as a cross-domain entropy proxy ($β= 0.924$, $R^2 = 0.746$) and collapse detector ($ρ= +0.454$, $p < 0.0001$) across 4~domains, 2~temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
cs.CL / 37 / 2610.01511
GAW-PO: Preference Optimization with Gradient-Aligned Token Weights
Abstract
Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter $β$ decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.
cs.CL / 38 / 2610.01560
AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
Abstract
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.
cs.CL / 39 / 2610.01627
What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language
Abstract
Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.
cs.CL / 40 / 2610.01688
Compound interpretation is based on analogy
Abstract
How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.
cs.CL / 41 / 2610.01702
Task-Oriented Rank Adaptation for Continual Learning in Text Classification
Abstract
Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.
cs.CL / 42 / 2610.01828
The Asymptotics of Language Model Alignment with Memory
Abstract
Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.
cs.CL / 43 / 2610.01921
Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
Abstract
Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
cs.CL / 44 / 2610.02019
Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization
Abstract
The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .
cs.CL / 45 / 2610.02040
Typological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages
Abstract
Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.
cs.CL / 46 / 2610.02076
LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
Abstract
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
cs.CL / 47 / 2610.02122
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Abstract
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
cs.CL / 48 / 2610.02142
Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Abstract
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
cs.CL / 49 / 2610.02163
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
Abstract
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
cs.CL / 50 / 2610.02206
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Abstract
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
cs.CL / 51 / 2610.00809
Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics
Abstract
Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single $2\times8$ image grid built via simple shot-transition detection approaches full-video understanding ($κ$ within~.05), at $\sim 15\%$ of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.
cs.CL / 52 / 2610.01873
Where LLMs Fail with Visualization DSLs
Abstract
As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.
cs.CL / 53 / 2610.00964
RPTune: Learned Context Curation for LLM Catalog Search
Abstract
For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average.
cs.CL / 54 / 2610.01139
Do Multilingual Encoders Produce Language-Consistent Semantic IDs?
Abstract
Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.
cs.CL / 55 / 2610.01553
From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search
Abstract
Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.
cs.CL / 56 / 2610.01492
Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Abstract
Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.
cs.CL / 57 / 2610.01846
Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
Abstract
Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.
cs.CL / 58 / 2610.01450
Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Abstract
Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.
多智能体系统 (cs.MA)
5
cs.MA / 1 / 2610.00980
Can AI Scientists Coordinate at Runtime?
Abstract
Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets. Code is available at https://github.com/systemind-team/Runtime-AI-Scientist.
cs.MA / 2 / 2610.01364
LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing
Abstract
Factories are shifting toward smaller lot sizes with high product customization, requiring frequent re-programming of flexible and reconfigurable automation systems. LLM-based agents can be deployed in two complementary roles: Offline, they generate deterministic production sequences, reducing programming effort; online, they operate live machines and handle unforeseen runtime faults that static programs cannot anticipate. We propose a solution in which each factory module is paired with a dedicated LLM-based agent and an MCP tool server that exposes the module's skills via OPC UA method calls, with agents coordinating over MQTT and grounded by real-time updates of the factory state. We compare three agent architectures (orchestrator, peer-to-peer, and monolithic) across nine production challenges of increasing complexity in a simulation of a physical six-module hexagonal factory, including silent hardware fault detection. The monolithic and peer-to-peer architectures both achieve the highest mean solve rate (93\%), while the orchestrator uniquely resolves a silent conveyor-belt fault in all ten runs by autonomously rerouting plates around the blocked segment. All architectures exhibit emergent fault-diagnosis behavior without any explicit failure-handling logic, establishing standardized MCP tooling, MQTT-based inter-agent communication, and real-time state injection as a viable and reproducible foundation for LLM-programmed smart manufacturing.
cs.MA / 3 / 2610.01569
Managing Context and Communication in Distributed Agentic UAV Swarms
Abstract
Unmanned aerial vehicle (UAV) swarms increasingly rely on language-model agents to provide adaptive mission-level reasoning in uncertain environments. Fully distributed control, in which each UAV hosts an independent Small Language Model (SLM), removes reliance on a centralized coordinator but introduces an information-management problem: long-running interaction histories can degrade the reasoning context, while indiscriminate information dissemination increases communication and inference overhead. We address these challenges with a distributed UAV-agent architecture that enables continuous local SLM control through an event-driven reason-act-observe lifecycle. Runtime knowledge is represented as structured atomic notes and organized into core, local, and peer-specific memory. A deterministic interest-aware gossip engine selectively disseminates these notes according to recipient-specific semantic novelty and recency. We evaluate the architecture using ten UAVs in a simulated search-and-rescue mission. Our approach completes all experimental runs, whereas unrestricted flooding messages completes only 70-85\%, and delegating forwarding decisions to the SLM prevents mission completion in every run. Compared with unrestricted flooding, our approach approximately halves inference-token consumption, reduces transmitted data, and achieves lower survivor-count error.
cs.MA / 4 / 2610.01630
After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning
Abstract
Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.
cs.MA / 5 / 2610.02118
Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation
Abstract
This paper presents a decentralized power-optimal coordination framework for magnetically actuated spacecraft swarms. Swarms that form large space structures overcome the aperture limit set by the launch vehicle and hold their shape on solar-generated power alone. Magnetic actuation is propellant-free and generated by a magnetorquer, which is commonly used for attitude control. However, every spacecraft interacts with every other within range, and its effect depends on the actuation power and a carrier frequency. We therefore design a decentralized power-optimal framework to jointly derive the interaction graph, frequency grouping, and controller gains. Our decentralized controller preserves angular momentum, which is a nonholonomic constraint. Then, this framework for connected groups whose memberships overlap across carriers guarantees that the relative position errors, the absolute attitude errors, and the imbalance of the reaction-wheel momenta converge to the desired states under the decentralized power-optimal allocation. A closed-loop simulation of a thousand spacecraft with the complete alternating-current interaction confirms the framework. A fast approximate integration with a proven error bound extends the framework to a long-horizon orbital reconfiguration held with high precision.
软件工程 (cs.SE)
9
cs.SE / 1 / 2610.00885
FORALL-LEAN-AGENT for Auditable Reasoning in Formal Mathematics and Software Verification
Abstract
Coding agents increasingly automate Lean proof development, but successful compilation alone does not establish that a candidate proves the intended statement under acceptable assumptions. We present FORALL-LEAN-AGENT, a frontend-agnostic framework for auditable reasoning in formal mathematics and software verification. The framework combines isolated workspaces, Lean tools, and fresh review with statement comparison, axiom audits, and independent proof checking where supported. Verification evidence and reviewer decisions are bound to the same candidate artifact, making acceptance traceable. We evaluate the framework on VeriSoftBench, PutnamBench, and both problems in the Lean Eval softwareverification track. On the 100-task VeriSoftBench subset, integration with FORALLLEAN-AGENT raises benchmark-rule success from 93 to 100 for GPT-5.6 Sol at low effort while reducing cost from $69 to $62. The PutnamBench evaluation accepts all 672 problems at an average of $4.72 each. These results show that agent harness design can improve correctness and efficiency while providing evidence beyond aggregate solve counts.
cs.SE / 2 / 2610.00905
Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems
Abstract
With the advancement of LLM-based multi-agent systems (MAS), an increasing number of opensource projects are adopting multi-agent architectures as the foundation of their core functionality. Although research and practice on MAS have attracted considerable attention, limited studies have explored the challenges faced by practitioners of open-source LLM-based MAS, the causes of these challenges, and potential solutions. To address this gap,we conducted an empirical study to understand the issues that practitioners encounter when developing and using open-source LLM-based MAS, the possible causes of these issues, and potential solutions. We collected 22,848 closed issues from 21 open-source LLM-basedMASand applied a mixed automated and manual filtering approach to reduce the dataset to 944 issues related to LLM-based MAS.We then analyzed these issues to understand the frequent issues encountered by practitioners, their underlying causes, and potential solutions. Our study results show that (1) Orchestration & Execution Issue is the most common issue faced by practitioners, (2) Workflow Problem, Tool Integration Problem, and Memory Problem are identified as the most frequent causes of the issues, and (3) Optimize Workflow is the predominant solution to the issues. Based on the study results, we derive empirically grounded implications for practitioners and researchers aimed at improving orchestration, tool integration, and memory mechanisms in LLM-based MAS.
cs.SE / 3 / 2610.01023
Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
Abstract
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.
cs.SE / 4 / 2610.01073
Safety Must Survive Self-Improvement: Why Failures Persist and How Agents Recover
Abstract
Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization tasks, where fixed LLM editors optimize executable agent components and independent traces record their effects. Paired interventions separate which revisions pass validation, which program continues running, and which program the editor revises next. After a new authorization dependency invalidates previously tested optimizations, historical scores preserve the same unsafe programs in 22 of 48 framework histories despite a correct alternative in every affected archive. Refreshing scores restores correctness on the original suite, with residual failures on independently composed tests. Failures also persist under an unchanged contract when all proposals are rejected and the failed incumbent remains active. Starting from shared failures, editing the initial implementation instead of the failed one improves recovery, although the advantage varies across editors. Full validation and validated rollback end with fully correct programs in the core trajectory study while retaining over 43% deployment savings. Preserving agent safety requires checking what will run under current conditions and choosing which implementation to edit next.
cs.SE / 5 / 2610.01203
CompProv Produces Machine Readable Graphs Encoding Microscopic Algebraic Provenance for Reproducible Computation
Abstract
The reliability of computational results in scientific research and financial modeling increasingly depends on verifiable traceability, not merely on trust in a reported output. Existing provenance systems operate at the granularity of files, datasets, or pipeline stages, leaving the internal sequence of algebraic transformations connecting an algorithm's inputs to its outputs unrecorded; a rounding error, an undocumented substitution, or a missing intermediate value can propagate to a final result with no recoverable trace. This work presents CompProv, a Java-based, audit-oriented provenance framework that captures lineage at the atomic level of individual algebraic operations by encapsulating numerical values in high-precision wrapper objects, producing a serializable Calculation Provenance Graph (CPG) that persists as a self-contained artifact rather than a discarded byproduct. The framework is evaluated through three heterogeneous case studies: a decentralized-finance NAV calculation, a reconstruction of an interferometric gauge-block calibration in metrology under explicitly documented input assumptions, and a hydrological model performance evaluation. Deterministic replay reproduced each result exactly in a fresh environment, and CPG-based input substitution supported sensitivity analysis without exposing the underlying source code. These results indicate that, provided the CompProv runtime and wrapper classes are available in the replay environment, a CPG allows numerical integrity to be audited without disclosing proprietary business logic, and that its self-contained structure supports temporal auditability once an execution environment has become deprecated. This framework establishes an operation-level foundation for auditable-by-design computational systems, with scaling to high-throughput computing identified as a direction for future work.
cs.SE / 6 / 2610.01376
Refactoring React Component Hierarchies to Eliminate Prop Drilling
Abstract
In React front-end development, Prop Drilling is the practice of propagating data through component properties across multiple levels of the component hierarchy. Despite being discouraged by React documentation and characterized as a code smell by recent research, little is known about its prevalence, complexity, and potential for automated refactoring. In this work, we propose a static analysis method for identifying Prop Drilling instances in a React codebase and eliminating them using two automated refactoring strategies based on Context API and Component Composition. The proposed method is implemented as a Node.js command line tool, ReactRefactor, and empirically evaluated on a benchmark dataset of open-source React applications. The main findings of the empirical evaluation indicate (a) the frequent occurrence of prop drillings across benchmark projects, irrespective of project size; (b) their generally low to moderate complexity in terms of data propagation depth, contrasted with the more prevalent complexity arising from the simultaneous forwarding of multiple properties along the same path; and, (c) the potential for automated elimination of a substantial share (76.3%) of prop drillings, primarily through refactoring to the Context API.
cs.SE / 7 / 2610.01611
Architectural Degradation: How to Measure and to Remediate
Abstract
Context. Architectural degradation undermines software maintainability, evolvability, and quality. However, existing research remains fragmented across measurement approaches, metrics, tools, and remediation strategies, limiting our understanding of how these elements relate across the degradation lifecycle. Aim. We consolidate the state of the art on architectural degradation by examining how researchers measure it, which metrics and tools support its assessment, and how existing approaches address remediation. Method. We conducted a Multivocal Literature Review of 284 peer-reviewed and grey-literature studies. We supported screening, data extraction, and classification with a locally executed LLM-assisted pipeline combining Retrieval-Augmented Generation, multi-model validation, and human adjudication. We then analyzed the resulting taxonomies and their cross-dimensional relationships. Results and Conclusions. We identified 277 measurement approaches, 357 metrics, 238 tools, and 395 remediation approaches. Research strongly concentrates on static and structural analysis, structural metrics, and detection-oriented tools. In contrast, remediation spans heterogeneous code-level, architectural, and organizational interventions and shows substantially less consolidation. Overall, the field has developed a mature diagnostic apparatus but has made less progress in connecting degradation detection with effective remediation. Our results provide a structured view of the available techniques and identify the diagnosis-remediation gap as a key direction for future research.
cs.SE / 8 / 2610.01664
Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
Abstract
Code detectors can become obsolete as code-generating models evolve: a detector validated on one generation of models may not transfer to the next. We call this limited useful life a detector half-life. We evaluate eight general-purpose LLM judges and three dedicated detectors on human-written code and code produced by seven generators across C++, Java, and Python. Our results reveal two problems. First, performance varies considerably across generators and prompting strategies, suggesting that some detectors rely on generator-specific patterns rather than general evidence of code provenance. Second, accuracy can conceal severe prediction bias. DetectCodeGPT and GPT-Sniffer achieved an accuracy of 0.50 but an F1 score of 0.00 across all generators because they classified almost every sample as AI-generated. However, general-purpose LLM judges achieved stronger accuracy and F1 scores. Our results show that general-purpose LLMs are promising training-free judges of code provenance and can outperform dedicated detectors. However, their reliability depends on the judge model, the code generator, and the prompting strategy. We therefore recommend evaluating LLM judges across multiple generators and reporting macro-F1 alongside class-specific precision and recall.
cs.SE / 9 / 2610.01769
CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Abstract
Coding agents can generate code that appears correct but implements behavior the user never intended. This mismatch can arise when an agent silently resolves underspecified requirements through its own assumptions. As subsequent development builds on these assumptions, correcting the resulting behavior can become increasingly costly. Early clarification can help prevent such mismatches, but unnecessary questions can interrupt developers and slow down development. Existing methods struggle to identify key clarification questions while avoiding unnecessary ones. Therefore, we propose CONTRA, a training-free method that combines broad question discovery with semantic and execution-based question qualification. CONTRA first generates candidate questions and filters out those unrelated to required behavior or already resolved by the requirement. For each remaining question, it generates programs conditioned on two plausible answers and checks for stable behavioral differences on shared inputs. It then uses the interaction history to select among qualified questions or stop asking. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points. With the same LLM and evaluation protocol, CONTRA also achieves higher clarification recall and F1 than the coding harnesses Claude Code and OpenHands. To support practical use, we also implement CONTRA as a Claude Code plugin that integrates selective clarification into everyday development.
操作系统 (cs.OS)
1
cs.OS / 1 / 2610.00714
MANTA: Machine Learning Augmented Tiering Advisor
Abstract
Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory. Its effectiveness depends on keeping useful pages in the fast tier, but existing heuristic policies can lag behind changing hot sets in phased or bursty workloads. To explore these limitations, we introduce ChOMP, a scalable offline optimizer that minimizes placement and bandwidth-sensitive migration costs. We then develop a trace-driven simulator that uses this reference to identify performance opportunities for online policies. Motivated by these results, MANTA predicts future page usefulness from runtime access features and integrates a lightweight learned model into ARMS. Across eight workloads on emulated CXL, MANTA achieves geometric-mean speedups over ARMS of 1.12$\times$ and 1.08$\times$ at 4~GB of fast memory on Linux 6.2 and 6.18, respectively; across six Optane workloads, it achieves 1.69$\times$. On individual workloads, MANTA is up to 1.25$\times$ faster than ARMS with emulated CXL on Linux 6.2, 1.21$\times$ on Linux 6.18, and 5.6$\times$ with Optane.
硬件架构 (cs.AR)
10
cs.AR / 1 / 2610.01186
From Physical Devices to RTL Models: Abstraction and Validation in Hardware Engineering
Abstract
This paper introduces the foundational principles underlying hardware engineering models and argues that abstraction is their defining characteristic. Because abstraction necessarily omits detail and constrains what engineers can build, models are inherently incomplete in specific respects - or, as George Box famously observed, "All models are wrong, but some are useful". At the same time, abstraction is essential for simplification, which is key to managing complexity. More abstract models also tend to simulate faster because fewer details must be considered. This paper subsequently examines a range of abstraction methods in digital design - sometimes referred to as design disciplines - including lumped models, value-discrete models, and time-discrete models. Together with constraints that define the validity of the abstraction and design guidelines, these abstraction methods establish design disciplines. This paper further relates these forms of abstraction to pre-clustered design elements such as transistors, gates, registers, and transfer functions. These pre-clustered elements define abstraction levels, such as the gate level, and are presented as a key enabler of increased design productivity.
cs.AR / 2 / 2610.01603
U-Sonic: An Open-Source 8-Channel Ultrasound Transmit IP in a 130 nm RISC-V SoC
Abstract
Miniaturized ultrasound (US) probes require programmable and synchronized transmit (TX) excitation across multiple elements, while existing compact platforms often rely on limited microcontroller (MCU) pulse generators or closed-source fixed-function pulser devices. We present U-Sonic, an open-source digital US TX peripheral integrated into a 32-bit RISC-V system-on-chip (SoC). The implemented SoC integrates 8 pulser cores, while the parameterized architecture supports up to 16 channels. Each core generates single- or dual-tone bursts with programmable period, duty cycle, pulse count, polarity, and idle level, together with optional inverted stop pulses for active damping. A shared memory-mapped Open Bus Interface (OBI) enables synchronous start and stop of arbitrary channel subsets and supports composite bipolar, gated, and three-level excitation schemes. Functional correctness was verified in Verilator against a Python golden model over 4379 checked cycles across directed and randomized configurations, and confirmed on a Terasic DE10-Lite field-programmable gate array (FPGA). The design was synthesized and placed-and-routed in IHP 130 nm. The post-layout area in kilo gate equivalents (kGE), scales as 1.65 kGE plus 1.66 kGE per channel. The 8-channel instance occupies 14.9 kGE, corresponding to approximately 14.3% of the 104 kGE SoC. The register-transfer level (RTL), register descriptions, verification collateral, and software support are released as open source.
cs.AR / 3 / 2610.01623
Open-Source Multi-Wire SPI Readout for Wearable Ultrasound Probes
Abstract
Wearable ultrasound probes must transfer increasingly large acquisition payloads while maintaining compact, low-power electronics. In TinyProbe, the current bottleneck in data transfer occurs between the acquisition FPGA and the wireless system controller. This work presents an open-source, multi-wire SPI readout interface that uses serial command and address phases followed by a build-time-selectable dual- or quad-lane payload phase that is intended to address this bottleneck by increasing the potential bandwidth over the wifi limit while retaining compatibility with the Microcontroller-centric wearable US architecture. The interface emulates a serial flash memory, enabling compatibility with a broad range of microcontroller families and their existing peripheral interfaces. On the FPGA, the data path connects the existing acquisition FIFOs to the SPI interface through clock-domain crossing, sample reshaping, and packing into 32-bit words. Dual-SPI readout is integrated into the existing IGLOO2/SiWG917 TinyProbe architecture and verified at an SCLK frequency of 5 MHz. A separate Kria K26 testbed is used to characterize the FPGA SPI interface independently of the acquisition and wireless subsystems, demonstrating error-free transfers at SCLK frequencies up to 66 MHz. These measurements identify the SiWG917 multi-lane SPI implementation as the next bandwidth-limiting component and motivate a future upgrade of the system controller. The HDL and MCU implementations are released under a permissive open-source license.
cs.AR / 4 / 2610.01867
ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference
Abstract
Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.
cs.AR / 5 / 2610.01918
Timing-Driven Logic Remapping with Local Physical Context
Abstract
The timing behavior of a mapped circuit depends on both its logic implementation and the physical environment in which that implementation is realized. Revisiting mapping decisions after placement therefore requires a search procedure that accounts for surrounding timing constraints, fanout loads, and interconnect effects. We study local remapping in this setting and develop a framework that couples discrete mapping search with physical implementation feedback. Timing-critical regions are isolated through bounded windows whose interfaces retain the context of the surrounding circuit. Within each window, a mixed-integer formulation jointly selects logic cuts, signal polarities, and library cells under a delay model informed by estimated locations and interconnect parasitics. A continuous relaxation filters the search space before discrete optimization produces alternative implementations with similar modeled timing and different structural choices. These implementations are reconstructed and assessed through legalization, routing-based parasitic estimation, and timing analysis. Physically validated improvements are incorporated into the design, and the updated context guides subsequent searches. The framework provides a systematic way to revisit local logic implementations while accounting for their interaction with an existing placement.
cs.AR / 6 / 2610.01975
CONFERM: Recurrence-Aware Temporal Mapping for Multi-Cycle Multi-Context CGRAs
Abstract
Throughput in DSP and machine learning workloads is often limited by two temporal structures, i.e., loop-carried recurrences and long-latency, multi-cycle compute nodes. On spatio-temporal coarse-grained reconfigurable arrays (CGRAs), both bottlenecks can be addressed by overlapping iterations across the multi-context modulo configurations. Yet, existing CGRA mappers schedule a fixed dataflow graph (DFG) that treats recurrence-aware scheduling and operator-level pipelining separately, limiting inter-iteration overlap and inflating routing pressure. To tackle this, we present CONFERM, a recurrence-aware temporal mapper that uses the dominant temporal con-straint to guide the DFG representation and expose opportunities for loop-carried pipelining. CONFERM identifies and prioritizes bottleneck regions during scheduling. The regular loop-carried offsets across interleaved iterations allow the emitted control sequence to repeat at a shorter cadence than the original initiation interval, thus delivering higher throughput with lower CGRA configuration overhead. Across ten benchmark kernels, CONFERM improves throughput by 2.18x over state-of-the-art mappers. Its uniform iteration offsets shorten the emitted initiation interval by 46%. CONFERM's mapper pass also converges faster by 5.07x on average with the same heuristic mapper backend.
cs.AR / 7 / 2610.02121
Catscan: Visualizing Pipelines of CPU Performance Simulation
Abstract
Processor pipeline visualization tools are routine inside industry CPU teams, but few of them are described or released publicly. As a result, students, researchers, and other practitioners rarely see the tooling that processor architects use to debug performance before silicon. This paper describes two pieces of Ampere Computing's performance- analysis infrastructure that we have released to the community as open source: event streams, a simulator-output format, and Catscan, an interactive viewer built around that format. Event streams record microarchitectural activity as typed events connected by transaction relationships, so a user can move between a symptom and the instruction, uop, or memory transaction that explains it. Catscan uses that structure to support resource- and transaction-oriented views, persistent highlighting, domain-specific search, comparative trace synchronization, and other workflows used during product development. In this paper we report the design choices that survived production use, the limitations we encountered, and the lessons we think are useful for future microarchitectural visualization tools.
cs.AR / 8 / 2610.01477
ALFRED: Requirement-driven development of an open-source mobile manipulator for long-term plant monitoring
Abstract
Tracking seasonal change in crops and forests requires observing the same plants repeatedly. Ground robots can do this at close range, and a manipulator gives their sensors more viewpoints. Yet the robots behind long-term field datasets are rarely released with their design files, and how a robot's own structure limits arm reach and occludes its sensors is seldom compared between builds. We present ALFRED, an open-source mobile manipulator built from commercially available components. It carries a six-degree-of-freedom arm, LiDAR, RGB-D cameras, RTK GNSS and an IMU on an Ackermann-steered base, all mounted on a reconfigurable aluminium strut frame, and runs containerised ROS software. It was developed through four builds against six requirements for repeated outdoor deployment: durability, modularity, repairability, sensing reach, endurance and reproducibility. Model-based analysis of the last three builds shows the usable share of the arm's reachable poses rising from 34.0% to 60.0% and then 66.1%, and ray casting shows that only the final build keeps the frame-mounted LiDAR's horizontal view clear both forwards and backwards. ALFRED completed a year of monthly forest surveys (528 traversals) without missing a scheduled collection. This was despite battery degradation, reconfiguration for another researcher's study, and the parallel development of ALFRED 2.0 for autonomous crop-row operation, with each switch between builds taking about six hours. The deployment also showed that mechanical modularity is only as dependable as the robot description that tracks it.
cs.AR / 9 / 2610.01602
Open-Source Live-Reconfigurable Multi-Mode Wearable Ultrasound
Abstract
Wearable ultrasound enables continuous deep-tissue monitoring, and a single programmable probe can operate in multiple complementary modes, such as structural A-mode and Doppler flow measurement. However, each operating mode requires dedicated measurement parameters and peripheral states, with no single configuration serving all modes on resource-constrained devices. Time multiplexing of operating modes introduces reconfiguration latency that lowers the effective mode repetition rate. To address this limitation, we present an open-source, transition-aware control stack for low-latency, in-session reconfiguration of the 32-channel TinyProbe wearable platform. Operating modes are described as hardware configurations, and host-side shadow registers track the peripheral states, enabling transition-specific register updates. Transition sequences are executed either by the host (over Wi-Fi 6) or by a firmware loop on the probe MCU. We validate the stack on a pulsatile-flow phantom by interleaving blocks of 25 to 100 pulsed-wave Doppler shots at 1.43 kHz PRF with single 16-channel A-mode acquisitions, changing channel configurations at every transition. Compared to full reconfiguration, the overhead per transition decreases from 30.2 ms to 11.6 ms (host-scheduled) and 3.1 ms (MCU-scheduled). For 75-shot Doppler blocks, the multi-mode repetition rate reaches 16.0 Hz (MCU-scheduled), 90.1% of the theoretical maximum of 17.7 Hz. Concurrent reconstruction of a Doppler spectrogram and a lumen-diameter trace demonstrates the functionality of time-multiplexed flow and structural monitoring.
cs.AR / 10 / 2610.00738
Q-MINO: A Minimal-Norm Method for Quantization-Aware Training
Abstract
The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates and parameter oscillations, particularly in ultra-low-bit regimes. We propose the Quantization-Aware Minimal-Norm Optimizer (Q-MINO), a temporal bundle method that combines gradient consensus, state-drift regularization, and an alignment constraint to construct stabilized, minimum-norm update directions from recent optimization states. Q-MINO solves the resulting constrained subproblem using a warm-started Frank--Wolfe procedure with a feasible fallback initialization. Theoretically, via a stochastic Lyapunov Kurdyka--Łojasiewicz (KL) framework, we show that Q-MINO achieves asymptotic neighborhood convergence. Moreover, we detail numerical experiments with Q-MINO at various quantizations.
密码学与安全 (cs.CR)
39
cs.CR / 1 / 2610.00519
Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS
Abstract
Ships broadcast their positions through the Automatic Identification System (AIS), and those reports can be false or missing. An operator who acts on an anomaly alert needs that alert to be attributable, reviewable, and recoverable after a failure. This paper describes Harbormaster, a production-shaped system on AWS that follows three rules. Physics checks run on every report before any learned model runs. The DynamoDB read store is an idempotent projection of PostgreSQL, and a guard on the log sequence number (LSN) of each change protects every write to it. A candidate model must pass a holdout gate and a shadow comparison before a canary gives it traffic, and a burn-rate check guards each canary step. The paper proves that replaying any prefix or suffix of the change log leaves each projected key at the value of its highest applied LSN. It also notes effects this result does not cover, such as repeated cache invalidations and repeated audit rows. In a bounded AWS window, a one-hour soak returned 35,999 HTTP 200 responses to 36,000 requests, with a 95th-percentile (p95) client latency of 142.751 ms. In a separate bounded AWS run of 900 s, the stream path received a burst of 400 records/s, and its consumer lag later drained to zero. Every number in the paper carries a label that says where it was measured, and the paper lists the parts of the design that were never built.
cs.CR / 2 / 2610.00695
Progressive-Resolution Secure Aggregation for Federated Learning
Abstract
Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the aggregate precision when clients upload. We introduce and formulate a new progressive-resolution secure-aggregation functionality in which clients upload once and successively finer resolutions of the same aggregate can later be authorized without renewed client participation. To realize this functionality, we propose progressive-resolution secure aggregation (PSA): each clipped, dithered update is represented by compatible nested-lattice digits; separately releasable layers are protected by secure aggregation and an additional aggregate pad that remains unavailable to the server until a non-colluding release controller authorizes that layer.
cs.CR / 3 / 2610.00759
Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents
Abstract
Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: https://anonymous.4open.science/r/RL-Transfer-between-env-4F47/.
cs.CR / 4 / 2610.00780
Made to Measure: Designing Image Watermarks to Specification
Abstract
Image watermarking supports provenance and attribution by embedding verifiable identity information into images. Practical deployments, however, must jointly satisfy requirements for attack resistance, false-positive rate (FPR), image quality, and latency. Existing watermarking methods are robust to different classes of transformations, so combining complementary methods can provide broader protection than any single watermark. Such composition is challenging, as additional fragments increase distortion and decoding cost and must share the same FPR budget. Therefore, we propose **TAILOR**, a request-conditioned watermark composition framework with three stages: (1) *offline characterization* measures fragment recovery, distortion, and runtime as response curves over embedding strength; (2) *joint configuration selection* encodes the request as an SMT model over these curves and solves for the lowest-distortion composition of fragments, strengths, order, and geometric recovery; and (3) *live calibration* validates the selected configuration on the user's images and refines predictions that fail to transfer. Experimental results across 7,321 distinct requests spanning five scenarios and 20 attack settings show that **TAILOR** achieves **96.21%** scenario-averaged request satisfaction with a mean PSNR of **41.02 dB**, outperforming existing methods in robustness while achieving consistently better image quality. Code is available at [https://github.com/aaFrostnova/Tailor](https://github.com/aaFrostnova/Tailor).
cs.CR / 5 / 2610.00787
Identity-Bound Governance Under Execution Uncertainty: An Accountability Proof Block for LLM Agent Persistent Halts, with Cryptographic Implementation and Cross-Model Calibration
Abstract
A correctly governed LLM agent can reach a state in which neither continuing execution nor automatically halting is admissible: the system has detected a persistent failure of its observability or drift-detection layer, but cannot itself decide who has the authority to resume, deny, or recalibrate the deployment. We call this an identity-bound governance event and formalise the mechanism that resolves it. We introduce the Accountability Proof Block (APB): a system-constructed evidence block, a human-supplied decision block, and an ed25519 signature binding both to a registered principal. The system cannot forge the signature, and the principal cannot alter the evidence undetected. We prove four theorems: Protocol-Bounded Governance Completeness, Non-Repudiability, Impossibility of Anonymous Re-Authorization, and Finite-Time APB Construction Termination. The implementation uses RFC 8785 JSON canonicalization and a UUID4-based replay predicate. Empirically, governance completeness holds over 3,812 halt events with zero unresolved cases; the verifier detects 100% of 1,800 attacks across 9 adversarial vectors; a k-of-n multi-principal variant shows 0 false acceptances in 2,000 single-key capture attempts. A study of six open LLMs finds the drift threshold T* stable within model (sigma/T* < 2%) but varying across models, refuting size-monotonicity: the largest model did not drift. T* must therefore be measured per deployment, and the APB is the vehicle by which that threshold yields accountable authority transfer.
cs.CR / 6 / 2610.00815
SafeDepth: Safety-Aware Token-Level Adaptive Computation
Abstract
Recent studies suggest that not every token needs to pass through all Transformer layers, motivating token-level adaptive models that selectively skip layers to reduce computation. Our experiments show that these execution choices also affect safety: existing token-level adaptive reasoning models exhibit higher harmful-response rates than their reference backbones. The safety effects depend on which layers are skipped and whether skipping occurs during prompt processing or answer generation. We further find that recognizing harmful requests and producing refusals depend on different layer-wise computations. We introduce SafeDepth, a lightweight, plug-in framework that uses selective layer execution to improve both safety and efficiency. SafeDepth learns to retain computations that support safe responses and bypass those that contribute to unsafe generation. A router selects layer execution based on token representations, layer position, and inference phase, while an adapter supports the skipped paths. We jointly train these modules to reduce computation and harmful responses while preserving performance on benign tasks, targeting a better safety-efficiency trade-off. The pretrained backbone remains frozen throughout, requiring no additional pretraining. Experimental results on Llama-3-8B-Instruct show that SafeDepth reduces computation relative to full-depth inference while largely preserving task performance. Compared with FlexiDepth, it lowers unsafe-response rates across all five harmful-request benchmarks, with the largest absolute decrease on HarmBench-HJ (from 55.44% to 25.25%), while reducing the XSTest false-refusal rate from 12.0% to 4.0%.
cs.CR / 7 / 2610.00850
AuraForge: Scaling Security Supervision for Training Coding Agents
Abstract
Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a critical risk: functional correctness alone does not guarantee a secure implementation. Despite growing attention to code security, training safer coding agents remains challenging because reliable security supervision is difficult to obtain at scale from real-world repositories. We introduce AuraForge to synthesize and validate executable security tests for training secure coding agents. Our approach combines attack-oriented test synthesis, language-extensible task construction, and safeguards against reward hacking. Using AuraForge, we construct AuraGym, a multi-language and multi-CWE executable training gym: 679 executable feature-implementation tasks from 344 real-world repositories across Python, JavaScript, and TypeScript, covering 177 CWE categories. On the subset with human-written security tests, AuraForge produces about 3 times as many test cases on average and reduces the false-positive rate by 83.23%, allowing alternative secure implementations to receive correct supervision. Training Qwen3.5-4B with synthesized security tests gains larger improvements than human-written security tests (average 19.7 FuncPass and 6.2 SecPass vs. 14.9 FuncPass and 4.4 SecPass) on three languages. These results demonstrate that AuraForge provides more diverse and reliable security supervision to train secure coding agents.
cs.CR / 8 / 2610.00977
ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications
Abstract
Broken access control, the failure of authorization, is one of the most prevalent web security risks. Unlike injection, a flow of untrusted input into a dangerous operation, authorization is a relation: who may act on what, not how data moves. Each application decides that relation for itself, so no rule written in advance carries to the next. An LLM agent can infer it from the code, but with no systematic way to cover the application and prioritize what to inspect, its search stays undirected and access-control flaws go undetected. We present ABSENTIA, a security scaffolding that turns general LLM agents into systematic vulnerability detectors for the backend of web applications, run as an audit by the developers and security engineers who maintain the code. Under its direction, the agents build a graph that maps the application's routes to the code behind them. ABSENTIA then works route by route, applying invariant falsification: it infers the properties the code is meant to satisfy, and where one is not enforced, reports the route for maintainer review. We also release BAC-Bench, a benchmark of 30 broken access control advisories across 25 repositories, 3 languages, and 9 frameworks, each published in 2025 or later, verified by a human auditor, and paired with its fixing commit, so credit requires flagging the vulnerable version and not the fixed one. ABSENTIA recalls 19 of them, 17 under paired credit, and an LLM verifier confirms 51% of its findings. CodeQL and Semgrep recall none, and an unstructured agent on the same model recalls 3. In the OWASP Benchmark injection categories, ABSENTIA leads the dedicated analyzers in Python and trails only CodeQL and IRIS in Java.
cs.CR / 9 / 2610.01009
Helol Tunnel: Covert Channel Exploitation of TLS Extensibility & Privacy Features
Abstract
Covert channels exploiting network protocols for data exfiltration and command-and-control (C2) are integral parts of modern cyberattacks. In search of a significant covert channel within the fabric of the Internet, we targeted the combinatorial properties of the Client Hello (CHLO) packets in the ubiquitous Transport Layer Security (TLS) protocol. The proposed Helol tunnel is a novel covert approach to embedding information in TLS Client Hello packets, which involves the strategic rearrangement of their cryptographic information elements. To sustain TLS protocol extensibility, the recent anti-ossification TLS compliance measures encourage the interactive middleboxes and next-generation firewalls (NGFWs) to preserve the parameter configuration in the Client Hello packets. Furthermore, to improve user privacy, popular Internet applications are varying their TLS CHLO parameter configurations to resist TLS fingerprinting by third-party network entities. We demonstrate the strength of the Helol tunnel to exploit these recent developments to evade NGFWs with interactive proxy and comprehensive threat protection. We also numerically show the efficacy of Helol tunneling over state-of-the-art covert channels that exploit TLS through the use of real traffic captures and public TLS fingerprinting data.
cs.CR / 10 / 2610.01184
ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning
Abstract
Remote multimodal models offer strong numerical reasoning capabilities over charts and speech, but sending private inputs risks exposing sensitive content. Text-only sanitization cannot directly satisfy fixed media interfaces, while identity anonymization leaves the underlying task content exposed. We introduce ReCast, an agentic plug-in framework that replaces source-specific content while preserving task-relevant relations and the required input modality. ReCast locally converts inputs into a shared textual evidence-query record, jointly rewrites entities and topics with a distilled 4B model, and substitutes values through a locally invertible, role-aware numerical map. A reconstruction agent generates and validates the required media from the protected record. The remote solver returns a program whose protected operands are restored locally before execution. On 4,000 held-out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of unprotected remote accuracy, while a model-based audit flags source-content leakage in 7.95% of solver-bound requests. It outperforms all evaluated local baselines, preserving the benefit of remote reasoning while reducing source-content exposure under existing media interfaces.
cs.CR / 11 / 2610.01250
A Resource-Aware Behavior Reconstruction and Hierarchical Semantic Learning Framework for Host Intrusion Detection
Abstract
System calls (syscalls) record key interactions between running programs and the operating system kernel, providing fine-grained and minimally intrusive data for host-based intrusion detection systems (HIDS) deployed in cloud and other modern computing environments. However, existing methods often model syscalls in their original execution order, where sequences from different processes are interleaved, making informative patterns difficult to extract and raising two questions: whether raw syscall sequences can be reorganized in a way that yields more discriminative representations, and how complex attack patterns can be effectively learned from the reorganized sequences. We propose ReSHID, a resource-aware behavior reconstruction and hierarchical semantic learning framework for host intrusion detection. It reconstructs semantically continuous sequences by leveraging syscall semantic invariants to cast subject identity and relationship resolution across PID namespaces as a bipartite matching problem and tracking file descriptor (FD) lifecycles to associate descriptors referring to the same resource. Additionally, features extracted from these sequences are organized into a lightweight subject behavior graph incorporating inter-subject relationships, where GATv2 captures key coordination patterns to model complex attacks involving multiple subjects. Experimental results show that sequence reconstruction combined with the detection method can improve HIDS performance. Even with a lightweight linear classifier, the proposed method achieves the best results among all compared methods in terms of F1-score (98.64%), ROC-AUC (99.80%), and PR-AUC (98.10%), while reducing the number of n-gram features by approximately 75.2% and 44.1% compared with the raw sequences and MGFE, respectively.
cs.CR / 12 / 2610.01294
GNSS Spoofing in Mobile Devices: A Survey on Impact and Countermeasures
Abstract
Smartphones rely on Global Navigation Satellite System (GNSS)-based positioning for many of the functions they execute everyday. The GNSS receivers embedded in smartphones are susceptible to anthropogenic radio frequency interference attacks in the forms of jamming and spoofing due to the low-power and open-architecture signals they receive from the satellite constellations. While jamming is a practice that denies a GNSS receiver the ability to form a position, velocity, and time (PVT) solution, spoofing represents a more insidious threat by using forged satellite signals that aim at causing the victim receiver to compute a false PVT solution. The ubiquity of smartphones and the sensitive geolocation data they hold make them a primary target for malicious spoofing. However, their hardware constraints and the lack of deep visibility into the GNSS receiver processing chain create significant hurdles for effective countermeasures. Existing surveys comprehensively explore general spoofing countermeasures but fail to address these mobile-specific limitations. This article fills that gap with a novel survey focused on techniques viable within the unique constraints of smartphone architectures. Specifically, we establish a taxonomy for defining GNSS spoofing attack effects and countermeasures, provide a historical review of smartphone vulnerability characterization, and provide an overview of techniques proposed to detect and counteract smartphone spoofing threats, offering a comparative framework to weigh their respective pros and cons on mobile platforms.
cs.CR / 13 / 2610.01300
A Systematization of Knowledge on DeFi Vaults: Architectures, Curation Mechanisms, and Strategy Design
Abstract
Decentralized finance (DeFi) vaults are smart-contract-based asset management systems that pool deposits, execute programmable strategies, and mint tokenized shares representing claims on underlying assets and strategy performance. As vault designs have evolved from early yield aggregators to modular, actively managed systems, a new control layer, curation, has emerged to select strategies, configure risk parameters, and coordinate operational execution, introducing principal-agent dynamics and new failure modes. This paper systematizes DeFi vault architectures and curator-mediated control planes through (i) a unified system model and formal definitions for share accounting, roles, and operational dependencies, and (ii) three complementary taxonomies covering vault exposures and objectives, curator governance and accountability mechanisms, and strategy execution patterns together with their failure modes. We further map a representative set of production protocols to the proposed dimensions. The frameworks in this work aim to support rigorous analysis and safer design of blockchain-based financial applications.
cs.CR / 14 / 2610.01385
Is it Possible to Generate Irreversible PolyProtected Templates from Face Embeddings using System-Specific Keys?
Abstract
This work aims to answer the question of whether it is possible to generate irreversible protected templates when the PolyProtect biometric template protection method is applied to face embeddings using system-specific keys (i.e., the same C and E parameters, which define the transform, are applied to all subjects' face embeddings), instead of the traditional subject-specific keys (i.e., each subject has their own C and E parameters). This is important for determining whether we can perform de-duplication of face identities in the PolyProtected domain, which is not possible in the subject-specific key scenario due to the clash with PolyProtect's unlinkability property (i.e., one could generate multiple protected templates belonging to the same identity, using different C and E parameters, such that those templates cannot be linked to each other). We present experiments (reproducible using our open-source code) to prove that there exist at least three ways of systematically selecting system-specific keys that produce irreversible PolyProtected templates: (i) from pre-selected subject-specific keys, (ii) by applying a previously proposed key selection algorithm to random vectors, and (iii) by approximating a "good" C/E pair distribution from which system-specific keys can be constructed. Our findings thus point to the conclusion that it is, indeed, possible to safely operate PolyProtect in the system-specific key scenario without degrading the template protection potential. This opens up the possibility for identity de-duplication in the PolyProtected domain.
cs.CR / 15 / 2610.01386
Evidence Coverage for Intent-Bound Execution: Scope, Obligations, and Cutoff Reasoning
Abstract
A verifier may authenticate every available record and still lack grounds to call an execution account complete. Such a claim requires a justified account of which records were due for the execution being assessed. We present an analytical model for retrospective coverage of declared execution-evidence obligations. Its scope binds a structured Intent, an exact Candidate, a selected analytical attempt, an execution and evidence boundary, a stage horizon, a fixed record-obligation profile, a named verifier, and an assessment cutoff. Branch and trigger premises determine obligation instances; source competence, content, integrity, and object and stage bindings determine admissibility. We distinguish closure of the obligation inventory from closure of the relevant verifier view, and define three reporting results: COMPLETE_WITHIN_SCOPE, INCOMPLETE, and UNKNOWN. These results concern current coverage of obligations due at the cutoff. Execution progress, external outcome knowledge, and historical delivery timeliness are reported separately. Constructed service-principal-disablement cases demonstrate complete dispatch and refusal branches, a due but missing final-result record, subsequent coverage after late delivery, and the limits of extending one selected attempt's coverage to all attempts. The contribution is an execution-specific composition of scope, branch, horizon, obligations, admissibility, view, and cutoff. Completeness remains conditional on the declared profile and assessment premises.
cs.CR / 16 / 2610.01484
Key-Reuse Vulnerability of Phase-Keyed Fourier-Curve Modulation: Relation Leakage and Key-Refresh Cost on Coded Links
Abstract
The security of keyed modulation is often argued from the key-space size and the error rate of a key-less receiver. This evidence fails when the key is reused and the waveform is harmonically coupled. For a phase-keyed Fourier-curve constellation, whose $k$ tones share one data parameter, integer relations among the harmonic indices yield data-cancelling mixed moments of the received tones that expose key characters. A modular relation lattice characterizes the exposed characters; for consecutive harmonics, third-order moments recover the relative phases and a fourth-order moment completes the key up to cyclic relabeling whenever its coefficient is nonzero, as in all evaluated settings. A non-data-aided relation-moment estimator turns this leakage into an attack that never enumerates the key space. On a regular $(3,6)$ LDPC-coded link, one key per 168-symbol codeword leaves the eavesdropper a block error rate below $0.04$ at the middle noise level, and the attack meets a predeclared $0.1$ compromise criterion in eleven of twelve operating points. Tangent artificial noise and a harmonic set without relations below order four raise her measured error rate at intermediate reuse lengths but do not remove the one-codeword vulnerability. For a grid of $2^{128}$ protocol keys at the middle noise level, equal-length refresh schedules that keep a $95\%$ lower confidence bound of her block error rate above $0.9$ consume at least $1.52$ fresh key bits per information bit, $1.52$ times the entropy rate of a one-time pad on the data.
cs.CR / 17 / 2610.01508
OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
Abstract
LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.
cs.CR / 18 / 2610.01535
False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift
Abstract
Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.
cs.CR / 19 / 2610.01564
Chaining Skills to Hijack LLM Agents
Abstract
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.
cs.CR / 20 / 2610.01580
Protocol Integration of Physical Layer Deception into EAP-TEAP Wi-Fi Authentication
Abstract
Credential-based Extensible Authentication Protocol (EAP) authentication cannot distinguish a legitimate credential holder from an adversary using compromised credentials. Physical Layer Deception (PLD) complements credential-based authentication by exposing a deceptive primary object over a primary transport while a separate recovery object travels with differentiated reliability over a secondary channel. Existing PLD studies remain, to our knowledge, at the physical/link-model level; using PLD's activation/deactivation mechanism as an authentication gate creates an authentication-specific design requirement, since an all-inactive attempt would exercise no recovery path. We present a batched PLD-based re-verification step for Enterprise Wi-Fi's TEAP/RADIUS/IEEE 802.11 authentication chain, implemented end to end across the server, access point, and device in the open-source hostap 2.12 codebase. Each attempt carries three rounds, at least one active, with no dedicated activation flag. Across four campaigns totaling 1593 attempts, the prototype evaluates batched recovery behavior, rejects the implemented naive credential-bearing attacker in all 30 attempts, measures successful-path latency, and evaluates the security-reliability trade-off for one, two, and three active rounds under two modeled recovery regimes. The evaluation exercises the protocol and software-MAC behavior directly and analyzes informed and retry-seeking attackers under the software recovery model.
cs.CR / 21 / 2610.01650
Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability
Abstract
The increasing prevalence of decentralized data has led to a growing interest in federated learning, which enables collaborative model training without clients sharing their sensitive local data. However, FL alone does not sufficiently protect sensitive training data and is generally coupled with privacy-preserving techniques, such as differential privacy and homomorphic encryption. Although powerful, these techniques address separate concerns via different mechanisms, so relying on just one might prove insufficient or impractical for addressing challenges associated with federated learning. In this work, we propose a privacy-preserving federated learning framework that combines homomorphic encryption-based training with differential privacy-based model inspection and release. We adopt a Markov chain Monte Carlo-based Bayesian privacy estimation method to estimate the privacy of our proposed framework. Our results show that this method improves both model utility and estimated privacy over the baseline method that relies solely on differential privacy for training. In our experiments with the FEMNIST dataset, by the end of training, our method reaches a test loss of $1.09$, compared to $2.37$ for the differential privacy-only approach, while providing stronger estimated privacy protection, with the estimated posterior mean of the privacy parameter $ε$ of $4.32$, compared to $7.26$ for the differential privacy-only approach. We also show that intermittent model monitoring can preserve the encrypted training trajectory while, under our evaluated experimental setting, providing estimated privacy comparable to or stronger than the differential privacy-only approach.
cs.CR / 22 / 2610.01736
The Achilles' Heel of Partial Reconfiguration: Optical Side-Channel Leakage on the 7-Series ICAP
Abstract
Major FPGA manufacturers have incorporated bitstream encryption to protect sensitive configuration data. However, for the most widely used FPGA families, multiple attacks against unpatchable protection schemes hard-wired into the devices can bypass or fully break them, making patchable schemes desirable. In this work, we present a proof-of-concept implementation of an AMD-proposed asymmetric key encryption scheme for bitstream protection for 7-Series FPGAs, using partial reconfiguration from the Programmable Logic. We analyze the security implications and hardware overhead of this implementation. We then propose and demonstrate an optical side-channel attack that is able to recover plain-text configuration data during the dynamic reconfiguration process. This attack leverages Photon Emission Microscopy and Electro-Optical Probing to first locate and then contactlessly extract the plain-text data from the ICAP interface, which internally connects the Programmable Logic with the configuration logic. We located the ICAP buses in an AMD XC7A200T device and show that the data on it can be extracted with Electro-Optical Probing. We claim that even advanced encryption schemes utilizing Partial Reconfiguration and custom cryptographic engines are vulnerable to optical attacks, as reconfiguration is only possible via hard-wired, vulnerable configuration interfaces.
cs.CR / 23 / 2610.01756
SoK: Decentralized Agent Economic Infrastructure
Abstract
Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides little evidence that the delivered work actually satisfied the task. We systematize this problem across the full lifecycle of an agent task. Our study organizes security and economic requirements into 17 property families over six stages, with receipt soundness and completeness assessed separately. We examine 12 systems and standards, five reusable mechanism families, and four classical baselines. We introduce guarantee closure, a task-relative criterion for determining whether guarantees established at one stage remain available and constrain the later decisions that depend on them. We apply the criterion to controlled and native workflows, covering 840 matched executions and an exhaustive 11,648-case check over a finite objective-task domain. Our results expose recurring failures between verification and settlement, where conforming work can remain unaccepted or valid evidence can be ignored. Public records and model judgments further distinguish recorded approval from evidence of task conformance, while economic analysis identifies the report, penalty, and shared-error assumptions behind these guarantees. These findings show where end-to-end guarantees fail and what must be repaired to preserve them across the workflow.
cs.CR / 24 / 2610.01768
The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching
Abstract
With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclosure of sensitive data to chatbot providers. While various solutions were developed to address these risks, including input structuring to prevent prompt injections or deploying local LLMs to avoid sharing confidential data with chatbot operators, LLMs also pose the risk of leaking confidential data to third parties. In this paper, we demonstrate with LLMLeak a novel attack vector where malicious software that runs locally but cannot communicate directly with the internet abuses LLMs to establish a covert channel. While inputs that instruct the LLM to send data directly via generated code are easy to detect and network libraries are typically restricted, LLMLeak relies only on the LLM's tool to fetch websites for further information. A malicious software component on the client side embeds a secret into a URL. It presents the referenced website as providing information required for a benign task, such as migrating a software library. When the LLM accesses the URL, the attacker receives the encoded secret through an attacker-controlled DNS or web server. We perform an extensive evaluation on eleven open-parameter models, observe an attack success rate of 79.7%, and also conduct a case study on real-world chatbots, demonstrating the relevance of LLMLeak.
cs.CR / 25 / 2610.01872
From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures
Abstract
While the literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for Internet of Things (IoT) and Industrial Internet of Things (IIoT) networks is mature, existing systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Response (EDR) and Extended Detection and Response (XDR) architectures, and they conflate blockchain's distinct functional roles into a single monolithic category. This Systematization of Knowledge (SoK) addresses these gaps by proposing a three-axis taxonomy that classifies proposals by detection-system class (NIDS, HIDS, EDR/XDR), blockchain functional role, and response-automation maturity. Synthesizing research published in high-impact venues between 2019 and 2026, we provide a rigorous gap analysis exposing why a genuine per-endpoint blockchain-anchored response loop remains nearly nonexistent due to latency, deployment, and community mismatches. Furthermore, we evaluate structural, cross-cutting challenges persisting across the literature, including consensus latency on constrained devices, post-quantum cryptographic vulnerability, smart-contract attack surfaces, and the adversarial vulnerability of evolving LLM-based detection engines. Finally, we outline a comprehensive research agenda centered on hybrid on-chain/off-chain orchestration to bridge the gap between decentralized trust and rapid response automation.
cs.CR / 26 / 2610.01893
A Structured State Space Sequence Model for Multi-Class Classification of Malware
Abstract
By 2030, Internet of Things (IoT) devices are projected to reach 40 billion, with fast-paced technological advancements in fields such as industry, healthcare, agriculture, automobiles, and building/home automation systems. This expansion has created a large attack surface for cybercrime, as the majority of these devices open the door for cybercriminals to exploit vulnerabilities, as they lack adequate built-in security. Cybercriminals launch malware attacks to compromise systems or steal sensitive data, and once a system is compromised, a ransom is typically demanded for its release. Current cybersecurity measures in place are being outpaced by the rapid growth of the IoT, which is accompanied by a subsequent growth in malware variants being created per day. Recognizing this pitfall, this research examines and proposes a novel approach to malware detection and classification to safeguard devices from further attacks and make IoT systems more robust and secure. The framework proposed utilizes a Structured State Space Sequence (S4) model, which discretizes sequences of malware samples in a sequence and captures long-range dependencies, essentially identifying the "cause" and "effect" hidden within malware execution flow. This study presents two novel contributions: the first empirical application of the S4 model for malware analysis, and a comprehensive comparison of its performance against other deep learning architectures, laying the stepping stone for future research in this new paradigm.
cs.CR / 27 / 2610.01907
Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler
Abstract
Differential privacy implementations rely on precise sampling from noise distributions to provide formal privacy guarantees. We report the discovery of systematic artifacts in OpenDP's discrete Laplace sampler that manifest as periodic distortions in the output distribution. Through systematic testing, we trace these artifacts to a faulty implementation in the rational arithmetic library used by the bernoulli_exp1 function, a low-level primitive that implements sampling from Bernoulli(e^(-x)) distributions. We present a diagnostic methodology that isolates the faulty component in the nested sampling hierarchy and propose an alternative implementation based on exact rational arithmetic that eliminates the artifacts. Statistical validation with 10^6 samples confirms that the corrected sampler produces outputs indistinguishable from the theoretical distribution at the tested precision level.
cs.CR / 28 / 2610.01949
A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders
Abstract
Ransomware has emerged as a major cybersecurity threat, with incidents increasing in frequency and impact across critical sectors. These attacks are typically launched through phishing emails, malicious downloads, or exploitation of software vulnerabilities to gain system access. Once inside, the malware encrypts files and demands a ransom, often in cryptocurrency, for the decryption key. Conventional detection methods often struggle with novel or scarce samples, leaving systems vulnerable. To address these challenges, this paper proposes a hybrid deep learning framework that combines an Autoencoder Feature Extractor (AFE) with a Model Agnostic Meta Learning (MAML) classifier for few shot malware detection. The AFE generates compact latent features that reduce noise and dimensionality, while the MAML classifier rapidly adapts to new threats using limited labeled data. Experiments conducted on the Ransomware Dataset 2024 demonstrate the effectiveness of the framework in binary classification tasks. Across one to fifty shot settings, the proposed model consistently achieves high accuracy, F1 score, and Matthews Correlation Coefficient values, maintaining reliable classification even under extreme scarcity. These results highlight the model's robustness and effectiveness in adapting to limited data scenarios, demonstrating the potential of combining feature extraction with meta learning to enhance resilience against malware, particularly in sectors such as healthcare, manufacturing, and public infrastructure, where cyberattacks can cause significant operational and financial disruption.
cs.CR / 29 / 2610.01960
System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7
Abstract
Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetic-kernel improvements, assembly tuning, register allocation, and instruction scheduling. Using the Module-Lattice-Based Key-Encapsulation Mechanism (ML-KEM) on an Arm Cortex-M7 as a case study, we examine the additional gains available from memory-hierarchy utilization, tightly coupled memory placement, peripheral integration, clock configuration, and deterministic public-data reuse. The evaluation starts from a state-of-the-art SLOTHY-optimized implementation and covers all three ML-KEM parameter sets. Without modifying the cryptographic algorithm or standardized wire formats, the evaluated profiles without auxiliary public state reduce cycles by up to 2.5%. A selected public-data-reuse profile reduces encapsulation and decapsulation cycles by up to 74.6% and 58.8%, respectively. These results demonstrate that substantial deployment gains remain after arithmetic-kernel optimization and motivate a two-stage methodology that also examines the surrounding execution system.
cs.CR / 30 / 2610.01924
Supersingularity and Superspeciality Verification of Abelian Surfaces
Abstract
Supersingular abelian surfaces are essential in isogeny-based cryptography. Despite this, we have no efficient algorithm to verify if a given abelian surface is supersingular. In this work, we initiate this research topic by giving an efficient Monte Carlo algorithm to verify if an abelian surface over $\mathbb{F}_p$ is supersingular in $O(\log p)$ with negligible failure probability, and an efficient conclusive algorithm if the order is smooth. We derive this algorithm by a careful analysis on the structure of supersingular Jacobians over $\mathbb{F}_p$. Furthermore, we derive efficient algorithms to verify if an abelian variety of any dimension is minimal or maximal, and to verify if a Jacobian of any dimension is superspecial.
cs.CR / 31 / 2610.00517
Blind Unforgeability implies Plus-One Unforgeability
Abstract
We prove that quantum blind-unforgeability implies quantum plus-one unforgeability, settling an open question in the literature. Since plus-one unforgeability is known not to imply blind-unforgeability, our result establishes that blind-unforgeability is \emph{strictly stronger} than plus-one unforgeability in the quantum setting. The implication is proven using a new \emph{smoothing} technique that coherently rescales the query histories of a quantum-query algorithm in a black-box manner. This enables us to transform any successful plus-one attacker into a blind-unforgeability attacker with only inverse-polynomial loss. Our proof represents a conceptual shift from merely extracting a useful query history as in earlier attempts, which we show inevitably fails in general, to first coherently reshaping how query histories interfere.
cs.CR / 32 / 2610.01025
Device-Independent Conference Keys from Parity-Extended Games
Abstract
Device-independent conference key agreement (DI-CKA) lets a group of parties establish a shared secret key from untrusted quantum devices, with security certified by non-locality. Existing DI-CKA protocols are each built around a single Bell inequality, typically a multiparty variant of the CHSH game. DI-QKD protocols, in contrast, have been built from a much richer landscape of non-local games, and it has remained unclear how to carry this landscape over to the conference setting. We introduce $\textit{Parity-$G$ games}$, which extend any two-player game $G$ to $N$ players, for every $N$, provided $G$ has an optimal strategy in which one player measures Pauli observables. The extension preserves the quantum and classical values of $G$, and the security of the resulting $N$-party protocol follows from an analysis of the two-player game alone. Our framework recovers the Parity-CHSH game of Ribeiro, Murta and Wehner (Phys. Rev. A, 2018) as a special case. Applied to the Mermin--Peres Magic Square Game, it yields a new $N$-player pseudo-telepathy game, the $\textit{Parity Magic Square Game}$, which ideal devices win in every round. We use it to construct the $\textit{first}$ DI-CKA protocol based on a pseudo-telepathy game. We prove the protocol secure against coherent attacks. It produces up to two key bits per round, and at low noise its key rate exceeds that of the DI-CKA protocol based on the Parity-CHSH game.
cs.CR / 33 / 2610.01421
Robust and leakage-resilient device-independent oblivious transfer in MiniQCryp
Abstract
Assuming post-quantum one-way functions, we construct device-independent (DI) oblivious transfer (OT) and bit commitment: honest parties use only trusted classical computation to operate untrusted quantum devices, which may share arbitrary entanglement and behave non-IID. Security is simulation-based against quantum polynomial-time adversaries and composes sequentially with efficient simulators. One protocol skeleton serves both, in two regimes. With isolated laboratories and coordinate-local measurements in the honest receiver's device, it tolerates a constant rate of honest-device faults. With polylogarithmically many qubits of adaptive leakage between the laboratories and arbitrary joint measurements, it tolerates an inverse-polylogarithmic rate. Each elementary DI call uses a fresh, isolated batch of polylogarithmically many device coordinates, and total device use in the compiled OT protocol is polynomial. The commitment has efficient simulators against both parties and yields DI coin tossing with abort. Because OT is complete for secure computation, the construction yields a DI protocol, with abort, for every efficiently computable classical functionality on a fixed number of parties, secure against static corruption of any proper subset of them. The commitment's extractor changes a public parity relation through classical equivocation and leaves the device execution, hence its leakage, unchanged. A commit-and-prove functionality, disjoint audits, and an affine consistency check link the certified correlations to ideal OT. Sender security rests on a selector-aware parallel-repetition bound for the Magic Square game, which we derive from the two-round threshold theorem of Kundu and Tan.
cs.CR / 34 / 2610.01777
QUFIG: GNN-Based Prediction of Quantum Fault Injection Vulnerabilities with Gate-Level Precision
Abstract
The growing scale and accessibility of quantum hardware exposed new reliability and security challenges in the quantum computing workflow, such as the run-time fault injection attacks in cloud-based quantum computing platforms. However, existing works fail to identify vulnerabilities with gate-level precision or adapt to run-time environments. In this work, we formulate gate-level fault analysis as a learning-guided prioritization problem under restricted fidelity budgets. The framework uses a circuit-DAG-based GNN backbone to predict the vulnerability score of each gate to each type of injected fault, defined as the impact of the gate-fault pair on circuit fidelity. The gate-fault pairs are then ranked by their vulnerability score. Experiments on QASMbench and HamLib MaxCut show that QUFIG recovers high-impact vulnerable gate-fault pairs with fewer inspections than random and depth-based heuristics. Our results show that QUFIG can reduce the number of gates requiring inspection by 2.9--19.8% while maintaining effective fault identification, allowing quantum circuit designers to identify vulnerabilities and apply targeted defenses more efficiently.
cs.CR / 35 / 2610.01792
Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift
Abstract
Quantum key distribution (QKD) links are provisioned from security analyses of stationary channels, whereas the devices that determine the channel drift between recalibrations. Whether an eavesdropper who cannot alter the channel's own noise gains by following that drift has not been quantified. Adaptive eavesdropping is posed here as a constrained Markov decision process in which the attacker selects one circuit per round while the noise level follows an Ornstein--Uhlenbeck process and the abort condition is a budget over each block of rounds. The value of adaptation is bounded by the best fixed circuit and a dynamic-programming upper bound. The actions are learnt attacks. Whereas Decker et al. trained a parametrised circuit on a fixed gate template against a fixed channel, here the gate structure and rotation angles are searched jointly. This yields circuits compact enough to form a discrete action set, extending the construction to noise models lacking a known template, including the amplitude damping channel. On device-independent E91 under bilateral depolarising noise, a reinforcement-learning attacker raises her Holevo information from $0.135$ for the best fixed circuit to $0.348$ at zero detection, $98\%$ of the upper bound. On BB84 under a drifting bit-flip channel, she exceeds a conservative noise-indexed rule by $0.024$ in fidelity, reaching $99\%$ of the upper bound. Under stationary noise, the attacker's gain from basis asymmetry changes sign between an averaged and a per-basis error-rate constraint. The search, started from random gate sequences, recovers the analytical cloners and the collective-attack key rate, and meets the lower bound of the Winick--Lütkenhaus--Coles objective from above.
cs.CR / 36 / 2610.01848
Trapdoored Clifford Operators and Applications
Abstract
Random Clifford operators have numerous applications in quantum computing, including randomized benchmarking, classical shadows, and quantum authentication. However, sampling and implementing uniformly random $n$-qubit Clifford incur near-quadratic complexity due to the size of Clifford group. We introduce a cryptographic way to overcome these barriers: trapdoored Clifford operator distributions whose samples are computationally indistinguishable from uniformly random Cliffords, yet implementing them can be much faster given the trapdoor. We construct a distribution of trapdoored Clifford operators whose elements can be sampled and implemented in near-linear time under a variant of the learning parity with noise assumption. Our constructions allow fast tableau action on Pauli labels for classical simulation, and also can be optimized to admit polylogarithmic-depth implementation. Along the way, we construct trapdoored matrices over finite fields that support efficient multiplication by both a matrix and its inverse, resolving an open question left by Vaikuntanathan and Zamir [SODA'26]. We use these constructions to obtain faster protocols based on random Cliffords. We also explore their applications to the worst-case to average-case reductions for matrix and Clifford problems including the iterated matrix multiplication and Clifford circuit synthesis. In particular, we show the hardness of batching Clifford circuits: synthesizing circuits that apply the same Clifford to multiple registers is at least as hard as worst-case matrix multiplication, even when synthesis succeeds on a small constant fraction of random Cliffords. This extends to approximate implementations by general quantum circuits.
cs.CR / 37 / 2610.02100
On the pseudorandomness of simple quantum processes
Abstract
Can simple processes appear highly complex? Gowers (Comb. Prob. Comp. '96) conjectured that repeatedly composing local random reversible operations can yield global permutations that are indistinguishable from random. In this work, we study the unitary quantum analog of this question, in an attempt to make new progress on this longstanding conjecture. Our first result shows that statistical moment matching in the form of unitary designs does not generically lead to pseudorandomness---even for the simplest quantum processes: for every fixed $t$, we give an efficiently samplable family $\{ν_n\}_n$ of distributions on one- and two-qubit gates such that, after $T=O_t(n^2\log^2 n)$ independent steps, the resulting $n$-qubit ensemble is an approximate unitary $t$-design with negligible error $\exp(-Ω(\log^2 n))$, yet an efficient quantum algorithm distinguishes it from random using only $O_t(\log^2 n)$ queries. This refutes the unitary analog of the Hoory--Magen--Myers--Rackoff conjecture (ICALP '04) for permutations. Our second result is a stronger separation between unitary designs and pseudorandom unitaries at polynomially bounded moments; our counterexample, however, requires highly structured ensembles, in contrast with the simple local walks from before. This suggests caution when using unitary designs to model information scrambling in black-hole physics, as even maximally scrambled systems can exhibit structure which is accessible to efficient experiments. Motivated by these findings, we then propose new conjectures for how pseudorandomness can plausibly emerge within simple quantum processes, such as random quantum circuits.
cs.CR / 38 / 2610.02101
Time-space lower bounds for breaking quantum cryptography
Abstract
We prove near-optimal time-space lower bounds for breaking quantum cryptography in the random oracle model. Specifically, we show that a $T$-query adversary with $S$ qubits of non-uniform advice can recover a random key $k$ from the $n$-qubit binary phase state $|ψ_k\rangle \propto \sum_{x} R(k,x) |x\rangle$ with probability at most $O(\frac{T^2 + \sqrt{ST}}{N})$ for $N=2^n$. In contrast, the best known bound for post-quantum one-way functions is $O(\frac{T^2 + ST}{N})$, with a trivial attack at $S = N$. This demonstrates a new advantage of quantum cryptography over classical cryptography: $n$ qubits of communication suffice for security against preprocessing attacks with space up to $N^2$ rather than $N$. Our methodology is simple: express the optimal preprocessing attack as the operator norm of a random matrix, and bound this value in expectation over the random oracle via the trace-moment method. These trace moments have a natural interpretation using compressed oracles [Zhandry, Crypto 2019], which we then analyze. This can be viewed as a simplification and generalization of the approach of Liu [Eurocrypt 2023] for proving time-space tradeoffs for breaking post-quantum cryptography. We also prove the following results: (1) We tighten Liu's analysis of post-quantum PRGs in QROM, achieving a distinguishing advantage bound of $O(\frac{T^2}N + \sqrt{\frac{ST}N})$. (2) For unitary synthesis, we extend the one-query lower bound of Lombardi-Ma-Wright [STOC 2024] to hold against adversaries that can make one arbitrary function query along with polynomially many (adaptive) queries to the random oracle, either before or after the function query. This also interprets the original LMW24 result in terms of compressed oracles. (3) Finally, we prove a tight $O(\frac{\sqrt{S}}N)$ bound for the pseudorandomness of random binary phase states against space $S$ distinguishers.
cs.CR / 39 / 2610.02113
Quantum Advantage for Two-Party Differential Privacy
Abstract
We introduce information-theoretically private quantum protocols for two-party Hamming distance when both parties must output the same estimate. Classically, for input length $n$, information-theoretic protocols require $Ω(\sqrt{n})$ error under pure differential privacy and $Ω(\sqrt{n}/\log n)$ error under strong approximate differential privacy, whereas computational security permits $O(1)$ error. In Klauck's honest, nonpreemptive, message-preserving model, we give an $O(n)$-communication quantum protocol with pure $\varepsilon$ quantum differential privacy (QDP) and expected error at most $\frac{2}{\sinh \varepsilon}+γ$, for every $γ>0$. For approximate $(\varepsilon, δ)$ QDP, an exact hockey-stick divergence calculation yields strictly smaller error, while preserving the $O(1)$-versus-$Ω(\sqrt{n}/\log n)$ separation for $δ=o(1/n)$. Thus, quantum communication achieves $O(1)$ information-theoretic error, matching the accuracy available classically only under computational assumptions. The main construction uses a guarded coherent round trip and an equal-Gram rigidity principle that prevents an honest player from retaining input-dependent complementary information. We separate this model from weaker prescribed-channel privacy, which already admits an exact classical realization, and from fully retention-robust security, against which measurement-and-abort attacks remain possible. Therefore, we identify preservation of non-orthogonal quantum messages as a resource for privacy.